Text Classification
fastText
Sinhala
Pali
Sanskrit
language-identification
continual-learning
leaf-surgery
table-2-target-only
sinhala
pali
sanskrit
Instructions to use script-langid/fasttext-leaf-surgery-target-only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-target-only with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-target-only", "model.bin")) - Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: cc-by-nc-4.0
|
| 3 |
+
language:
|
| 4 |
+
- sin
|
| 5 |
+
- pli
|
| 6 |
+
- san
|
| 7 |
+
tags:
|
| 8 |
+
- language-identification
|
| 9 |
+
- fasttext
|
| 10 |
+
- continual-learning
|
| 11 |
+
- leaf-surgery
|
| 12 |
+
- table-2-target-only
|
| 13 |
+
- sinhala
|
| 14 |
+
- pali
|
| 15 |
+
- sanskrit
|
| 16 |
+
pipeline_tag: text-classification
|
| 17 |
+
base_model: facebook/fasttext-language-identification
|
| 18 |
+
---
|
| 19 |
+
|
| 20 |
+
# fastText Continual Leaf Surgery: Target-Only Training (Table 2)
|
| 21 |
+
|
| 22 |
+
This model represents our **Continual Learning (Leaf Surgery)** architecture trained strictly on the target dataset (Table 2 in the research paper):
|
| 23 |
+
> **"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"**
|
| 24 |
+
> *University of Moratuwa, Department of Computer Science & Engineering*
|
| 25 |
+
|
| 26 |
+
## 1. Architecture & Innovation
|
| 27 |
+
Instead of fine-tuning with a replacement flat softmax head (which ruins fastText's Huffman tree and causes catastrophic forgetting), our method surgically splits the Sinhala (`si`) leaf node in the binary Huffman tree into modern Sinhala and canonical Pali branches.
|
| 28 |
+
|
| 29 |
+
All original 176 language decision nodes are retained intact, and the model is trained strictly on target text (`train.csv`).
|
| 30 |
+
|
| 31 |
+
## 2. Checkpoint Files
|
| 32 |
+
- `weights.pt`: PyTorch weights containing inherited input embedding and expanded output hierarchical decision parameters (124.5 MB).
|
| 33 |
+
- `config.json`: Tree paths, Huffman bit codes, and label mapping.
|
| 34 |
+
- `vocab.json`: Subword vocabulary and hashes.
|
| 35 |
+
|
| 36 |
+
## 3. How to Load and Run Inference in Python
|
| 37 |
+
```python
|
| 38 |
+
from data_pipeline.fasttext_continual.model import ContinualLID
|
| 39 |
+
|
| 40 |
+
# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
|
| 41 |
+
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-target-only")
|
| 42 |
+
|
| 43 |
+
# Inference example:
|
| 44 |
+
text = "මනොපුබ්බඞ්ගමා ධම්මා මනොසෙට්ඨා මනොමයා"
|
| 45 |
+
prediction, score = model.predict(text)
|
| 46 |
+
print(f"Language: {prediction}, Score: {score:.4f}")
|
| 47 |
+
# Output: Language: pi, Score: 0.99...
|
| 48 |
+
```
|