Text Classification
fastText
Sinhala
Pali
Sanskrit
language-identification
continual-learning
leaf-surgery
table-2-target-only
sinhala
pali
sanskrit
Instructions to use script-langid/fasttext-leaf-surgery-target-only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-target-only with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-target-only", "model.bin")) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from script-langid/fasttext-leaf-surgery-target-only: direct link, hf CLI and curl.
- Browser
- Download file 1.97 kB
-
https://huggingface.co/script-langid/fasttext-leaf-surgery-target-only/resolve/main/README.md
- Command line
-
hf download hf://script-langid/fasttext-leaf-surgery-target-only/README.md
-
curl -L -o README.md https://huggingface.co/script-langid/fasttext-leaf-surgery-target-only/resolve/main/README.md
1.97 kB
| license: cc-by-nc-4.0 | |
| language: | |
| - sin | |
| - pli | |
| - san | |
| tags: | |
| - language-identification | |
| - fasttext | |
| - continual-learning | |
| - leaf-surgery | |
| - table-2-target-only | |
| - sinhala | |
| - pali | |
| - sanskrit | |
| pipeline_tag: text-classification | |
| base_model: facebook/fasttext-language-identification | |
| # fastText Continual Leaf Surgery: Target-Only Training (Table 2) | |
| This model represents our **Continual Learning (Leaf Surgery)** architecture trained strictly on the target dataset (Table 2 in the research paper): | |
| > **"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"** | |
| > *University of Moratuwa, Department of Computer Science & Engineering* | |
| ## 1. Architecture & Innovation | |
| Instead of fine-tuning with a replacement flat softmax head (which ruins fastText's Huffman tree and causes catastrophic forgetting), our method surgically splits the Sinhala (`si`) leaf node in the binary Huffman tree into modern Sinhala and canonical Pali branches. | |
| All original 176 language decision nodes are retained intact, and the model is trained strictly on target text (`train.csv`). | |
| ## 2. Checkpoint Files | |
| - `weights.pt`: PyTorch weights containing inherited input embedding and expanded output hierarchical decision parameters (124.5 MB). | |
| - `config.json`: Tree paths, Huffman bit codes, and label mapping. | |
| - `vocab.json`: Subword vocabulary and hashes. | |
| ## 3. How to Load and Run Inference in Python | |
| ```python | |
| from data_pipeline.fasttext_continual.model import ContinualLID | |
| # Loads weights.pt, config.json, and vocab.json automatically from Hugging Face | |
| model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-target-only") | |
| # Inference example: | |
| text = "මනොපුබ්බඞ්ගමා ධම්මා මනොසෙට්ඨා මනොමයා" | |
| prediction, score = model.predict(text) | |
| print(f"Language: {prediction}, Score: {score:.4f}") | |
| # Output: Language: pi, Score: 0.99... | |
| ``` | |