Instructions to use script-langid/fasttext-leaf-surgery-11lang with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-11lang with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-11lang", "model.bin")) - Notebooks
- Google Colab
- Kaggle
FastText Continual Leaf Surgery (11-Language Rehearsal)
This model is a continual fine-tuning of Meta's fastText LID-176 utilizing Hierarchical Softmax Leaf Surgery.
Architecture & Innovation
- Base Architecture: Meta
fastText LID-176(100-dim dense subword embeddings with Huffman tree hierarchical softmax). - Leaf Surgery: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (
si) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation. - Continual Rehearsal: Trained on the 11-language uniform rehearsal buffer (
train_11lang_uniform.csv) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari).
Benchmark Performance
| Benchmark | Macro F1 | Sinhala F1 | Pali F1 | Sanskrit F1 | Overall Accuracy |
|---|---|---|---|---|---|
| WiLI-2018 (11-Lang) | 0.9811 | 0.9766 | 0.9881 | 0.9895 | 98.00% |
| CommonLID (11-Lang) | 0.9486 | 0.9766 | 0.9881 | 0.9695 | 96.82% |
| FLORES+ (11-Lang) | 0.9395 | 0.9766 | 0.9881 | 0.9817 | 92.71% |
How to Load in Python
from data_pipeline.fasttext_continual.model import ContinualLID
# Loads config.json, vocab.json, and weights.pt automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")
# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction = model.predict(text)
print(prediction) # 'pi' (Pali)
Repository & Research
- GitHub: Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit
- Research: University of Moratuwa, Department of Computer Science & Engineering.
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support