Instructions to use script-langid/fasttext-leaf-surgery-11lang with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-11lang with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-11lang", "model.bin")) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from script-langid/fasttext-leaf-surgery-11lang: direct link, hf CLI and curl.
- Browser
- Download file 2.53 kB
-
https://huggingface.co/script-langid/fasttext-leaf-surgery-11lang/resolve/main/README.md
- Command line
-
hf download hf://script-langid/fasttext-leaf-surgery-11lang/README.md
-
curl -L -o README.md https://huggingface.co/script-langid/fasttext-leaf-surgery-11lang/resolve/main/README.md
2.53 kB
metadata
language:
- sin
- pli
- san
- en
- ta
- hi
- bn
- ar
- fr
- de
license: cc-by-4.0
tags:
- language-identification
- fasttext
- continual-learning
- leaf-surgery
- sinhala
- pali
- sanskrit
datasets:
- script-langid/sinhala-pali-sanskrit-target
metrics:
- f1
- accuracy
FastText Continual Leaf Surgery (11-Language Rehearsal)
This model is a continual fine-tuning of Meta's fastText LID-176 utilizing Hierarchical Softmax Leaf Surgery.
Architecture & Innovation
- Base Architecture: Meta
fastText LID-176(100-dim dense subword embeddings with Huffman tree hierarchical softmax). - Leaf Surgery: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (
si) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation. - Continual Rehearsal: Trained on the 11-language uniform rehearsal buffer (
train_11lang_uniform.csv) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari).
Benchmark Performance
| Benchmark | Macro F1 | Sinhala F1 | Pali F1 | Sanskrit F1 | Overall Accuracy |
|---|---|---|---|---|---|
| WiLI-2018 (11-Lang) | 0.9811 | 0.9766 | 0.9881 | 0.9895 | 98.00% |
| CommonLID (11-Lang) | 0.9486 | 0.9766 | 0.9881 | 0.9695 | 96.82% |
| FLORES+ (11-Lang) | 0.9395 | 0.9766 | 0.9881 | 0.9817 | 92.71% |
How to Load in Python
from data_pipeline.fasttext_continual.model import ContinualLID
# Loads config.json, vocab.json, and weights.pt automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")
# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction = model.predict(text)
print(prediction) # 'pi' (Pali)
Repository & Research
- GitHub: Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit
- Research: University of Moratuwa, Department of Computer Science & Engineering.