Instructions to use script-langid/fasttext-leaf-surgery-11lang with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-11lang with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-11lang", "model.bin")) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from script-langid/fasttext-leaf-surgery-11lang: direct link, hf CLI and curl.
- Browser
- Download file 2.53 kB
-
https://huggingface.co/script-langid/fasttext-leaf-surgery-11lang/resolve/main/README.md
- Command line
-
hf download hf://script-langid/fasttext-leaf-surgery-11lang/README.md
-
curl -L -o README.md https://huggingface.co/script-langid/fasttext-leaf-surgery-11lang/resolve/main/README.md
2.53 kB
| language: | |
| - sin | |
| - pli | |
| - san | |
| - en | |
| - ta | |
| - hi | |
| - bn | |
| - ar | |
| - fr | |
| - de | |
| license: cc-by-4.0 | |
| tags: | |
| - language-identification | |
| - fasttext | |
| - continual-learning | |
| - leaf-surgery | |
| - sinhala | |
| - pali | |
| - sanskrit | |
| datasets: | |
| - script-langid/sinhala-pali-sanskrit-target | |
| metrics: | |
| - f1 | |
| - accuracy | |
| # FastText Continual Leaf Surgery (11-Language Rehearsal) | |
| This model is a continual fine-tuning of Meta's `fastText LID-176` utilizing **Hierarchical Softmax Leaf Surgery**. | |
| ## Architecture & Innovation | |
| - **Base Architecture**: Meta `fastText LID-176` (100-dim dense subword embeddings with Huffman tree hierarchical softmax). | |
| - **Leaf Surgery**: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (`si`) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation. | |
| - **Continual Rehearsal**: Trained on the 11-language uniform rehearsal buffer (`train_11lang_uniform.csv`) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari). | |
| ## Benchmark Performance | |
| | Benchmark | Macro F1 | Sinhala F1 | Pali F1 | Sanskrit F1 | Overall Accuracy | | |
| |---|:---:|:---:|:---:|:---:|:---:| | |
| | **WiLI-2018 (11-Lang)** | **0.9811** | 0.9766 | 0.9881 | 0.9895 | **98.00%** | | |
| | **CommonLID (11-Lang)** | **0.9486** | 0.9766 | 0.9881 | 0.9695 | **96.82%** | | |
| | **FLORES+ (11-Lang)** | **0.9395** | 0.9766 | 0.9881 | 0.9817 | **92.71%** | | |
| ## How to Load in Python | |
| ```python | |
| from data_pipeline.fasttext_continual.model import ContinualLID | |
| # Loads config.json, vocab.json, and weights.pt automatically from Hugging Face | |
| model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang") | |
| # Inference example: | |
| text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස" | |
| prediction = model.predict(text) | |
| print(prediction) # 'pi' (Pali) | |
| ``` | |
| ## Repository & Research | |
| - **GitHub**: [Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit](https://github.com/Maleesha-K/Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit) | |
| - **Research**: University of Moratuwa, Department of Computer Science & Engineering. | |