--- language: - sin - pli - san - en - ta - hi - bn - ar - fr - de license: cc-by-4.0 tags: - language-identification - fasttext - continual-learning - leaf-surgery - sinhala - pali - sanskrit datasets: - script-langid/sinhala-pali-sanskrit-target metrics: - f1 - accuracy --- # FastText Continual Leaf Surgery (11-Language Rehearsal) This model is a continual fine-tuning of Meta's `fastText LID-176` utilizing **Hierarchical Softmax Leaf Surgery**. ## Architecture & Innovation - **Base Architecture**: Meta `fastText LID-176` (100-dim dense subword embeddings with Huffman tree hierarchical softmax). - **Leaf Surgery**: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (`si`) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation. - **Continual Rehearsal**: Trained on the 11-language uniform rehearsal buffer (`train_11lang_uniform.csv`) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari). ## Benchmark Performance | Benchmark | Macro F1 | Sinhala F1 | Pali F1 | Sanskrit F1 | Overall Accuracy | |---|:---:|:---:|:---:|:---:|:---:| | **WiLI-2018 (11-Lang)** | **0.9811** | 0.9766 | 0.9881 | 0.9895 | **98.00%** | | **CommonLID (11-Lang)** | **0.9486** | 0.9766 | 0.9881 | 0.9695 | **96.82%** | | **FLORES+ (11-Lang)** | **0.9395** | 0.9766 | 0.9881 | 0.9817 | **92.71%** | ## How to Load in Python ```python from data_pipeline.fasttext_continual.model import ContinualLID # Loads config.json, vocab.json, and weights.pt automatically from Hugging Face model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang") # Inference example: text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස" prediction = model.predict(text) print(prediction) # 'pi' (Pali) ``` ## Repository & Research - **GitHub**: [Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit](https://github.com/Maleesha-K/Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit) - **Research**: University of Moratuwa, Department of Computer Science & Engineering.