fastText
language-identification
continual-learning
leaf-surgery
sinhala
pali
sanskrit
MaleeshaK's picture
Upload folder using huggingface_hub
bc3cca2 verified
|
Raw History Blame Contribute Delete
2.53 kB
---
language:
- sin
- pli
- san
- en
- ta
- hi
- bn
- ar
- fr
- de
license: cc-by-4.0
tags:
- language-identification
- fasttext
- continual-learning
- leaf-surgery
- sinhala
- pali
- sanskrit
datasets:
- script-langid/sinhala-pali-sanskrit-target
metrics:
- f1
- accuracy
---
# FastText Continual Leaf Surgery (11-Language Rehearsal)
This model is a continual fine-tuning of Meta's `fastText LID-176` utilizing **Hierarchical Softmax Leaf Surgery**.
## Architecture & Innovation
- **Base Architecture**: Meta `fastText LID-176` (100-dim dense subword embeddings with Huffman tree hierarchical softmax).
- **Leaf Surgery**: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (`si`) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation.
- **Continual Rehearsal**: Trained on the 11-language uniform rehearsal buffer (`train_11lang_uniform.csv`) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari).
## Benchmark Performance
| Benchmark | Macro F1 | Sinhala F1 | Pali F1 | Sanskrit F1 | Overall Accuracy |
|---|:---:|:---:|:---:|:---:|:---:|
| **WiLI-2018 (11-Lang)** | **0.9811** | 0.9766 | 0.9881 | 0.9895 | **98.00%** |
| **CommonLID (11-Lang)** | **0.9486** | 0.9766 | 0.9881 | 0.9695 | **96.82%** |
| **FLORES+ (11-Lang)** | **0.9395** | 0.9766 | 0.9881 | 0.9817 | **92.71%** |
## How to Load in Python
```python
from data_pipeline.fasttext_continual.model import ContinualLID
# Loads config.json, vocab.json, and weights.pt automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")
# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction = model.predict(text)
print(prediction) # 'pi' (Pali)
```
## Repository & Research
- **GitHub**: [Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit](https://github.com/Maleesha-K/Sinhala-Script-Language-Identification-LangID-for-Sinhala-Pali-and-Sanskrit)
- **Research**: University of Moratuwa, Department of Computer Science & Engineering.