Text Classification
fastText
language-identification
continual-learning
leaf-surgery
table-3-multilingual-rehearsal
Instructions to use script-langid/fasttext-leaf-surgery-11lang with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use script-langid/fasttext-leaf-surgery-11lang with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("script-langid/fasttext-leaf-surgery-11lang", "model.bin")) - Notebooks
- Google Colab
- Kaggle
File size: 2,015 Bytes
7225712 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 | ---
license: cc-by-nc-4.0
language:
- sin
- pli
- san
- en
- ta
- hi
- bn
- ar
- fr
- de
tags:
- language-identification
- fasttext
- continual-learning
- leaf-surgery
- table-3-multilingual-rehearsal
pipeline_tag: text-classification
base_model: facebook/fasttext-language-identification
---
# fastText Continual Leaf Surgery: 11-Language Multilingual Rehearsal (Table 3)
This model represents our **Continual Learning (Leaf Surgery)** architecture trained with the 11-language balanced rehearsal buffer (Table 3 in the research paper):
> **"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"**
> *University of Moratuwa, Department of Computer Science & Engineering*
## 1. Architecture & Innovation
- **Hierarchical Softmax Leaf Surgery**: The Sinhala (`si`) leaf node in the official Facebook `lid.176.bin` Huffman tree was split into Sinhala and Pali branches.
- **Multilingual Rehearsal**: Fine-tuned on the balanced 11-language Aya dataset buffer (3 target languages in Sinhala script + 8 global/anchor languages: English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit Devanagari) to preserve cross-lingual representations.
## 2. Benchmark Performance (11-Language Evaluation)
- **WiLI-2018 (Macro-F1):** **0.9811** (Accuracy: 98.00%)
- **CommonLID (Macro-F1):** **0.9486** (Accuracy: 96.82%)
- **FLORES+ (Macro-F1):** **0.9395** (Accuracy: 92.71%)
## 3. How to Load and Run Inference in Python
```python
from data_pipeline.fasttext_continual.model import ContinualLID
# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")
# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction, score = model.predict(text)
print(f"Language: {prediction}, Score: {score:.4f}")
# Output: Language: pi, Score: 0.99...
```
|