fastText
language-identification
continual-learning
leaf-surgery
sinhala
pali
sanskrit

FastText Continual Leaf Surgery (11-Language Rehearsal)

This model is a continual fine-tuning of Meta's fastText LID-176 utilizing Hierarchical Softmax Leaf Surgery.

Architecture & Innovation

  • Base Architecture: Meta fastText LID-176 (100-dim dense subword embeddings with Huffman tree hierarchical softmax).
  • Leaf Surgery: Instead of retraining the classification head from scratch (which shuffles the Huffman tree and causes catastrophic forgetting of 176 pre-trained languages), the original hierarchical softmax binary decision paths are preserved. The Sinhala (si) leaf node is surgically split into an internal decision node branching into modern Sinhala and canonical Pali, with Sanskrit adaptation.
  • Continual Rehearsal: Trained on the 11-language uniform rehearsal buffer (train_11lang_uniform.csv) balancing the 3 target languages in Sinhala script with 8 global and regional anchor languages (English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit in Devanagari).

Benchmark Performance

Benchmark Macro F1 Sinhala F1 Pali F1 Sanskrit F1 Overall Accuracy
WiLI-2018 (11-Lang) 0.9811 0.9766 0.9881 0.9895 98.00%
CommonLID (11-Lang) 0.9486 0.9766 0.9881 0.9695 96.82%
FLORES+ (11-Lang) 0.9395 0.9766 0.9881 0.9817 92.71%

How to Load in Python

from data_pipeline.fasttext_continual.model import ContinualLID

# Loads config.json, vocab.json, and weights.pt automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")

# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction = model.predict(text)
print(prediction)  # 'pi' (Pali)

Repository & Research

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support