MaleeshaK's picture
Upload README.md with huggingface_hub
ffa4865 verified
|
Raw History Blame Contribute Delete
1.97 kB
metadata
license: cc-by-nc-4.0
language:
  - sin
  - pli
  - san
tags:
  - language-identification
  - fasttext
  - continual-learning
  - leaf-surgery
  - table-2-target-only
  - sinhala
  - pali
  - sanskrit
pipeline_tag: text-classification
base_model: facebook/fasttext-language-identification

fastText Continual Leaf Surgery: Target-Only Training (Table 2)

This model represents our Continual Learning (Leaf Surgery) architecture trained strictly on the target dataset (Table 2 in the research paper):

"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"
University of Moratuwa, Department of Computer Science & Engineering

1. Architecture & Innovation

Instead of fine-tuning with a replacement flat softmax head (which ruins fastText's Huffman tree and causes catastrophic forgetting), our method surgically splits the Sinhala (si) leaf node in the binary Huffman tree into modern Sinhala and canonical Pali branches.

All original 176 language decision nodes are retained intact, and the model is trained strictly on target text (train.csv).

2. Checkpoint Files

  • weights.pt: PyTorch weights containing inherited input embedding and expanded output hierarchical decision parameters (124.5 MB).
  • config.json: Tree paths, Huffman bit codes, and label mapping.
  • vocab.json: Subword vocabulary and hashes.

3. How to Load and Run Inference in Python

from data_pipeline.fasttext_continual.model import ContinualLID

# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-target-only")

# Inference example:
text = "මනොපුබ්බඞ්ගමා ධම්මා මනොසෙට්ඨා මනොමයා"
prediction, score = model.predict(text)
print(f"Language: {prediction}, Score: {score:.4f}")
# Output: Language: pi, Score: 0.99...