MaleeshaK commited on
Commit
ffa4865
·
verified ·
1 Parent(s): 0adc755

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +48 -0
README.md ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-nc-4.0
3
+ language:
4
+ - sin
5
+ - pli
6
+ - san
7
+ tags:
8
+ - language-identification
9
+ - fasttext
10
+ - continual-learning
11
+ - leaf-surgery
12
+ - table-2-target-only
13
+ - sinhala
14
+ - pali
15
+ - sanskrit
16
+ pipeline_tag: text-classification
17
+ base_model: facebook/fasttext-language-identification
18
+ ---
19
+
20
+ # fastText Continual Leaf Surgery: Target-Only Training (Table 2)
21
+
22
+ This model represents our **Continual Learning (Leaf Surgery)** architecture trained strictly on the target dataset (Table 2 in the research paper):
23
+ > **"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"**
24
+ > *University of Moratuwa, Department of Computer Science & Engineering*
25
+
26
+ ## 1. Architecture & Innovation
27
+ Instead of fine-tuning with a replacement flat softmax head (which ruins fastText's Huffman tree and causes catastrophic forgetting), our method surgically splits the Sinhala (`si`) leaf node in the binary Huffman tree into modern Sinhala and canonical Pali branches.
28
+
29
+ All original 176 language decision nodes are retained intact, and the model is trained strictly on target text (`train.csv`).
30
+
31
+ ## 2. Checkpoint Files
32
+ - `weights.pt`: PyTorch weights containing inherited input embedding and expanded output hierarchical decision parameters (124.5 MB).
33
+ - `config.json`: Tree paths, Huffman bit codes, and label mapping.
34
+ - `vocab.json`: Subword vocabulary and hashes.
35
+
36
+ ## 3. How to Load and Run Inference in Python
37
+ ```python
38
+ from data_pipeline.fasttext_continual.model import ContinualLID
39
+
40
+ # Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
41
+ model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-target-only")
42
+
43
+ # Inference example:
44
+ text = "මනොපුබ්බඞ්ගමා ධම්මා මනොසෙට්ඨා මනොමයා"
45
+ prediction, score = model.predict(text)
46
+ print(f"Language: {prediction}, Score: {score:.4f}")
47
+ # Output: Language: pi, Score: 0.99...
48
+ ```