IndicSLM β a from-scratch India-specific small language model
A ~22M-parameter decoder-only transformer built fully from scratch β no HF Trainer, no
tokenizers library, no library RoPE/RMSNorm/SwiGLU/attention implementations β for Indian
government/legal-domain Q&A (RTI Act 2005, Consumer Protection Act 2019, myscheme.gov.in scheme
eligibility) in English, Hindi, and Marathi. This is a portfolio/applied-research project demonstrating
an end-to-end LLM pipeline (tokenizer β architecture β pretraining β SFT β DPO) at small scale on
free-tier compute (Kaggle T4s), not a production information source.
Read this before using any output. This model frequently answers with the wrong scheme's details, sometimes invents plausible-looking but fake
myscheme.gov.inURLs, and does not know when it doesn't know (a France-capital question gets a confident, fluent, irrelevant answer). See Evaluation below for measured rates. Do not use this for real RTI/scheme/legal guidance.
Stages in this repo
| Stage | Checkpoint | What it can do |
|---|---|---|
| 1. Pretrained | checkpoints/stage1_pretrained.pt |
Free-text continuation only. No instruction-following. |
| 2. SFT (English) | checkpoints/stage2_sft_english.pt |
Instruction/response format, English only. |
| 3. SFT (multilingual) | checkpoints/stage3_sft_multilingual.pt |
Same, extended to Hindi + Marathi via machine-translated response data. |
| 4. DPO | checkpoints/stage4_dpo.pt |
Direct Preference Optimization on top of Stage 3, 1,171 preference pairs. Measured, modest improvement β see Evaluation below. |
Architecture
Decoder-only transformer: RoPE (rotate-half), RMSNorm, SwiGLU, Grouped-Query Attention, KV-cache β
all raw PyTorch tensor ops. vocab_size=16,000 (+2 SFT special tokens), d_model=384, n_layers=10,
n_heads=6, n_kv_heads=2 (GQA 3:1), head_dim=64, ffn_hidden_dim=1024, max_seq_len=512
(pretrained) / 1024 (SFT), tied embeddings, no Linear biases. 21,880,704 parameters.
Honest framing on GQA: implemented for architectural fidelity to production LLM design and to demonstrate the technique β not because this model has a real KV-cache memory problem. At this scale, full multi-head attention's KV cache would already be a few MB.
Tokens-per-parameter β 4.6:1 against Chinchilla's ~20:1 rule of thumb β deliberately undertrained, a scale/coherence tradeoff, not an oversight.
Tokenizer
Byte-level BPE implemented from scratch (Sennrich et al. algorithm). 16,000 merges trained on the project's ~915MB corpus. On held-out Hindi/Marathi text: 3.58 chars/token vs GPT-2's 0.83 β GPT-2's vocab has essentially no Devanagari merges and fragments toward byte-level; this tokenizer doesn't.
Training data
- Pretraining (~115M tokens): Hindi + Marathi Wikipedia (full dumps, ~64M + ~18M tokens) + a
domain slice (RTI Act, Consumer Protection Act, ~3,400 myscheme.gov.in scheme entries from the
jainamgada45/indian-government-schemesKaggle dataset), repeated ~7.5x to reach ~20% domain fraction. There is no general-domain English pretraining text β English exposure is entirely the repeated domain slice, so English generation skews toward legal/bureaucratic register. - SFT (~20,400 rows in English; extended to Hindi/Marathi via IndicTrans2-200M machine translation, ~17,000-17,400 rows per language after dropping MT rows with corrupted numbers/URLs or lost the English scheme name β ~85% keep rate, biased toward simpler responses). ~15% of every SFT batch is unmasked pretraining-replay to limit catastrophic forgetting.
Evaluation
Measured with this project's own eval harness (paired prompts across en/hi/mr, same scheme +
category in every language, greedy + sampled decoding, n=30/language/split). Full methodology and
per-axis numbers are in this project's eval/ scripts β headline results, Stage 3
checkpoint, greedy decoding, unseen ("val") schemes:
| Axis | hi | mr | en |
|---|---|---|---|
Ends at <eos> (300-token cap) |
73% | 70% | 80% |
| Correct-script response | 100% | 100% | 100% |
| English scheme name copied | 97% | 93% | 93% |
| Repetition loop (β₯30% repeated 4-grams) | 50% | 43% | 10% |
| Numbers matching the reference answer | 14% | 33% | 32% |
Separately, of 110 generated URLs pooled across evaluation runs (Stage 3): 0 matched the correct
reference URL for their prompt; 67% used a myscheme.gov.in/schemes/<slug> shape with a slug that does
not exist in the real 3,400-scheme catalog. The model has learned the shape of an
authoritative-looking reference, not real ones β treat any URL it produces as fabricated until verified.
(This check has not yet been re-run for Stage 4 β see the DPO note below.)
Format-following (script match, scheme-name copy, template structure) is reliable. Factual grounding (the right scheme's details) is not β greedy word-overlap with the reference answer on schemes the model did train on (0.41) was not meaningfully higher than on unseen schemes (0.48), i.e. no clear memorization advantage was observed at this scale (n=90, not a large sample).
Stage 4 (DPO) vs Stage 3, same unseen-scheme prompts, both decode modes β trained on 1,171 preference pairs sampled from Stage 3 (rejected = a repetition loop, a missing scheme name, or low overlap with the gold answer):
| Axis | DPO greedy | Stage 3 greedy | DPO sampled | Stage 3 sampled |
|---|---|---|---|---|
| Word-overlap F1 vs reference | 0.51 | 0.48 | 0.45 | 0.46 |
| Repetition-loop rate | 37.8% | 34.4% | 4.4% | 10.0% |
A small, real improvement on the axis DPO was trained against (reference overlap), and fewer loops under sampled decoding β but greedy-decoding loops did not improve (in fact ticked up slightly), and the fabricated-URL rate has not been re-measured on this checkpoint. Read this as "DPO helped a little, on a small pair set, not as a fix for grounding or hallucinated URLs" β not as a solved problem.
Usage
import sys
sys.path.insert(0, ".") # this repo's root, after cloning/downloading
from stages import run_stage
output = run_stage("sft_multi", "PM-KISAN ΰ€―ΰ₯ΰ€ΰ€¨ΰ₯ΰ€Έΰ€Ύΰ€ ΰ₯ ΰ€ΰ₯ΰ€£ ΰ€ͺΰ€Ύΰ€€ΰ₯ΰ€° ΰ€ΰ€Ήΰ₯?", max_new_tokens=200)
print(output)
stages.py handles tokenizer loading, the base-vs-instruction prompt-format difference, and KV-cache
generation. Requires torch and regex (pip install torch regex).
Known limitations (do not treat as solved)
- Wrong-scheme retrieval: asked about one scheme, frequently answers with another's facts.
- Fabricated URLs: see Evaluation above β this is a specific, checkable-looking false claim, a more serious risk class than generic fluent-but-wrong hallucination.
- No abstention: no SFT row teaches "say you don't know" β out-of-domain questions get confident, fluent, irrelevant answers.
- Marathi is the weakest language, especially the
rti_procedurecategory (only 2 seed training rows vs ~6,800 per scheme category) β degenerates into repetition loops more than Hindi or English. - SFT/MT data has not had a native-speaker accuracy pass. The Hindi/Marathi response text was produced by IndicTrans2-200M machine translation plus automated integrity filtering (numbers, URLs, scheme-name preservation) β not verified by a native speaker. This affects training data quality, not just generation quality.
If you plan to redistribute or build on this model, check these sources' actual terms yourself before doing so β do not assume permissive licensing.
- Downloads last month
- 246