IndicSLM β€” a from-scratch India-specific small language model

A ~22M-parameter decoder-only transformer built fully from scratch β€” no HF Trainer, no tokenizers library, no library RoPE/RMSNorm/SwiGLU/attention implementations β€” for Indian government/legal-domain Q&A (RTI Act 2005, Consumer Protection Act 2019, myscheme.gov.in scheme eligibility) in English, Hindi, and Marathi. This is a portfolio/applied-research project demonstrating an end-to-end LLM pipeline (tokenizer β†’ architecture β†’ pretraining β†’ SFT β†’ DPO) at small scale on free-tier compute (Kaggle T4s), not a production information source.

Read this before using any output. This model frequently answers with the wrong scheme's details, sometimes invents plausible-looking but fake myscheme.gov.in URLs, and does not know when it doesn't know (a France-capital question gets a confident, fluent, irrelevant answer). See Evaluation below for measured rates. Do not use this for real RTI/scheme/legal guidance.

Stages in this repo

Stage Checkpoint What it can do
1. Pretrained checkpoints/stage1_pretrained.pt Free-text continuation only. No instruction-following.
2. SFT (English) checkpoints/stage2_sft_english.pt Instruction/response format, English only.
3. SFT (multilingual) checkpoints/stage3_sft_multilingual.pt Same, extended to Hindi + Marathi via machine-translated response data.
4. DPO checkpoints/stage4_dpo.pt Direct Preference Optimization on top of Stage 3, 1,171 preference pairs. Measured, modest improvement β€” see Evaluation below.

Architecture

Decoder-only transformer: RoPE (rotate-half), RMSNorm, SwiGLU, Grouped-Query Attention, KV-cache β€” all raw PyTorch tensor ops. vocab_size=16,000 (+2 SFT special tokens), d_model=384, n_layers=10, n_heads=6, n_kv_heads=2 (GQA 3:1), head_dim=64, ffn_hidden_dim=1024, max_seq_len=512 (pretrained) / 1024 (SFT), tied embeddings, no Linear biases. 21,880,704 parameters.

Honest framing on GQA: implemented for architectural fidelity to production LLM design and to demonstrate the technique β€” not because this model has a real KV-cache memory problem. At this scale, full multi-head attention's KV cache would already be a few MB.

Tokens-per-parameter β‰ˆ 4.6:1 against Chinchilla's ~20:1 rule of thumb β€” deliberately undertrained, a scale/coherence tradeoff, not an oversight.

Tokenizer

Byte-level BPE implemented from scratch (Sennrich et al. algorithm). 16,000 merges trained on the project's ~915MB corpus. On held-out Hindi/Marathi text: 3.58 chars/token vs GPT-2's 0.83 β€” GPT-2's vocab has essentially no Devanagari merges and fragments toward byte-level; this tokenizer doesn't.

Training data

  • Pretraining (~115M tokens): Hindi + Marathi Wikipedia (full dumps, ~64M + ~18M tokens) + a domain slice (RTI Act, Consumer Protection Act, ~3,400 myscheme.gov.in scheme entries from the jainamgada45/indian-government-schemes Kaggle dataset), repeated ~7.5x to reach ~20% domain fraction. There is no general-domain English pretraining text β€” English exposure is entirely the repeated domain slice, so English generation skews toward legal/bureaucratic register.
  • SFT (~20,400 rows in English; extended to Hindi/Marathi via IndicTrans2-200M machine translation, ~17,000-17,400 rows per language after dropping MT rows with corrupted numbers/URLs or lost the English scheme name β€” ~85% keep rate, biased toward simpler responses). ~15% of every SFT batch is unmasked pretraining-replay to limit catastrophic forgetting.

Evaluation

Measured with this project's own eval harness (paired prompts across en/hi/mr, same scheme + category in every language, greedy + sampled decoding, n=30/language/split). Full methodology and per-axis numbers are in this project's eval/ scripts β€” headline results, Stage 3 checkpoint, greedy decoding, unseen ("val") schemes:

Axis hi mr en
Ends at <eos> (300-token cap) 73% 70% 80%
Correct-script response 100% 100% 100%
English scheme name copied 97% 93% 93%
Repetition loop (β‰₯30% repeated 4-grams) 50% 43% 10%
Numbers matching the reference answer 14% 33% 32%

Separately, of 110 generated URLs pooled across evaluation runs (Stage 3): 0 matched the correct reference URL for their prompt; 67% used a myscheme.gov.in/schemes/<slug> shape with a slug that does not exist in the real 3,400-scheme catalog. The model has learned the shape of an authoritative-looking reference, not real ones β€” treat any URL it produces as fabricated until verified. (This check has not yet been re-run for Stage 4 β€” see the DPO note below.)

Format-following (script match, scheme-name copy, template structure) is reliable. Factual grounding (the right scheme's details) is not β€” greedy word-overlap with the reference answer on schemes the model did train on (0.41) was not meaningfully higher than on unseen schemes (0.48), i.e. no clear memorization advantage was observed at this scale (n=90, not a large sample).

Stage 4 (DPO) vs Stage 3, same unseen-scheme prompts, both decode modes β€” trained on 1,171 preference pairs sampled from Stage 3 (rejected = a repetition loop, a missing scheme name, or low overlap with the gold answer):

Axis DPO greedy Stage 3 greedy DPO sampled Stage 3 sampled
Word-overlap F1 vs reference 0.51 0.48 0.45 0.46
Repetition-loop rate 37.8% 34.4% 4.4% 10.0%

A small, real improvement on the axis DPO was trained against (reference overlap), and fewer loops under sampled decoding β€” but greedy-decoding loops did not improve (in fact ticked up slightly), and the fabricated-URL rate has not been re-measured on this checkpoint. Read this as "DPO helped a little, on a small pair set, not as a fix for grounding or hallucinated URLs" β€” not as a solved problem.

Usage

import sys
sys.path.insert(0, ".")  # this repo's root, after cloning/downloading
from stages import run_stage

output = run_stage("sft_multi", "PM-KISAN ΰ€―ΰ₯‹ΰ€œΰ€¨ΰ₯‡ΰ€Έΰ€Ύΰ€ ΰ₯€ ΰ€•ΰ₯‹ΰ€£ ΰ€ͺΰ€Ύΰ€€ΰ₯ΰ€° ΰ€†ΰ€Ήΰ₯‡?", max_new_tokens=200)
print(output)

stages.py handles tokenizer loading, the base-vs-instruction prompt-format difference, and KV-cache generation. Requires torch and regex (pip install torch regex).

Known limitations (do not treat as solved)

  • Wrong-scheme retrieval: asked about one scheme, frequently answers with another's facts.
  • Fabricated URLs: see Evaluation above β€” this is a specific, checkable-looking false claim, a more serious risk class than generic fluent-but-wrong hallucination.
  • No abstention: no SFT row teaches "say you don't know" β€” out-of-domain questions get confident, fluent, irrelevant answers.
  • Marathi is the weakest language, especially the rti_procedure category (only 2 seed training rows vs ~6,800 per scheme category) β€” degenerates into repetition loops more than Hindi or English.
  • SFT/MT data has not had a native-speaker accuracy pass. The Hindi/Marathi response text was produced by IndicTrans2-200M machine translation plus automated integrity filtering (numbers, URLs, scheme-name preservation) β€” not verified by a native speaker. This affects training data quality, not just generation quality.

If you plan to redistribute or build on this model, check these sources' actual terms yourself before doing so β€” do not assume permissive licensing.

Downloads last month
246
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using Atharva31/indicslm-from-scratch 1