YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Agri-SLM β inference
A 134M-parameter Qwen3-style decoder trained from scratch on Indian agriculture text (11.8B tokens, 3 epochs). Two variants differing only in tokenizer.
agri_slm.py is self-contained: it needs torch and tokenizers only, and has
no dependency on the training code. It is verified bit-for-bit identical to the
training implementation.
Install
pip install torch tokenizers
pip install huggingface_hub # only if loading from the Hub
Run
# download and generate
python agri_slm.py --prompt "Rice blast disease is caused by"
# interactive, streaming
python agri_slm.py --interactive
# local files
python agri_slm.py --ckpt final.pt --tokenizer tokenizer_final.json --interactive
As a library:
from agri_slm import AgriSLM
slm = AgriSLM.from_pretrained("luffy19/custom_tokenizer")
print(slm.generate("Integrated pest management in cotton involves"))
# or from local files
slm = AgriSLM.from_files("final.pt", "tokenizer_final.json")
Decoding β read this before reporting results
This model loops badly under greedy decoding. That is a property of greedy
search on small base models, not a defect in the weights: argmax has no way out
of a high-probability cycle. A repetition penalty fixes it completely.
Measured over 300-token generations on 4 prompts:
| setting | repeat-4gram | longest verbatim loop | distinct-2 | sentences well-formed |
|---|---|---|---|---|
| greedy, no penalty | 0.636 | 37 tokens | 0.285 | 91% |
| greedy + rep 1.2 | 0.023 | 0 | 0.884 | 84% |
| greedy + rep 1.2 + no-repeat-4gram | 0.000 | 0 | 0.939 | 78% |
| T 0.7 | 0.173 | 0 | 0.633 | 89% |
| T 0.7 + rep 1.15 | 0.011 | 0.8 | 0.919 | 80% |
| T 0.9 + rep 1.15 + ng4 | 0.000 | 0 | 0.902 | 85% |
| T 1.0 + rep 1.1 | 0.006 | 0 | 0.910 | 87% |
A repetition penalty alone takes greedy from a 37-token loop to essentially zero repetition β a 27x reduction.
Presets
python agri_slm.py --preset balanced # T=1.0 top_p=0.95 rep=1.10 (default)
python agri_slm.py --preset safe # T=0.9 top_p=0.95 rep=1.15 ngram=4 (zero loops)
python agri_slm.py --preset focused # T=0.7 top_p=0.90 rep=1.15 (deterministic-ish)
python agri_slm.py --preset greedy # T=0 rep=1.20 ngram=3 (reproducible)
python agri_slm.py --preset raw-greedy # no penalty β demonstrates the loop
Any preset value can be overridden: --temperature, --top-k, --top-p,
--repetition-penalty, --no-repeat-ngram.
balanced is the recommended default. Use safe when you need a hard
guarantee of no repetition; it costs a little naturalness. raw-greedy exists
only to reproduce the failure.
Models
| repo | tokenizer | notes |
|---|---|---|
luffy19/custom_tokenizer |
hybrid 40k (KG + AGROVOC entity injection) | hybrid40k/final.pt |
luffy19/agri-slm_qwen40k |
Qwen BPE pruned 151kβ40k | qwen40k/final.pt |
Both: 133,476,480 params, 40,000 vocab, 20 layers, 640-dim, GQA (10 heads / 5 KV groups), SwiGLU, RMSNorm, RoPE base 1e6, rank-256 factorized embedding tied to the output head.
Measured differences
On 2,000 identical held-out documents (3.72M characters):
| metric | Qwen 40k | Hybrid 40k | delta |
|---|---|---|---|
| bits per character | 0.9208 | 0.9168 | β0.43% |
| fertility (tok/1k chars) | 229.16 | 212.76 | β7.16% |
| tokens for same text | 853,383 | 792,312 | β7.16% |
| agriculture probe | 14/15 | 15/15 | +1 |
| perplexity | 16.20 | 19.82 | not comparable |
Do not compare the two models by perplexity. It is per-token, and the hybrid's tokens each carry ~7% more text, so it is predicting harder units. Bits-per-char normalises this and reverses the apparent ordering. The 7.16% fertility gain is the large, robust effect; the 0.43% BPC difference is small and comes from single runs with one seed, so treat it as directional rather than significant.
What this model is and is not
It is a base language model: it continues text. It is not instruction-tuned and does not answer questions. Prompt it with sentence openings a textbook would contain ("Rice blast disease is caused by"), not questions ("What causes rice blast?").
Known limitations, all measured:
- Confidently wrong on some facts. It defined Minimum Support Price as "a measure of the quality of a product", which is incorrect. Domain-term recall is strong (14β15 of 15 probes) but definitions can be wrong.
- Invents citations. ~11% of the training corpus is citation-dense academic text, so it produces plausible-looking references like "(Khan et al., 2015)" that may not exist.
- Drifts into thesis format. ~13% of the corpus is dissertation text, so long generations can wander into "REVIEW OF LITERATURE / Chapter II".
- Reproduces a PDF extraction artifact. ~1% of training documents contain doubled letters from bad PDF extraction ("RREEVVIIEEWW"); the model occasionally emits this.
- No KV cache. Generation recomputes the full sequence each step, so cost grows with length. Fine to a few hundred tokens; slow beyond that.
Do not use its output as agricultural advice without verification.
Performance
~25 tokens/sec on an H200 at ~1 GB in bf16. Runs on CPU (fp32) at a few tokens per second.