custom_tokenizer / README.md
luffy19's picture
Upload README.md with huggingface_hub
ffd7a98 verified
|
Raw History Blame Contribute Delete
5.35 kB
# Agri-SLM β€” inference
A 134M-parameter Qwen3-style decoder trained from scratch on Indian agriculture
text (11.8B tokens, 3 epochs). Two variants differing only in tokenizer.
`agri_slm.py` is self-contained: it needs `torch` and `tokenizers` only, and has
no dependency on the training code. It is verified bit-for-bit identical to the
training implementation.
## Install
```bash
pip install torch tokenizers
pip install huggingface_hub # only if loading from the Hub
```
## Run
```bash
# download and generate
python agri_slm.py --prompt "Rice blast disease is caused by"
# interactive, streaming
python agri_slm.py --interactive
# local files
python agri_slm.py --ckpt final.pt --tokenizer tokenizer_final.json --interactive
```
As a library:
```python
from agri_slm import AgriSLM
slm = AgriSLM.from_pretrained("luffy19/custom_tokenizer")
print(slm.generate("Integrated pest management in cotton involves"))
# or from local files
slm = AgriSLM.from_files("final.pt", "tokenizer_final.json")
```
## Decoding β€” read this before reporting results
**This model loops badly under greedy decoding.** That is a property of greedy
search on small base models, not a defect in the weights: `argmax` has no way out
of a high-probability cycle. A repetition penalty fixes it completely.
Measured over 300-token generations on 4 prompts:
| setting | repeat-4gram | longest verbatim loop | distinct-2 | sentences well-formed |
|---|---|---|---|---|
| greedy, no penalty | 0.636 | **37 tokens** | 0.285 | 91% |
| greedy + rep 1.2 | 0.023 | 0 | 0.884 | 84% |
| greedy + rep 1.2 + no-repeat-4gram | 0.000 | 0 | 0.939 | 78% |
| T 0.7 | 0.173 | 0 | 0.633 | 89% |
| T 0.7 + rep 1.15 | 0.011 | 0.8 | 0.919 | 80% |
| T 0.9 + rep 1.15 + ng4 | 0.000 | 0 | 0.902 | 85% |
| **T 1.0 + rep 1.1** | **0.006** | **0** | **0.910** | **87%** |
A repetition penalty alone takes greedy from a 37-token loop to essentially zero
repetition β€” a 27x reduction.
### Presets
```bash
python agri_slm.py --preset balanced # T=1.0 top_p=0.95 rep=1.10 (default)
python agri_slm.py --preset safe # T=0.9 top_p=0.95 rep=1.15 ngram=4 (zero loops)
python agri_slm.py --preset focused # T=0.7 top_p=0.90 rep=1.15 (deterministic-ish)
python agri_slm.py --preset greedy # T=0 rep=1.20 ngram=3 (reproducible)
python agri_slm.py --preset raw-greedy # no penalty β€” demonstrates the loop
```
Any preset value can be overridden: `--temperature`, `--top-k`, `--top-p`,
`--repetition-penalty`, `--no-repeat-ngram`.
**`balanced` is the recommended default.** Use `safe` when you need a hard
guarantee of no repetition; it costs a little naturalness. `raw-greedy` exists
only to reproduce the failure.
## Models
| repo | tokenizer | notes |
|---|---|---|
| `luffy19/custom_tokenizer` | hybrid 40k (KG + AGROVOC entity injection) | `hybrid40k/final.pt` |
| `luffy19/agri-slm_qwen40k` | Qwen BPE pruned 151k→40k | `qwen40k/final.pt` |
Both: 133,476,480 params, 40,000 vocab, 20 layers, 640-dim, GQA (10 heads /
5 KV groups), SwiGLU, RMSNorm, RoPE base 1e6, rank-256 factorized embedding tied
to the output head.
### Measured differences
On 2,000 identical held-out documents (3.72M characters):
| metric | Qwen 40k | Hybrid 40k | delta |
|---|---|---|---|
| bits per character | 0.9208 | **0.9168** | **βˆ’0.43%** |
| fertility (tok/1k chars) | 229.16 | **212.76** | **βˆ’7.16%** |
| tokens for same text | 853,383 | **792,312** | **βˆ’7.16%** |
| agriculture probe | 14/15 | **15/15** | +1 |
| perplexity | 16.20 | 19.82 | *not comparable* |
**Do not compare the two models by perplexity.** It is per-token, and the hybrid's
tokens each carry ~7% more text, so it is predicting harder units. Bits-per-char
normalises this and reverses the apparent ordering. The 7.16% fertility gain is
the large, robust effect; the 0.43% BPC difference is small and comes from single
runs with one seed, so treat it as directional rather than significant.
## What this model is and is not
It is a **base language model**: it continues text. It is not instruction-tuned
and does not answer questions. Prompt it with sentence openings a textbook would
contain ("Rice blast disease is caused by"), not questions ("What causes rice
blast?").
Known limitations, all measured:
- **Confidently wrong on some facts.** It defined Minimum Support Price as "a
measure of the quality of a product", which is incorrect. Domain-term recall is
strong (14–15 of 15 probes) but definitions can be wrong.
- **Invents citations.** ~11% of the training corpus is citation-dense academic
text, so it produces plausible-looking references like "(Khan et al., 2015)"
that may not exist.
- **Drifts into thesis format.** ~13% of the corpus is dissertation text, so long
generations can wander into "REVIEW OF LITERATURE / Chapter II".
- **Reproduces a PDF extraction artifact.** ~1% of training documents contain
doubled letters from bad PDF extraction ("RREEVVIIEEWW"); the model occasionally
emits this.
- **No KV cache.** Generation recomputes the full sequence each step, so cost
grows with length. Fine to a few hundred tokens; slow beyond that.
Do not use its output as agricultural advice without verification.
## Performance
~25 tokens/sec on an H200 at ~1 GB in bf16. Runs on CPU (fp32) at a few tokens
per second.