|
Download README.md from luffy19/custom_tokenizer: direct link, hf CLI and curl.
- Browser
- Download file 5.35 kB
-
https://huggingface.co/luffy19/custom_tokenizer/resolve/main/README.md
- Command line
-
hf download hf://luffy19/custom_tokenizer/README.md
-
curl -L -o README.md https://huggingface.co/luffy19/custom_tokenizer/resolve/main/README.md
5.35 kB
| # Agri-SLM β inference | |
| A 134M-parameter Qwen3-style decoder trained from scratch on Indian agriculture | |
| text (11.8B tokens, 3 epochs). Two variants differing only in tokenizer. | |
| `agri_slm.py` is self-contained: it needs `torch` and `tokenizers` only, and has | |
| no dependency on the training code. It is verified bit-for-bit identical to the | |
| training implementation. | |
| ## Install | |
| ```bash | |
| pip install torch tokenizers | |
| pip install huggingface_hub # only if loading from the Hub | |
| ``` | |
| ## Run | |
| ```bash | |
| # download and generate | |
| python agri_slm.py --prompt "Rice blast disease is caused by" | |
| # interactive, streaming | |
| python agri_slm.py --interactive | |
| # local files | |
| python agri_slm.py --ckpt final.pt --tokenizer tokenizer_final.json --interactive | |
| ``` | |
| As a library: | |
| ```python | |
| from agri_slm import AgriSLM | |
| slm = AgriSLM.from_pretrained("luffy19/custom_tokenizer") | |
| print(slm.generate("Integrated pest management in cotton involves")) | |
| # or from local files | |
| slm = AgriSLM.from_files("final.pt", "tokenizer_final.json") | |
| ``` | |
| ## Decoding β read this before reporting results | |
| **This model loops badly under greedy decoding.** That is a property of greedy | |
| search on small base models, not a defect in the weights: `argmax` has no way out | |
| of a high-probability cycle. A repetition penalty fixes it completely. | |
| Measured over 300-token generations on 4 prompts: | |
| | setting | repeat-4gram | longest verbatim loop | distinct-2 | sentences well-formed | | |
| |---|---|---|---|---| | |
| | greedy, no penalty | 0.636 | **37 tokens** | 0.285 | 91% | | |
| | greedy + rep 1.2 | 0.023 | 0 | 0.884 | 84% | | |
| | greedy + rep 1.2 + no-repeat-4gram | 0.000 | 0 | 0.939 | 78% | | |
| | T 0.7 | 0.173 | 0 | 0.633 | 89% | | |
| | T 0.7 + rep 1.15 | 0.011 | 0.8 | 0.919 | 80% | | |
| | T 0.9 + rep 1.15 + ng4 | 0.000 | 0 | 0.902 | 85% | | |
| | **T 1.0 + rep 1.1** | **0.006** | **0** | **0.910** | **87%** | | |
| A repetition penalty alone takes greedy from a 37-token loop to essentially zero | |
| repetition β a 27x reduction. | |
| ### Presets | |
| ```bash | |
| python agri_slm.py --preset balanced # T=1.0 top_p=0.95 rep=1.10 (default) | |
| python agri_slm.py --preset safe # T=0.9 top_p=0.95 rep=1.15 ngram=4 (zero loops) | |
| python agri_slm.py --preset focused # T=0.7 top_p=0.90 rep=1.15 (deterministic-ish) | |
| python agri_slm.py --preset greedy # T=0 rep=1.20 ngram=3 (reproducible) | |
| python agri_slm.py --preset raw-greedy # no penalty β demonstrates the loop | |
| ``` | |
| Any preset value can be overridden: `--temperature`, `--top-k`, `--top-p`, | |
| `--repetition-penalty`, `--no-repeat-ngram`. | |
| **`balanced` is the recommended default.** Use `safe` when you need a hard | |
| guarantee of no repetition; it costs a little naturalness. `raw-greedy` exists | |
| only to reproduce the failure. | |
| ## Models | |
| | repo | tokenizer | notes | | |
| |---|---|---| | |
| | `luffy19/custom_tokenizer` | hybrid 40k (KG + AGROVOC entity injection) | `hybrid40k/final.pt` | | |
| | `luffy19/agri-slm_qwen40k` | Qwen BPE pruned 151kβ40k | `qwen40k/final.pt` | | |
| Both: 133,476,480 params, 40,000 vocab, 20 layers, 640-dim, GQA (10 heads / | |
| 5 KV groups), SwiGLU, RMSNorm, RoPE base 1e6, rank-256 factorized embedding tied | |
| to the output head. | |
| ### Measured differences | |
| On 2,000 identical held-out documents (3.72M characters): | |
| | metric | Qwen 40k | Hybrid 40k | delta | | |
| |---|---|---|---| | |
| | bits per character | 0.9208 | **0.9168** | **β0.43%** | | |
| | fertility (tok/1k chars) | 229.16 | **212.76** | **β7.16%** | | |
| | tokens for same text | 853,383 | **792,312** | **β7.16%** | | |
| | agriculture probe | 14/15 | **15/15** | +1 | | |
| | perplexity | 16.20 | 19.82 | *not comparable* | | |
| **Do not compare the two models by perplexity.** It is per-token, and the hybrid's | |
| tokens each carry ~7% more text, so it is predicting harder units. Bits-per-char | |
| normalises this and reverses the apparent ordering. The 7.16% fertility gain is | |
| the large, robust effect; the 0.43% BPC difference is small and comes from single | |
| runs with one seed, so treat it as directional rather than significant. | |
| ## What this model is and is not | |
| It is a **base language model**: it continues text. It is not instruction-tuned | |
| and does not answer questions. Prompt it with sentence openings a textbook would | |
| contain ("Rice blast disease is caused by"), not questions ("What causes rice | |
| blast?"). | |
| Known limitations, all measured: | |
| - **Confidently wrong on some facts.** It defined Minimum Support Price as "a | |
| measure of the quality of a product", which is incorrect. Domain-term recall is | |
| strong (14β15 of 15 probes) but definitions can be wrong. | |
| - **Invents citations.** ~11% of the training corpus is citation-dense academic | |
| text, so it produces plausible-looking references like "(Khan et al., 2015)" | |
| that may not exist. | |
| - **Drifts into thesis format.** ~13% of the corpus is dissertation text, so long | |
| generations can wander into "REVIEW OF LITERATURE / Chapter II". | |
| - **Reproduces a PDF extraction artifact.** ~1% of training documents contain | |
| doubled letters from bad PDF extraction ("RREEVVIIEEWW"); the model occasionally | |
| emits this. | |
| - **No KV cache.** Generation recomputes the full sequence each step, so cost | |
| grows with length. Fine to a few hundred tokens; slow beyond that. | |
| Do not use its output as agricultural advice without verification. | |
| ## Performance | |
| ~25 tokens/sec on an H200 at ~1 GB in bf16. Runs on CPU (fp32) at a few tokens | |
| per second. | |