This is a fun research project, not a useful model. I wanted to see how quickly a small model could be trained from scratch on one GPU at home, all the way from pretraining to a tool-using agent with reinforcement learning. The model works only on the narrow synthetic tasks it was trained on. It fails at anything else, and it can't read ordinary English text well.
What it cost: only my power bill. About 33 GPU-hours on one Intel Arc Pro B70 (275 W cap), or roughly 13–15 kWh for the whole PC. At ~€0.30/kWh that is about €4–5.
By @imdariotoo
tiny-agent-112m
A 112M-parameter decoder trained from scratch in about 30 hours of wall-clock time on a single
Intel Arc Pro B70 (32 GB). The pipeline: pretrain on code, math and English; distill from a
Qwen3.8-27B teacher; run GRPO in a sandbox with bash, read, edit, write and submit tools.
The architecture borrows from DeepSeek-V4.1-Flash: a tiny shared key/value cache and Engram n-gram
memory.
Everything is here: weights (base and RL), the full training code, and an honest account of what worked and what didn't.
Results at a glance
| Synthetic workspace tasks it trained on (14 kinds) | ~91% pass@1 |
| The same tasks, reworded in ways it never saw | 68–74% |
| Held-out task kinds (combining skills in new ways) | ~0% |
| Reading real prose (HotpotQA, given the right paragraphs) | ~2% exact match |
| Invented answers (wrong and seen in no tool output) | ~2.5% of episodes |
Try it
pip install torch tokenizers safetensors numpy huggingface_hub # Python 3.11+
hf download darioooooo0o/tiny-agent-112m --local-dir tiny-agent-112m && cd tiny-agent-112m
python code/scripts/demo.py --model rl --tokenizer tokenizer.json # random task
python code/scripts/demo.py --model rl --kind fix_test --seed 7
python code/scripts/demo.py --model rl --question "Which service has the most replicas?"
It generates a small fake project (configs, code, tests, docs, logs, CSVs) in a sandbox directory,
asks the model a question about it, and prints every thought, tool call and result. It runs on CPU,
Intel XPU or CUDA. The model's shell commands run inside bubblewrap
(no network, read-only system, only the temporary workspace writable), so the demo needs Linux
with bwrap installed. The custom architecture lives in code/tiny_agent/; there is no
transformers integration.
Files
| Path | What |
|---|---|
rl/ |
RL model (GRPO run "r2", step 300). Use this one. |
base/ |
Base model after pretraining and LR decay, before RL |
tokenizer.json |
32k BPE tokenizer trained on the pretraining data |
code/ |
Full code: model, data, training, teacher distillation, GRPO, evals, dashboard |
Model
| Layers / width / heads | 18 / 768 / 12 (head dim 64, 2 KV heads) |
| Parameters | 112M backbone + 50M embeddings + 269M Engram tables (lookup memory; adds no compute) |
| Attention | Simplified DeepSeek-V4.1 "CSA": one shared global K/V per group of 4 layers, plus a 128-token sliding window per layer, in a single softmax (FlexAttention) |
| KV cache | ~2.5 KB per token of context (standard multi-head attention at this depth: ~55 KB) |
| Engram | Hashed 2- and 3-gram embedding memory at layer 1 (Cheng et al. 2026, as used in V4.1) |
| Other | ReLU² MLP, RoPE, logit soft-cap, Muon optimizer with Sinkhorn-normalized momentum, WSD schedule |
| Context | Trained at 2k, LR decay at 4k |
How it was trained (timeline)
| Stage | GPU time | What |
|---|---|---|
| Sizing and tuning | ~3 h | Throughput benchmarks for 4 sizes, LR sweep, 60-min Engram A/B |
| Pretraining | 11.1 h | 2.01B tokens: code 40%, math 25%, English 25%, grounded-reading QA 10%. ~50k tok/s |
| Teacher distillation | 10.1 h | Qwen3.8-27B (vLLM on the same GPU) solved 11.9k sandbox tasks; 9.3k usable episodes kept |
| LR decay | 3.1 h | 0.5B tokens, mix shifted toward tool use: teacher episodes, scripted episodes, reading data |
| RL (GRPO) | 4.1 h | Continuous-batching rollouts, ~30 s per step (256 episodes) |
| Warm start + evals | ~1.5 h | |
| Total | ~33 h | 2.51B pretraining tokens |
RL details.
- Group-relative advantages (8 attempts per task), with a token-level importance ratio between the sampling and training policies.
- Reward shaping:
- a MiMo-style length penalty;
- extra weight on turns with malformed or repeated tool calls;
- an invented-answer penalty: −0.3 for a wrong answer that appears in no tool output.
- The final run also rewords training questions at random, and adds two task kinds that teach writing files and following pointers.
What I learned
- RL polishes; it doesn't add new skills here. RL lifted the trained tasks from 86% to ~91% and then saturated. Through supervised training and 300 RL steps, the held-out task kinds stayed near 0%.
- Fixed-template evals flatter small models. Rewording the questions cost 25–40 points.
- Plain RL did nothing for that: 51% → 54% on fresh wording.
- Training on reworded demos, then RL with reworded questions, got 51% → 74%.
- The model learned recipes, not reasoning. Transcripts on held-out tasks show it stitching together phrases from its scripted training demos out of context. Once it wrote an error message into the answer file.
- Reading is the wall for tiny search agents. BM25 over 5.2M Wikipedia abstracts finds both
needed paragraphs within two perfect hops 64% of the time. But the model answers only ~2% of
HotpotQA questions even when handed the right paragraphs, and ~28% of single-paragraph SQuAD
questions. It saw only ~0.5B tokens of English prose; the sandbox tasks work because their answers
sit in grep-able
key: valuelines. - Teacher format matters. Asked to emit our JSON tool calls, Qwen3.8 blended them with its native XML format, and only ~20% of episodes were usable. Prompting in its native format and converting afterwards gave ~90%. (The thinking-off mode avoided Qwen's famously long thinking: ~50 tokens per call.)
- The honesty penalty worked without being gamed. Invented answers stayed at 1–3%. Answering NOT_FOUND on answerable questions stayed at ~1%, and copying wrong text from tool output fell from 6% to 2.4%.
- Engram didn't help at this scale, at least in a 60-minute equal-time A/B (val loss 4.45 with it, 4.39 without). It was kept anyway, for the fun of it.
- Intel GPU note. The
xedriver kills compute jobs that run longer than 5 s, and the training process then hangs instead of exiting. A liveness watchdog (code/scripts/rl_watchdog.sh) handles it.
Evaluation details
Fixed eval: 150 synthetic workspace tasks, 8 samples each, temperature 1.0.
- 12 trained kinds: config/code/CSV/doc/log lookups, running code, counting, editing, fixing a failing test, answering NOT_FOUND.
- 3 held-out kinds: two-file lookup, counting log lines, writing a looked-up value to a file.
| base | after warm start | RL r1 (step 125) | RL r2 (step 300, rl/) |
|
|---|---|---|---|---|
| Trained kinds | 86.1% | 88.1% | 91.4% | 90.7% |
| Held-out kinds | 0.0% | 0.4% | 0.4% | 0.0% |
| Fresh wording (72 tasks) | 51% | 61% | 54% | 68% (74% at step 125) |
Training data
- FineWeb-Edu (ODC-By)
- FineMath (ODC-By)
- Stack-Edu via Common Pile (openly licensed code)
- OpenMathInstruct-2 (CC-BY-4.0)
- GSM8K (MIT)
- SQuAD v2 and HotpotQA (CC-BY-SA-4.0)
- Teacher trajectories from Qwen3.8-27B (Apache-2.0), released as darioooooo0o/tiny-agent-qwen3.8-trajectories
- Synthetic sandbox tasks generated by
code/tiny_agent/tasks.py
Limitations
- Works only on the synthetic sandbox tasks; general instructions and chat don't work.
- Little world knowledge, by design.
- Weak at reading natural-language text.
- Outputs can be wrong or invented. Don't use it for anything that matters.
- Bias and safety were not evaluated.