This is a fun research project, not a useful model. I wanted to see how quickly a small model could be trained from scratch on one GPU at home, all the way from pretraining to a tool-using agent with reinforcement learning. The model works only on the narrow synthetic tasks it was trained on. It fails at anything else, and it can't read ordinary English text well.

What it cost: only my power bill. About 33 GPU-hours on one Intel Arc Pro B70 (275 W cap), or roughly 13–15 kWh for the whole PC. At ~€0.30/kWh that is about €4–5.

By @imdariotoo

tiny-agent-112m

A 112M-parameter decoder trained from scratch in about 30 hours of wall-clock time on a single Intel Arc Pro B70 (32 GB). The pipeline: pretrain on code, math and English; distill from a Qwen3.8-27B teacher; run GRPO in a sandbox with bash, read, edit, write and submit tools. The architecture borrows from DeepSeek-V4.1-Flash: a tiny shared key/value cache and Engram n-gram memory.

Everything is here: weights (base and RL), the full training code, and an honest account of what worked and what didn't.

Results at a glance

Synthetic workspace tasks it trained on (14 kinds) ~91% pass@1
The same tasks, reworded in ways it never saw 68–74%
Held-out task kinds (combining skills in new ways) ~0%
Reading real prose (HotpotQA, given the right paragraphs) ~2% exact match
Invented answers (wrong and seen in no tool output) ~2.5% of episodes

Try it

pip install torch tokenizers safetensors numpy huggingface_hub   # Python 3.11+
hf download darioooooo0o/tiny-agent-112m --local-dir tiny-agent-112m && cd tiny-agent-112m
python code/scripts/demo.py --model rl --tokenizer tokenizer.json            # random task
python code/scripts/demo.py --model rl --kind fix_test --seed 7
python code/scripts/demo.py --model rl --question "Which service has the most replicas?"

It generates a small fake project (configs, code, tests, docs, logs, CSVs) in a sandbox directory, asks the model a question about it, and prints every thought, tool call and result. It runs on CPU, Intel XPU or CUDA. The model's shell commands run inside bubblewrap (no network, read-only system, only the temporary workspace writable), so the demo needs Linux with bwrap installed. The custom architecture lives in code/tiny_agent/; there is no transformers integration.

Files

Path What
rl/ RL model (GRPO run "r2", step 300). Use this one.
base/ Base model after pretraining and LR decay, before RL
tokenizer.json 32k BPE tokenizer trained on the pretraining data
code/ Full code: model, data, training, teacher distillation, GRPO, evals, dashboard

Model

Layers / width / heads 18 / 768 / 12 (head dim 64, 2 KV heads)
Parameters 112M backbone + 50M embeddings + 269M Engram tables (lookup memory; adds no compute)
Attention Simplified DeepSeek-V4.1 "CSA": one shared global K/V per group of 4 layers, plus a 128-token sliding window per layer, in a single softmax (FlexAttention)
KV cache ~2.5 KB per token of context (standard multi-head attention at this depth: ~55 KB)
Engram Hashed 2- and 3-gram embedding memory at layer 1 (Cheng et al. 2026, as used in V4.1)
Other ReLU² MLP, RoPE, logit soft-cap, Muon optimizer with Sinkhorn-normalized momentum, WSD schedule
Context Trained at 2k, LR decay at 4k

How it was trained (timeline)

Stage GPU time What
Sizing and tuning ~3 h Throughput benchmarks for 4 sizes, LR sweep, 60-min Engram A/B
Pretraining 11.1 h 2.01B tokens: code 40%, math 25%, English 25%, grounded-reading QA 10%. ~50k tok/s
Teacher distillation 10.1 h Qwen3.8-27B (vLLM on the same GPU) solved 11.9k sandbox tasks; 9.3k usable episodes kept
LR decay 3.1 h 0.5B tokens, mix shifted toward tool use: teacher episodes, scripted episodes, reading data
RL (GRPO) 4.1 h Continuous-batching rollouts, ~30 s per step (256 episodes)
Warm start + evals ~1.5 h
Total ~33 h 2.51B pretraining tokens

RL details.

  • Group-relative advantages (8 attempts per task), with a token-level importance ratio between the sampling and training policies.
  • Reward shaping:
    • a MiMo-style length penalty;
    • extra weight on turns with malformed or repeated tool calls;
    • an invented-answer penalty: −0.3 for a wrong answer that appears in no tool output.
  • The final run also rewords training questions at random, and adds two task kinds that teach writing files and following pointers.

What I learned

  1. RL polishes; it doesn't add new skills here. RL lifted the trained tasks from 86% to ~91% and then saturated. Through supervised training and 300 RL steps, the held-out task kinds stayed near 0%.
  2. Fixed-template evals flatter small models. Rewording the questions cost 25–40 points.
    • Plain RL did nothing for that: 51% → 54% on fresh wording.
    • Training on reworded demos, then RL with reworded questions, got 51% → 74%.
  3. The model learned recipes, not reasoning. Transcripts on held-out tasks show it stitching together phrases from its scripted training demos out of context. Once it wrote an error message into the answer file.
  4. Reading is the wall for tiny search agents. BM25 over 5.2M Wikipedia abstracts finds both needed paragraphs within two perfect hops 64% of the time. But the model answers only ~2% of HotpotQA questions even when handed the right paragraphs, and ~28% of single-paragraph SQuAD questions. It saw only ~0.5B tokens of English prose; the sandbox tasks work because their answers sit in grep-able key: value lines.
  5. Teacher format matters. Asked to emit our JSON tool calls, Qwen3.8 blended them with its native XML format, and only ~20% of episodes were usable. Prompting in its native format and converting afterwards gave ~90%. (The thinking-off mode avoided Qwen's famously long thinking: ~50 tokens per call.)
  6. The honesty penalty worked without being gamed. Invented answers stayed at 1–3%. Answering NOT_FOUND on answerable questions stayed at ~1%, and copying wrong text from tool output fell from 6% to 2.4%.
  7. Engram didn't help at this scale, at least in a 60-minute equal-time A/B (val loss 4.45 with it, 4.39 without). It was kept anyway, for the fun of it.
  8. Intel GPU note. The xe driver kills compute jobs that run longer than 5 s, and the training process then hangs instead of exiting. A liveness watchdog (code/scripts/rl_watchdog.sh) handles it.

Evaluation details

Fixed eval: 150 synthetic workspace tasks, 8 samples each, temperature 1.0.

  • 12 trained kinds: config/code/CSV/doc/log lookups, running code, counting, editing, fixing a failing test, answering NOT_FOUND.
  • 3 held-out kinds: two-file lookup, counting log lines, writing a looked-up value to a file.
base after warm start RL r1 (step 125) RL r2 (step 300, rl/)
Trained kinds 86.1% 88.1% 91.4% 90.7%
Held-out kinds 0.0% 0.4% 0.4% 0.0%
Fresh wording (72 tasks) 51% 61% 54% 68% (74% at step 125)

Training data

Limitations

  • Works only on the synthetic sandbox tasks; general instructions and chat don't work.
  • Little world knowledge, by design.
  • Weak at reading natural-language text.
  • Outputs can be wrong or invented. Don't use it for anything that matters.
  • Bias and safety were not evaluated.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train darioooooo0o/tiny-agent-112m