Download code/README.md from darioooooo0o/tiny-agent-112m: direct link, hf CLI and curl.
- Browser
- Download file 5.71 kB
-
https://huggingface.co/darioooooo0o/tiny-agent-112m/resolve/main/code/README.md
- Command line
-
hf download hf://darioooooo0o/tiny-agent-112m/code/README.md
-
curl -L -o README.md https://huggingface.co/darioooooo0o/tiny-agent-112m/resolve/main/code/README.md
tiny-agent
A small from-scratch LLM for research on the Intel Arc Pro B70: about 24 hours of pretraining, then RL. It is built to be weak at recalling facts but good at tool calls and copying found information. Everything runs locally.
Design
| Part | Choice |
|---|---|
| Attention | DeepSeek V4.1 CSA2, simplified. Compression ratio 1, no indexer. In each group of 4 layers, 1 Full layer writes a shared global K/V and 3 Reuse layers reuse it. Every layer also has its own local K/V with a 128-token sliding window. Each query does one softmax over the union of global and local keys (FlexAttention). |
| KV cache | One global K/V per group plus one 128-slot ring buffer per layer. About 2.7 KB/token at 8k context for size L, roughly 20× smaller than plain attention. |
| Engram | Hashed 2- and 3-gram tables over compressed token ids, with a context-aware sigmoid gate, at layer 1. No short conv, as in V4.1. Behind --engram. |
| Backbone | Dense, ReLU² MLP, QK-norm, RoPE, logit softcap 15, parameter-free RMSNorm, zero-init output projections. |
| Optimizer | Muon for matrices (head-wise on Q). V4.1 momentum + Sinkhorn update for embedding, head and Engram tables. |
| Schedule | Warmup-stable-decay. touch DECAY starts the decay at any point; --max_minutes decays automatically to finish on time. |
| Tokenizer | 32k byte-level BPE trained on the mix, digits split one at a time, Hermes-style <tool_call>/<tool_response> tokens. |
| Data (stable) | Code 40% (openly licensed Stack-Edu: Python, Shell, JS/TS, Markdown, Go, Rust, C), math 25% (FineMath), English 25% (FineWeb-Edu), grounded reading 10%. |
| Data (decay) | More grounded reading, scripted tool-use trajectories, and math with <think> reasoning. |
| Tools | Pi's default four: bash, read, edit, write, plus submit. bash runs in bubblewrap: no network, read-only /usr, ulimits, timeout, tmpfs workspace. |
| Tasks | Workspaces with invented facts: configs in YAML, JSON, TOML or env, plus code, tests, docs, logs, CSVs and a long job log. 15 kinds: lookups, compute (run code, grep -c, count across files, sum a column), multi-step, edit, fix a failing test, write a file, and not-found. The answer is only in the files. |
| RL | GRPO directly on base snapshots, with no SFT stage. Continuous-batching rollouts: each cache slot is refilled the moment its episode ends, and episodes still running carry across updates (at most 2 updates stale) with a per-token importance ratio against the recorded sampling log-prob. Groups without signal are dropped. Rewards and credit follow MiMo-V2.6's released reference profile: 1 if correct, 0.5 if correct but never seen in a tool result; an in-group length penalty on passes when more than half the group passes; a per-turn penalty for malformed, unknown, bad-argument or repeated tool calls; loss averaged per prompt; no KL. Curriculum favors task kinds solved about half the time. |
Run
source env.sh
$TA_PY scripts/download.py # raw data -> $TA_DATA/raw
$TA_PY scripts/train_tokenizer.py # tokenizer.json + cid_map.npy (Engram compression)
$TA_PY scripts/tokenize_data.py # -> $TA_DATA/tok/{train,val}/<source>/*.bin
$TA_PY scripts/gen_synth.py # synth_tools / synth_grounded / synth_reasoning
$TA_PY scripts/bench_size.py # tok/s per size -> pick the size
$TA_PY scripts/train.py --size L --engram 1 --tokens 4e9 --max_minutes 1440 --out $TA_DATA/runs/l1
$TA_PY scripts/eval_agent.py --ckpt $TA_DATA/runs/l1/snap_2.00B.pt # pass@k gate
$TA_PY -m scripts.grpo --ckpt $TA_DATA/runs/main/final.pt --out $TA_DATA/rl/r1 # RL (--engine lockstep = fallback, --opt muon)
setsid nohup bash scripts/rl_after_decay.sh & # waits for the decay, runs the eval gate, then GRPO
$TA_PY -m pytest -q tests/
Measured on the B70
T=2048, 32k tokens/step, compiled, bf16.
| Size | Layers params | tok/s | +Engram tok/s | Tokens per 24h | Tokens/param |
|---|---|---|---|---|---|
| S (512×12) | 33M | 93k | 77k | 8.1B | 241 |
| M (640×16) | 69M | 59k | 51k | 5.1B | 74 |
| L (768×18) | 111M | 46k | 40k | 4.0B | 36 |
| XL (1024×20) | 216M | 29k | – | 2.5B | 11 |
Layout
tiny_agent/:
model.py: model, attention, Engramoptim.py: Muon, Sinkhorn updategenerate.py: KV-cache decodertools.py: sandboxed toolschat.py: ta-v1 formattasks.py: task generator, scripted solver, checkerrollout.py: batched episodes, rewardsdata.py: mixture loader
scripts/ holds the entry points above.
Watching and controlling a run
$TA_PY scripts/status.py # pipeline decisions, loss, val, tok/s, snapshots
touch $TA_DATA/runs/main/DECAY # start the LR decay now
touch $TA_DATA/runs/main/STOP # checkpoint and exit; re-running the same command resumes
The wall-clock budget is saved in the checkpoint, so a resume continues the same 23h clock. Nothing restarts a crashed run automatically.
Notes
Held-out task kinds:
multi_hop,log_countandwrite_factnever appear in the pretraining trajectories, and RL leaves them out by default.eval_agent.pyreports_in_distand_held_outseparately, so template recall can be told apart from tool use that transfers.Seed ranges keep data apart: eval tasks use seeds 0–99,999, RL uses 100,000–999,999, and pretraining trajectories use 1,000,000 and up.
The synthetic reasoning source includes OpenMathInstruct-2 (CC-BY-4.0; solutions written by Llama-3.1-405B). Drop it with
--only tools,groundedif you want only local or script-made data.The Triton/XPU compile needs the Level Zero loader paths set in
env.sh.