tiny-agent-112m / code /README.md
darioooooo0o's picture
tiny-agent-112m: base + RL weights, tokenizer, code, model card
4397e12 verified
|
Raw History Blame Contribute Delete
5.71 kB

tiny-agent

A small from-scratch LLM for research on the Intel Arc Pro B70: about 24 hours of pretraining, then RL. It is built to be weak at recalling facts but good at tool calls and copying found information. Everything runs locally.

Design

Part Choice
Attention DeepSeek V4.1 CSA2, simplified. Compression ratio 1, no indexer. In each group of 4 layers, 1 Full layer writes a shared global K/V and 3 Reuse layers reuse it. Every layer also has its own local K/V with a 128-token sliding window. Each query does one softmax over the union of global and local keys (FlexAttention).
KV cache One global K/V per group plus one 128-slot ring buffer per layer. About 2.7 KB/token at 8k context for size L, roughly 20× smaller than plain attention.
Engram Hashed 2- and 3-gram tables over compressed token ids, with a context-aware sigmoid gate, at layer 1. No short conv, as in V4.1. Behind --engram.
Backbone Dense, ReLU² MLP, QK-norm, RoPE, logit softcap 15, parameter-free RMSNorm, zero-init output projections.
Optimizer Muon for matrices (head-wise on Q). V4.1 momentum + Sinkhorn update for embedding, head and Engram tables.
Schedule Warmup-stable-decay. touch DECAY starts the decay at any point; --max_minutes decays automatically to finish on time.
Tokenizer 32k byte-level BPE trained on the mix, digits split one at a time, Hermes-style <tool_call>/<tool_response> tokens.
Data (stable) Code 40% (openly licensed Stack-Edu: Python, Shell, JS/TS, Markdown, Go, Rust, C), math 25% (FineMath), English 25% (FineWeb-Edu), grounded reading 10%.
Data (decay) More grounded reading, scripted tool-use trajectories, and math with <think> reasoning.
Tools Pi's default four: bash, read, edit, write, plus submit. bash runs in bubblewrap: no network, read-only /usr, ulimits, timeout, tmpfs workspace.
Tasks Workspaces with invented facts: configs in YAML, JSON, TOML or env, plus code, tests, docs, logs, CSVs and a long job log. 15 kinds: lookups, compute (run code, grep -c, count across files, sum a column), multi-step, edit, fix a failing test, write a file, and not-found. The answer is only in the files.
RL GRPO directly on base snapshots, with no SFT stage. Continuous-batching rollouts: each cache slot is refilled the moment its episode ends, and episodes still running carry across updates (at most 2 updates stale) with a per-token importance ratio against the recorded sampling log-prob. Groups without signal are dropped. Rewards and credit follow MiMo-V2.6's released reference profile: 1 if correct, 0.5 if correct but never seen in a tool result; an in-group length penalty on passes when more than half the group passes; a per-turn penalty for malformed, unknown, bad-argument or repeated tool calls; loss averaged per prompt; no KL. Curriculum favors task kinds solved about half the time.

Run

source env.sh
$TA_PY scripts/download.py                 # raw data -> $TA_DATA/raw
$TA_PY scripts/train_tokenizer.py          # tokenizer.json + cid_map.npy (Engram compression)
$TA_PY scripts/tokenize_data.py            # -> $TA_DATA/tok/{train,val}/<source>/*.bin
$TA_PY scripts/gen_synth.py                # synth_tools / synth_grounded / synth_reasoning
$TA_PY scripts/bench_size.py               # tok/s per size -> pick the size
$TA_PY scripts/train.py --size L --engram 1 --tokens 4e9 --max_minutes 1440 --out $TA_DATA/runs/l1
$TA_PY scripts/eval_agent.py --ckpt $TA_DATA/runs/l1/snap_2.00B.pt     # pass@k gate
$TA_PY -m scripts.grpo --ckpt $TA_DATA/runs/main/final.pt --out $TA_DATA/rl/r1   # RL (--engine lockstep = fallback, --opt muon)
setsid nohup bash scripts/rl_after_decay.sh &   # waits for the decay, runs the eval gate, then GRPO
$TA_PY -m pytest -q tests/

Measured on the B70

T=2048, 32k tokens/step, compiled, bf16.

Size Layers params tok/s +Engram tok/s Tokens per 24h Tokens/param
S (512×12) 33M 93k 77k 8.1B 241
M (640×16) 69M 59k 51k 5.1B 74
L (768×18) 111M 46k 40k 4.0B 36
XL (1024×20) 216M 29k – 2.5B 11

Layout

tiny_agent/:

  • model.py: model, attention, Engram
  • optim.py: Muon, Sinkhorn update
  • generate.py: KV-cache decoder
  • tools.py: sandboxed tools
  • chat.py: ta-v1 format
  • tasks.py: task generator, scripted solver, checker
  • rollout.py: batched episodes, rewards
  • data.py: mixture loader

scripts/ holds the entry points above.

Watching and controlling a run

$TA_PY scripts/status.py                    # pipeline decisions, loss, val, tok/s, snapshots
touch $TA_DATA/runs/main/DECAY              # start the LR decay now
touch $TA_DATA/runs/main/STOP               # checkpoint and exit; re-running the same command resumes

The wall-clock budget is saved in the checkpoint, so a resume continues the same 23h clock. Nothing restarts a crashed run automatically.

Notes

  • Held-out task kinds: multi_hop, log_count and write_fact never appear in the pretraining trajectories, and RL leaves them out by default. eval_agent.py reports _in_dist and _held_out separately, so template recall can be told apart from tool use that transfers.

  • Seed ranges keep data apart: eval tasks use seeds 0–99,999, RL uses 100,000–999,999, and pretraining trajectories use 1,000,000 and up.

  • The synthetic reasoning source includes OpenMathInstruct-2 (CC-BY-4.0; solutions written by Llama-3.1-405B). Drop it with --only tools,grounded if you want only local or script-made data.

  • The Triton/XPU compile needs the Level Zero loader paths set in env.sh.