|
Download code/README.md from darioooooo0o/tiny-agent-112m: direct link, hf CLI and curl.
- Browser
- Download file 5.71 kB
-
https://huggingface.co/darioooooo0o/tiny-agent-112m/resolve/main/code/README.md
- Command line
-
hf download hf://darioooooo0o/tiny-agent-112m/code/README.md
-
curl -L -o README.md https://huggingface.co/darioooooo0o/tiny-agent-112m/resolve/main/code/README.md
5.71 kB
| # tiny-agent | |
| A small from-scratch LLM for research on the Intel Arc Pro B70: about 24 hours of pretraining, | |
| then RL. It is built to be weak at recalling facts but good at tool calls and copying found | |
| information. Everything runs locally. | |
| ## Design | |
| | Part | Choice | | |
| |---|---| | |
| | Attention | DeepSeek V4.1 CSA2, simplified. Compression ratio 1, no indexer. In each group of 4 layers, 1 Full layer writes a shared global K/V and 3 Reuse layers reuse it. Every layer also has its own local K/V with a 128-token sliding window. Each query does one softmax over the union of global and local keys (FlexAttention). | | |
| | KV cache | One global K/V per group plus one 128-slot ring buffer per layer. About 2.7 KB/token at 8k context for size L, roughly 20× smaller than plain attention. | | |
| | Engram | Hashed 2- and 3-gram tables over compressed token ids, with a context-aware sigmoid gate, at layer 1. No short conv, as in V4.1. Behind `--engram`. | | |
| | Backbone | Dense, ReLU² MLP, QK-norm, RoPE, logit softcap 15, parameter-free RMSNorm, zero-init output projections. | | |
| | Optimizer | Muon for matrices (head-wise on Q). V4.1 momentum + Sinkhorn update for embedding, head and Engram tables. | | |
| | Schedule | Warmup-stable-decay. `touch DECAY` starts the decay at any point; `--max_minutes` decays automatically to finish on time. | | |
| | Tokenizer | 32k byte-level BPE trained on the mix, digits split one at a time, Hermes-style `<tool_call>`/`<tool_response>` tokens. | | |
| | Data (stable) | Code 40% (openly licensed Stack-Edu: Python, Shell, JS/TS, Markdown, Go, Rust, C), math 25% (FineMath), English 25% (FineWeb-Edu), grounded reading 10%. | | |
| | Data (decay) | More grounded reading, scripted tool-use trajectories, and math with `<think>` reasoning. | | |
| | Tools | Pi's default four: `bash`, `read`, `edit`, `write`, plus `submit`. bash runs in bubblewrap: no network, read-only /usr, ulimits, timeout, tmpfs workspace. | | |
| | Tasks | Workspaces with invented facts: configs in YAML, JSON, TOML or env, plus code, tests, docs, logs, CSVs and a long job log. 15 kinds: lookups, compute (run code, grep -c, count across files, sum a column), multi-step, edit, fix a failing test, write a file, and not-found. The answer is only in the files. | | |
| | RL | GRPO directly on base snapshots, with no SFT stage. Continuous-batching rollouts: each cache slot is refilled the moment its episode ends, and episodes still running carry across updates (at most 2 updates stale) with a per-token importance ratio against the recorded sampling log-prob. Groups without signal are dropped. Rewards and credit follow MiMo-V2.6's released reference profile: 1 if correct, 0.5 if correct but never seen in a tool result; an in-group length penalty on passes when more than half the group passes; a per-turn penalty for malformed, unknown, bad-argument or repeated tool calls; loss averaged per prompt; no KL. Curriculum favors task kinds solved about half the time. | | |
| ## Run | |
| ```bash | |
| source env.sh | |
| $TA_PY scripts/download.py # raw data -> $TA_DATA/raw | |
| $TA_PY scripts/train_tokenizer.py # tokenizer.json + cid_map.npy (Engram compression) | |
| $TA_PY scripts/tokenize_data.py # -> $TA_DATA/tok/{train,val}/<source>/*.bin | |
| $TA_PY scripts/gen_synth.py # synth_tools / synth_grounded / synth_reasoning | |
| $TA_PY scripts/bench_size.py # tok/s per size -> pick the size | |
| $TA_PY scripts/train.py --size L --engram 1 --tokens 4e9 --max_minutes 1440 --out $TA_DATA/runs/l1 | |
| $TA_PY scripts/eval_agent.py --ckpt $TA_DATA/runs/l1/snap_2.00B.pt # pass@k gate | |
| $TA_PY -m scripts.grpo --ckpt $TA_DATA/runs/main/final.pt --out $TA_DATA/rl/r1 # RL (--engine lockstep = fallback, --opt muon) | |
| setsid nohup bash scripts/rl_after_decay.sh & # waits for the decay, runs the eval gate, then GRPO | |
| $TA_PY -m pytest -q tests/ | |
| ``` | |
| ## Measured on the B70 | |
| T=2048, 32k tokens/step, compiled, bf16. | |
| | Size | Layers params | tok/s | +Engram tok/s | Tokens per 24h | Tokens/param | | |
| |---|---|---|---|---|---| | |
| | S (512×12) | 33M | 93k | 77k | 8.1B | 241 | | |
| | M (640×16) | 69M | 59k | 51k | 5.1B | 74 | | |
| | L (768×18) | 111M | 46k | 40k | 4.0B | 36 | | |
| | XL (1024×20) | 216M | 29k | – | 2.5B | 11 | | |
| ## Layout | |
| `tiny_agent/`: | |
| - `model.py`: model, attention, Engram | |
| - `optim.py`: Muon, Sinkhorn update | |
| - `generate.py`: KV-cache decoder | |
| - `tools.py`: sandboxed tools | |
| - `chat.py`: ta-v1 format | |
| - `tasks.py`: task generator, scripted solver, checker | |
| - `rollout.py`: batched episodes, rewards | |
| - `data.py`: mixture loader | |
| `scripts/` holds the entry points above. | |
| ## Watching and controlling a run | |
| ```bash | |
| $TA_PY scripts/status.py # pipeline decisions, loss, val, tok/s, snapshots | |
| touch $TA_DATA/runs/main/DECAY # start the LR decay now | |
| touch $TA_DATA/runs/main/STOP # checkpoint and exit; re-running the same command resumes | |
| ``` | |
| The wall-clock budget is saved in the checkpoint, so a resume continues the same 23h clock. | |
| Nothing restarts a crashed run automatically. | |
| ## Notes | |
| - Held-out task kinds: `multi_hop`, `log_count` and `write_fact` never appear in the pretraining | |
| trajectories, and RL leaves them out by default. `eval_agent.py` reports `_in_dist` and `_held_out` | |
| separately, so template recall can be told apart from tool use that transfers. | |
| - Seed ranges keep data apart: eval tasks use seeds 0–99,999, RL uses 100,000–999,999, and pretraining trajectories use 1,000,000 and up. | |
| - The synthetic reasoning source includes OpenMathInstruct-2 (CC-BY-4.0; solutions written by Llama-3.1-405B). Drop it with `--only tools,grounded` if you want only local or script-made data. | |
| - The Triton/XPU compile needs the Level Zero loader paths set in `env.sh`. | |