File size: 5,705 Bytes
4397e12
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
# tiny-agent

A small from-scratch LLM for research on the Intel Arc Pro B70: about 24 hours of pretraining,
then RL. It is built to be weak at recalling facts but good at tool calls and copying found
information. Everything runs locally.

## Design

| Part | Choice |
|---|---|
| Attention | DeepSeek V4.1 CSA2, simplified. Compression ratio 1, no indexer. In each group of 4 layers, 1 Full layer writes a shared global K/V and 3 Reuse layers reuse it. Every layer also has its own local K/V with a 128-token sliding window. Each query does one softmax over the union of global and local keys (FlexAttention). |
| KV cache | One global K/V per group plus one 128-slot ring buffer per layer. About 2.7 KB/token at 8k context for size L, roughly 20× smaller than plain attention. |
| Engram | Hashed 2- and 3-gram tables over compressed token ids, with a context-aware sigmoid gate, at layer 1. No short conv, as in V4.1. Behind `--engram`. |
| Backbone | Dense, ReLU² MLP, QK-norm, RoPE, logit softcap 15, parameter-free RMSNorm, zero-init output projections. |
| Optimizer | Muon for matrices (head-wise on Q). V4.1 momentum + Sinkhorn update for embedding, head and Engram tables. |
| Schedule | Warmup-stable-decay. `touch DECAY` starts the decay at any point; `--max_minutes` decays automatically to finish on time. |
| Tokenizer | 32k byte-level BPE trained on the mix, digits split one at a time, Hermes-style `<tool_call>`/`<tool_response>` tokens. |
| Data (stable) | Code 40% (openly licensed Stack-Edu: Python, Shell, JS/TS, Markdown, Go, Rust, C), math 25% (FineMath), English 25% (FineWeb-Edu), grounded reading 10%. |
| Data (decay) | More grounded reading, scripted tool-use trajectories, and math with `<think>` reasoning. |
| Tools | Pi's default four: `bash`, `read`, `edit`, `write`, plus `submit`. bash runs in bubblewrap: no network, read-only /usr, ulimits, timeout, tmpfs workspace. |
| Tasks | Workspaces with invented facts: configs in YAML, JSON, TOML or env, plus code, tests, docs, logs, CSVs and a long job log. 15 kinds: lookups, compute (run code, grep -c, count across files, sum a column), multi-step, edit, fix a failing test, write a file, and not-found. The answer is only in the files. |
| RL | GRPO directly on base snapshots, with no SFT stage. Continuous-batching rollouts: each cache slot is refilled the moment its episode ends, and episodes still running carry across updates (at most 2 updates stale) with a per-token importance ratio against the recorded sampling log-prob. Groups without signal are dropped. Rewards and credit follow MiMo-V2.6's released reference profile: 1 if correct, 0.5 if correct but never seen in a tool result; an in-group length penalty on passes when more than half the group passes; a per-turn penalty for malformed, unknown, bad-argument or repeated tool calls; loss averaged per prompt; no KL. Curriculum favors task kinds solved about half the time. |

## Run

```bash
source env.sh
$TA_PY scripts/download.py                 # raw data -> $TA_DATA/raw
$TA_PY scripts/train_tokenizer.py          # tokenizer.json + cid_map.npy (Engram compression)
$TA_PY scripts/tokenize_data.py            # -> $TA_DATA/tok/{train,val}/<source>/*.bin
$TA_PY scripts/gen_synth.py                # synth_tools / synth_grounded / synth_reasoning
$TA_PY scripts/bench_size.py               # tok/s per size -> pick the size
$TA_PY scripts/train.py --size L --engram 1 --tokens 4e9 --max_minutes 1440 --out $TA_DATA/runs/l1
$TA_PY scripts/eval_agent.py --ckpt $TA_DATA/runs/l1/snap_2.00B.pt     # pass@k gate
$TA_PY -m scripts.grpo --ckpt $TA_DATA/runs/main/final.pt --out $TA_DATA/rl/r1   # RL (--engine lockstep = fallback, --opt muon)
setsid nohup bash scripts/rl_after_decay.sh &   # waits for the decay, runs the eval gate, then GRPO
$TA_PY -m pytest -q tests/
```

## Measured on the B70

T=2048, 32k tokens/step, compiled, bf16.

| Size | Layers params | tok/s | +Engram tok/s | Tokens per 24h | Tokens/param |
|---|---|---|---|---|---|
| S (512×12) | 33M | 93k | 77k | 8.1B | 241 |
| M (640×16) | 69M | 59k | 51k | 5.1B | 74 |
| L (768×18) | 111M | 46k | 40k | 4.0B | 36 |
| XL (1024×20) | 216M | 29k | – | 2.5B | 11 |

## Layout

`tiny_agent/`:
- `model.py`: model, attention, Engram
- `optim.py`: Muon, Sinkhorn update
- `generate.py`: KV-cache decoder
- `tools.py`: sandboxed tools
- `chat.py`: ta-v1 format
- `tasks.py`: task generator, scripted solver, checker
- `rollout.py`: batched episodes, rewards
- `data.py`: mixture loader

`scripts/` holds the entry points above.

## Watching and controlling a run

```bash
$TA_PY scripts/status.py                    # pipeline decisions, loss, val, tok/s, snapshots
touch $TA_DATA/runs/main/DECAY              # start the LR decay now
touch $TA_DATA/runs/main/STOP               # checkpoint and exit; re-running the same command resumes
```

The wall-clock budget is saved in the checkpoint, so a resume continues the same 23h clock.
Nothing restarts a crashed run automatically.

## Notes

- Held-out task kinds: `multi_hop`, `log_count` and `write_fact` never appear in the pretraining
  trajectories, and RL leaves them out by default. `eval_agent.py` reports `_in_dist` and `_held_out`
  separately, so template recall can be told apart from tool use that transfers.

- Seed ranges keep data apart: eval tasks use seeds 0–99,999, RL uses 100,000–999,999, and pretraining trajectories use 1,000,000 and up.
- The synthetic reasoning source includes OpenMathInstruct-2 (CC-BY-4.0; solutions written by Llama-3.1-405B). Drop it with `--only tools,grounded` if you want only local or script-made data.
- The Triton/XPU compile needs the Level Zero loader paths set in `env.sh`.