GoldWorm Coder — model bundle (gen0)
A self-contained local-first coding agent whose brain is this repository's model: the goldworm byte-level upcycled MoE. Small aux models are allowed ONLY as sub-jobs (the output layer that phrases answers is qwen2.5-0.5b-GGUF via llama.cpp, never the brain; the agent always ranks with goldworm and verifies by running the tests).
Honest scope: the brain is a 38M-param byte-LM (upcycled 4-expert MoE, ~113.6M params with experts). It is measured, not hyped. Everything below is number-first, run-through-the-project's-own-gate evidence.
What makes goldworm special (measured, not claimed)
- Byte-level MoE with REAL domain routing: 100% held-out tail router
accuracy, gate peak 0.856, bijective topic→expert assignment. The mixture of
4 topic experts beats the best single expert (scaled
verify_moe.pymatrix: MoE mean tail 2.2385 vs best-single 2.2700, oracle 1.0734 for stories). - Gate-validated improvement, never wishful: every champion push goes through
rsi/gate.py(frozen thresholds, hash-chained audit log) — determinism proof, drift/KL guardrails, router identity, entropy collapse, contamination canary, Goodhart audit-set guard. Two strong candidates in Phase 5 were rejected by the gate even though they improved most axes (the ctx-1024 candidate still regressed TS-val; the code-heavy candidate dropped the router peak). The gate is the single door — that's the project's honesty, engineered. - Score-then-verify agent: deterministic candidate generation, goldworm BPC ranking, strict red-to-green test verification in a throwaway git worktree. The agent never trusts a patch until a test goes red→green.
Measured numbers (this exact bundle)
| metric | value | protocol |
|---|---|---|
| TinyStories val BPC | 0.8109 | train_scaled.eval_bpc, wins=128 |
| mean topic-tail BPC (stories/code/math/wiki) | 2.2385 | verified via verify_moe.bpc |
| mean gate peak | 0.856 (bijective) | verify_moe.gate_dist |
| router held-out accuracy | 4/4 (100%) | expert_router |
Agentic skills (product, 28/28 sub-checks, this session)
| skill | result |
|---|---|
| classify | 10/10 exact (fix/test/chat/help/generate/inspect) |
| fix (red→green) | 2/2 verified_by_tests |
| test | 2/2 |
| inspect/triage | 4/4 |
| chat (qwen2.5-0.5b output layer) | 5/5 (greet, 2+2, 7×8, capital, session memory) |
| remediation | 5/5 (clean/reset/confirm/401 gates) |
goldworm-hybrid vs qwen2.5:0.5b alone (same skills)
| skill | goldworm | qwen-alone |
|---|---|---|
| classify | 10/10 | 4/10 |
| fix | 2/2 | 1/2 (flaky) |
| inspect | 4/4 | 3/4 |
| chat | 5/5 | 5/5 (identical prompts) |
ARC-AGI probe (3 pinned games, disclosed selection)
0/3 for both arms (random ≈ 1/16-17). ARC grids are NOT this model's strength — recorded, not hidden.
Known honest limits
The underlying model: 0/5 arithmetic, ARC chance level, confident-but-wrong on free-form generation. The product works because it never asks a weak model to write — it asks goldworm to rank, route, and verify.
How to use
- Load via the project's load path:
rsi/gate.load_moe("moe.pt", dev)— never trust the checkpoint's embedded cfg (it's stale). - The full app: the self-contained Linux bundle (start.sh) — see the
Gold_LabSpace for a showcase + the agent source + benchmarks.
Contents
moe.pt— champion MoE checkpointexperts/{stories,code,math,wiki}.pt— topic experts (d512/L12)active_bundle.json,manifest.json— atomic promotion pointer + hashes