GoldWorm Coder — model bundle (gen0)

A self-contained local-first coding agent whose brain is this repository's model: the goldworm byte-level upcycled MoE. Small aux models are allowed ONLY as sub-jobs (the output layer that phrases answers is qwen2.5-0.5b-GGUF via llama.cpp, never the brain; the agent always ranks with goldworm and verifies by running the tests).

Honest scope: the brain is a 38M-param byte-LM (upcycled 4-expert MoE, ~113.6M params with experts). It is measured, not hyped. Everything below is number-first, run-through-the-project's-own-gate evidence.

What makes goldworm special (measured, not claimed)

  • Byte-level MoE with REAL domain routing: 100% held-out tail router accuracy, gate peak 0.856, bijective topic→expert assignment. The mixture of 4 topic experts beats the best single expert (scaled verify_moe.py matrix: MoE mean tail 2.2385 vs best-single 2.2700, oracle 1.0734 for stories).
  • Gate-validated improvement, never wishful: every champion push goes through rsi/gate.py (frozen thresholds, hash-chained audit log) — determinism proof, drift/KL guardrails, router identity, entropy collapse, contamination canary, Goodhart audit-set guard. Two strong candidates in Phase 5 were rejected by the gate even though they improved most axes (the ctx-1024 candidate still regressed TS-val; the code-heavy candidate dropped the router peak). The gate is the single door — that's the project's honesty, engineered.
  • Score-then-verify agent: deterministic candidate generation, goldworm BPC ranking, strict red-to-green test verification in a throwaway git worktree. The agent never trusts a patch until a test goes red→green.

Measured numbers (this exact bundle)

metric value protocol
TinyStories val BPC 0.8109 train_scaled.eval_bpc, wins=128
mean topic-tail BPC (stories/code/math/wiki) 2.2385 verified via verify_moe.bpc
mean gate peak 0.856 (bijective) verify_moe.gate_dist
router held-out accuracy 4/4 (100%) expert_router

Agentic skills (product, 28/28 sub-checks, this session)

skill result
classify 10/10 exact (fix/test/chat/help/generate/inspect)
fix (red→green) 2/2 verified_by_tests
test 2/2
inspect/triage 4/4
chat (qwen2.5-0.5b output layer) 5/5 (greet, 2+2, 7×8, capital, session memory)
remediation 5/5 (clean/reset/confirm/401 gates)

goldworm-hybrid vs qwen2.5:0.5b alone (same skills)

skill goldworm qwen-alone
classify 10/10 4/10
fix 2/2 1/2 (flaky)
inspect 4/4 3/4
chat 5/5 5/5 (identical prompts)

ARC-AGI probe (3 pinned games, disclosed selection)

0/3 for both arms (random ≈ 1/16-17). ARC grids are NOT this model's strength — recorded, not hidden.

Known honest limits

The underlying model: 0/5 arithmetic, ARC chance level, confident-but-wrong on free-form generation. The product works because it never asks a weak model to write — it asks goldworm to rank, route, and verify.

How to use

  1. Load via the project's load path: rsi/gate.load_moe("moe.pt", dev) — never trust the checkpoint's embedded cfg (it's stale).
  2. The full app: the self-contained Linux bundle (start.sh) — see the Gold_Lab Space for a showcase + the agent source + benchmarks.

Contents

  • moe.pt — champion MoE checkpoint
  • experts/{stories,code,math,wiki}.pt — topic experts (d512/L12)
  • active_bundle.json, manifest.json — atomic promotion pointer + hashes
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support