Starlight Small 1
North ML executor. Qwen3.5-0.8B with a causal agent-memory register after every full-attention layer. The register's output projection starts at zero, so the model matches Qwen at step 0.
| Benchmark | Score | Correct |
|---|---|---|
| MMLU (5-shot) | 45.5% | 6385/14042 |
| ARC-Challenge (25-shot) | 43.4% | 509/1172 |
| HellaSwag (0-shot) | 46.9% | 4709/10042 |
| GSM8K (5-shot) | 32.8% | 433/1319 |
MMLU and ARC-Challenge score the next-token letter on the test split. HellaSwag scores the length-normalized ending on the validation split. GSM8K is greedy generation, up to 192 new tokens, on the test split.
Load agent_register.pt with starlight_arch.py and hook it onto the full-attention layers before generation. model.safetensors alone is the Qwen backbone without the register.
- Downloads last month
- -