Starlight Pro 1
North ML executor. Qwen3.5-2B with a causal agent-memory register after every full-attention layer. The register's output projection starts at zero, so the model matches Qwen at step 0.
| Benchmark | Score | Correct |
|---|---|---|
| MMLU (5-shot) | 59.2% | 8311/14042 |
| ARC-Challenge (25-shot) | 80.0% | 938/1172 |
| HellaSwag (0-shot) | 57.2% | 5742/10042 |
MMLU and ARC-Challenge score the next-token letter on the test split. HellaSwag scores the length-normalized ending on the validation split.
Load agent_register.pt with starlight_arch.py and hook it onto the full-attention layers before generation. model.safetensors alone is the Qwen backbone without the register.
- Downloads last month
- -