Starlight Pro 1

North ML executor. Qwen3.5-2B with a causal agent-memory register after every full-attention layer. The register's output projection starts at zero, so the model matches Qwen at step 0.

Benchmark Score Correct
MMLU (5-shot) 59.2% 8311/14042
ARC-Challenge (25-shot) 80.0% 938/1172
HellaSwag (0-shot) 57.2% 5742/10042

MMLU and ARC-Challenge score the next-token letter on the test split. HellaSwag scores the length-normalized ending on the validation split.

Load agent_register.pt with starlight_arch.py and hook it onto the full-attention layers before generation. model.safetensors alone is the Qwen backbone without the register.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for North-ML1/starlight-pro-1

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(456)
this model