Starlight Small 1

North ML executor. Qwen3.5-0.8B with a causal agent-memory register after every full-attention layer. The register's output projection starts at zero, so the model matches Qwen at step 0.

Benchmark Score Correct
MMLU (5-shot) 45.5% 6385/14042
ARC-Challenge (25-shot) 43.4% 509/1172
HellaSwag (0-shot) 46.9% 4709/10042
GSM8K (5-shot) 32.8% 433/1319

MMLU and ARC-Challenge score the next-token letter on the test split. HellaSwag scores the length-normalized ending on the validation split. GSM8K is greedy generation, up to 192 new tokens, on the test split.

Load agent_register.pt with starlight_arch.py and hook it onto the full-attention layers before generation. model.safetensors alone is the Qwen backbone without the register.

Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for North-ML1/starlight-small-1

Finetuned
(459)
this model