YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

mt92 extract โ€” shipped working set (24 layers + outlier split)

r24 is the checkpoint behind the live artifact: Qwen3-0.6B at lr 8e-5, then the four least useful transformer blocks dropped by angular distance and retrained.

The two things that mattered

All the int8 damage lives in a handful of channels. down_proj inputs peak near 3800 while weights sit near 0.5, and onnxruntime scales activations per tensor, so one shared scale erases everything small. Pulling the 32 widest channels per layer into their own full-precision matmul drops the remaining peak to about 10 and restores full-precision quality for ~3 MB. split_onnx.py does it with Gather/MatMul/Add only, so validate_graph stays happy.

Depth is nearly free on this task and it is what the cost coordinate measures. Cost is prefill of a fixed 249-token probe. 28 -> 24 layers costs nothing measurable (0.9170 -> 0.9139) and 24 -> 20 costs 1.6 points (0.8984).

artifact micro-F1 (400) probe TTFT size
28L, down_proj in fp32 (previous) 0.9183 1579 ms 984 MiB
28L, outlier split 0.9170 721 ms 736 MiB
24L, outlier split (shipped) 0.9139 602 ms 675 MiB

Full-precision down_proj was buying the same quality at 2.6x the latency.

Dead ends

Qwen3-1.7B: even at 0.975 quality its cost puts it in a thin frontier slice -- 0.22 share against the real round-1237 field, versus 0.5-0.88 for this artifact. Static QDQ calibration, SmoothQuant migration, weight averaging, and a QAT phase on a properly tuned base all measured worse than doing nothing.

Rebuild: ./build_final.sh <model_dir> 32.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support