Commit History

v4: two-lane decoder (two pieces per infer call, grouped-query attention)
b8b7f8c
verified

lunks commited on

Declare the base-model relation as quantized
c54339a
verified

lunks commited on

Model card in the style of the MLX conversion's
f34b38d
verified

lunks commited on

Confucius4-R2T2 converted for Core ML on the Apple Neural Engine (encoder fp16, decoder LUT8 with 128-row prefill and verify head)
d0e38c6
verified

lunks commited on

initial commit
e1c4fea
verified

lunks commited on