Mike Ravkine PRO
mike-ravkine
AI & ML interests
LLM Research / Development / Evaluation
Recent Activity
liked a model 1 day ago
amd/Instella-MoE-16B-A3B-Think repliedto their post 2 days ago
I went into this expecting to find a ~Q2 garbage dumpster but https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf is a slick feat of QAT engineering.
It is weaker then FP16 on a handful of tasks where quants usually degrade, as per their own paper the loss is "concentrated on sustained chains of reasoning / agentic" and in ReasonScape this bites on Sort, Shuffle and Dates, but counter-acting this are some noticeable improvements to thinking length without accuracy loss on several other tasks (Shapes, Cars).
I haven't had a chance to run the Binary yet, but PQ2 + Bonsai QAT are confirmed to be pretty darn impressive. repliedto their post 2 days ago
I went into this expecting to find a ~Q2 garbage dumpster but https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf is a slick feat of QAT engineering.
It is weaker then FP16 on a handful of tasks where quants usually degrade, as per their own paper the loss is "concentrated on sustained chains of reasoning / agentic" and in ReasonScape this bites on Sort, Shuffle and Dates, but counter-acting this are some noticeable improvements to thinking length without accuracy loss on several other tasks (Shapes, Cars).
I haven't had a chance to run the Binary yet, but PQ2 + Bonsai QAT are confirmed to be pretty darn impressive. 

