Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
gxcsoccer 
posted an update 5 days ago
Post
107
I wrote up how DeskMind's local decision models work: a 0.8B and a 4B (Qwen3.5 fine-tunes, 8-bit MLX) that answer each desktop step as a multiple-choice question with a probability for every option. The 0.8B answers first, unsure steps go to the 4B, and when a task is ambiguous it asks the user instead of guessing.

The article also has three lessons from about twenty training rounds, including how smoothing the labels to 0.95 squeezed the 0.8B's confidence right onto the routing threshold.

Article: https://huggingface.co/blog/gxcsoccer/deskmind-local-decision-models
Models:
deskmind

Your JevBench table already prices the 0.96 gate.

0.8B alone 0.723, 4B alone 0.835, router 0.797 on the 231 public items. A gate that escalated a blind 70% of items would land at 0.801. Per tier it holds too: original 65 vs 65.1 blind, hard 71 vs 72.1 blind.

So on this set the confidence band carries no routing signal. That is what a 0.95 soft label predicts. The cross-entropy optimum puts the top option at 0.95, so a 0.96 threshold sits above the number the model was trained to output. G14 at 0.94 sat below it, and most steps stayed local at p50 0.57 s.

This assumes JevBench escalates near the desktop's 70%. Your routing records would give the real rate.

One check before the distillation round: what does the 4B alone score on bench v25? If it is also 39/39, the G14 to G18b gain is the 4B answering 70% of steps, not the router choosing well. And a 0.8B distilled from the 4B's probabilities takes the 4B's calibration, Brier 0.269, as its target.

Have you swept the threshold on the public items to see where the router actually beats the blind mixture?