Upload MIMO.md with huggingface_hub
Browse files
MIMO.md
ADDED
|
@@ -0,0 +1,181 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# MIMO β MiMo-V2.6-Distill-Qwen-9B workbench (2026-09-22)
|
| 2 |
+
|
| 3 |
+
Xiaomi MiMo SFT distillate of Qwen3.5-9B (77.4B SFT tokens, 27.2B
|
| 4 |
+
loss-bearing; 27% visual). Agentic generalist, not a code specialist.
|
| 5 |
+
SFT β continued-pretraining scale β deltas vs base are huge
|
| 6 |
+
(SWE Pro 32β44.6, TerminalBench 27β37.1).
|
| 7 |
+
|
| 8 |
+
Donors/files in `/mnt/Vsio/Downloads/`:
|
| 9 |
+
- `MiMo-V2.6-Distill-Qwen-9B-bf16.gguf` (17 GB)
|
| 10 |
+
- `MiMo-V2.6-Distill-Qwen-9B-imatrix.gguf` (own, 4.9 MB)
|
| 11 |
+
- `MiMo-V2.6-Distill-Qwen-9B-calibration-v6.txt` (own cal, 1.1 MB)
|
| 12 |
+
- `MiMo-V2.6-Distill-Qwen-9B-MERNIK-5100.gguf` (SMAPE, own lens)
|
| 13 |
+
- `MiMo-V2.6-Distill-Qwen-9B-Q6_K.gguf` (stock baseline)
|
| 14 |
+
- `RINIQ-N1-BF16.gguf` (Ox base + MiMo 15,16,17, via fuse_layers.py)
|
| 15 |
+
- `RINIQ-N1m-BF16.gguf` (MiMo base + Ox 15,16,17, mirror)
|
| 16 |
+
- `RINIQ-N1-MERNIK-5100.gguf` / `RINIQ-N1m-MERNIK-5100.gguf` (SMAPE-5100, dual-max lens)
|
| 17 |
+
|
| 18 |
+
MiMo ships its own jinja (macro-based, 3.9 KB) β NOT the Qwen
|
| 19 |
+
froggeric template (16 KB). `--jinja` works natively on all runs.
|
| 20 |
+
|
| 21 |
+
Sampling (Qwen family protocol): temp 1.0 / top_p 0.95 / top_k 20,
|
| 22 |
+
presence 0.0, max_tokens 2048. No Gemma leftovers (those lived only in
|
| 23 |
+
chain scripts). HE runner now logs battery time (start stamp + min/task).
|
| 24 |
+
|
| 25 |
+
## Imatrix compass (MiMo vs Ox vs Neo, 248 common tensors)
|
| 26 |
+
|
| 27 |
+
| Pair | Pearson | Rank agree |
|
| 28 |
+
|:-----|--------:|-----------:|
|
| 29 |
+
| mimoβox | 0.9989 | 0.9975 |
|
| 30 |
+
| mimoβneo | 0.9992 | 0.9977 |
|
| 31 |
+
| oxβneo | 0.9999 | 0.9999 |
|
| 32 |
+
|
| 33 |
+
Same law as FUSION: recipe dominates, finetune nearly invisible. But the
|
| 34 |
+
divergence clusters mid-layer FFN **15,16,17,19,13,27,14,18,23,11** β the
|
| 35 |
+
weight-compass "finetune soul" zone. Scale: MiMo cal runs 1.6Γ hotter
|
| 36 |
+
(mean 1.9e5 vs 1.2e5) β quantize on its own imatrix, never mixed.
|
| 37 |
+
(Dual-max(Ox,MiMo) is a fiction: max always picks MiMo's hotter scale.
|
| 38 |
+
N1 dry-run with dual lens == pure-MiMo geometry bit for bit.)
|
| 39 |
+
|
| 40 |
+
## Verdicts (slow ring: HumanEval pass@1, temp 1.0 / top_p 0.95 / top_k 20, presence 0.0, max 2048)
|
| 41 |
+
|
| 42 |
+
| Build | Size | PPL (ctx1024) | HE pass@1 | HE+ |
|
| 43 |
+
|:------|-----:|:-------------:|:---------:|:---:|
|
| 44 |
+
| MiMo-5100-SMAPE (own imatrix) | 5.0 GB | 8.7435 | 70.12% (115/164) | 66.5% |
|
| 45 |
+
| MiMo-4500-SMAPE (own imatrix, allow-q3) | 4.5 GB | 10.5343 | 29.88% (49/164) | 26.2 β lava zone (Neo-3650-Q2: 23.78%), sub-4 at ~4GB |
|
| 46 |
+
| MiMo-4500-SMSE (own imatrix, allow-q3) | 4.5 GB | **9.4092** | **54.88% (90/164)** | **53.0** β rescue +25pp over SMAPE, TOX-like save |
|
| 47 |
+
| MiMo-5100-SMSE (own imatrix) | 5.0 GB | 8.7389 | 71.34% (117/164) | 65.9 β 3 empties. Duel 0/3 columns (HE p=0.88, HE+ p=0.76 against, empties p=0.13): the +1.2pp is NOISE, crown demoted 2026-09-28. Decisiveness edge (3 vs 7) stands as observation |
|
| 48 |
+
| MiMo-6500-MSE (own imatrix) | 6.5 GB | 8.7655 | **75.00% (123/164)** | **68.9** β beats Q6_K by +3pp at β0.7 GB |
|
| 49 |
+
| MiMo-6500-SMSE (own imatrix) | 6.5 GB | 8.7275 | 72.56% (119/164) | 67.1 β LOSES to MSE by 2.4pp: PPL tie, verdict fail. Burn wins; refinement: SMSE had MORE Q8 (125 vs 108) yet lost β at big budgets the MIDDLE decides (MSE Q6 67 vs SMSE 41) |
|
| 50 |
+
| MiMo-Q6_K (stock) | 7.2 GB | 9.1362 | 71.95% (118/164) | 66.5% |
|
| 51 |
+
| Ox-SMAPE-5100 (ref) | 5.0 GB | 7.5670 | 88.41% (145/164) | β |
|
| 52 |
+
| Neo-SMAPE-5100 (ref) | 5.0 GB | 7.8012 | 82.93% (136/164) | β |
|
| 53 |
+
|
| 54 |
+
Fusions (N1/N1m) live in `RINIQ-NEXT.md`.
|
| 55 |
+
|
| 56 |
+
Geometry tells the rescue story (audit_tiers):
|
| 57 |
+
|
| 58 |
+
| Tier | 4500-SMAPE | 4500-SMSE | 5100-SMAPE | 5100-SMSE |
|
| 59 |
+
|:-----|----------:|----------:|----------:|----------:|
|
| 60 |
+
| sub-Q4 | **107** (104ΓIQ2_XXS!) | **66** (59ΓIQ2_XXS, rest IQ3/IQ4) | 0 | 0 |
|
| 61 |
+
| Q4 | 4 | 6 | 206 | 197 |
|
| 62 |
+
| Q5 | 2 | 27 | 11 | 29 |
|
| 63 |
+
| Q6 | 0 | 82 | 11 | 0 |
|
| 64 |
+
| Q8 penthouses | **118** | **0** | 0 | 0 |
|
| 65 |
+
| HE | 29.88% | 54.88% | 70.12% | 71.34% |
|
| 66 |
+
|
| 67 |
+
SMSE never affords a penthouse (Q8 = 0 at both budgets), halves the
|
| 68 |
+
dungeon, and rebuilds the middle (Q5+Q6 = 109 @4500). "Penthouses only
|
| 69 |
+
if affordable" β in numbers. Bonus catch: SMAPE-4500 evicted even the
|
| 70 |
+
free F16 residents (22β0), SMSE-4500 kept 8 β total war vs discipline.
|
| 71 |
+
Open risk @6500: MSE's 75% rides 108 Q8 penthouses; if SMSE arrives
|
| 72 |
+
with zero Q8 there and loses, the law extends to "penthouses required
|
| 73 |
+
at big budgets". Dry-run geometry at quant time will tell first.
|
| 74 |
+
|
| 75 |
+
Dry-run geometry @5100 (SMAPE, own lens): F16 177 / Q4 226 / Q5 13 /
|
| 76 |
+
Q6 11 / Q8 0 β all floor, zero penthouses (same as Ox-SMAPE-5100 shape).
|
| 77 |
+
|
| 78 |
+
Empties: MiMo 7 (Ox-like decisiveness, not Neo hesitation).
|
| 79 |
+
Speed: MiMo ~20 s/task (vs 40+ Gemma-12B, 60+ RINIQ duels) β decisive,
|
| 80 |
+
no thinking-chewing. Timer now logged per battery.
|
| 81 |
+
Battery times: N1 77.1 min (28.2 s/task); MiMo-5100 ~50 min (~18 s/task).
|
| 82 |
+
Fusion decisiveness (N1: 24 empties, seam friction) β `RINIQ-NEXT.md`.
|
| 83 |
+
|
| 84 |
+
HE+ rescore runs after every battery (CPU, EvalPlus 80Γ) β strictness
|
| 85 |
+
column. Layer swaps (15β17) may move other benchmarks (SWE, terminal,
|
| 86 |
+
vision) in either direction β those arenas are queued, not claimed.
|
| 87 |
+
If a build behaves weird on your hardware, open a discussion with setup
|
| 88 |
+
+ task β every scar goes in the ledger.
|
| 89 |
+
|
| 90 |
+
## Field notes (manual testing, 2026-09-22)
|
| 91 |
+
|
| 92 |
+
- Thinks a lot: long reasoning traces, lower t/s than OxCoder but FEELS
|
| 93 |
+
faster (decisive, no hesitation). Part of the gap was CPU contention
|
| 94 |
+
on the test box β clean comparison gives MiMo ~+1 tok/s over RINIQ/Ox.
|
| 95 |
+
- Agent-scaffold leak: in the pi agent framework, a bare "Privet" (hello
|
| 96 |
+
in Russian) makes it confabulate a task (portfolio brief) β operator
|
| 97 |
+
aborted manually upon noticing ("Operation aborted" was the operator,
|
| 98 |
+
not the model).
|
| 99 |
+
Other frameworks fine; RINIQ/Ox/Neo never do this. Read: MiMo's agent
|
| 100 |
+
training (TerminalBench/Toolathlon scaffolds) misfires on agent-style
|
| 101 |
+
system prompts β training artifact, harness-dependent, not quant damage.
|
| 102 |
+
- Temp sensitivity: temp 0.6 breaks tool-call format adherence (confuses
|
| 103 |
+
file-create vs bash); temp 1.0 works. Vendor temp 1.0 mandatory for
|
| 104 |
+
agentic tasks β lower temp commits to the wrong pattern confidently.
|
| 105 |
+
- `--reasoning-effort` WORKS (unique in family): low β <5k thinking
|
| 106 |
+
tokens, xhigh β ~15k. Controllable depth/speed dial out of the box.
|
| 107 |
+
Ox/Neo ignore these flags (thinking stripped).
|
| 108 |
+
- 3D-snake field test, first attempt BROKEN (critical: wall.push coords,
|
| 109 |
+
isHead args, food destructuring [] vs {}, drawObj syntax, undefined
|
| 110 |
+
dir0/tSince; logic: matrices rewritten, WebGL buffers, input).
|
| 111 |
+
OxCoder baseline: wrote 3D models but NEVER got it running.
|
| 112 |
+
Iterations-to-working TBD.
|
| 113 |
+
- low-effort FIRST TRY: WORKING 3D snake (renders, orbit camera, Russian
|
| 114 |
+
UI, score/game-over flow; not quite playable). Total 11k tokens vs
|
| 115 |
+
17k for the broken xhigh attempt.
|
| 116 |
+
CONFOUND (honest): prompts differed by ONE word (originally in
|
| 117 |
+
Russian) β xhigh got "make an HTML 3D snake...", low got "make a
|
| 118 |
+
WORKING HTML 3D snake...".
|
| 119 |
+
The win may belong to the word, not the effort level. Unconfounded
|
| 120 |
+
A/B (same prompt, low vs xhigh) still open.
|
| 121 |
+
- opencode harness: 400-line file, noticed a bug immediately, fixing
|
| 122 |
+
it himself. Harness-independent thinking, second framework confirmed.
|
| 123 |
+
- RINIQ-N1 field test: thinking depth lands BETWEEN parents (Ox shallow
|
| 124 |
+
< N1 medium < MiMo deep) β mid-layer blocks 15β17 appear to govern
|
| 125 |
+
reasoning depth. 3D-snake attempts: 1st = 2D game, 2nd = semi-working
|
| 126 |
+
3D, 3rd = working snake (ultra slow but works), each attempt <4k
|
| 127 |
+
tokens. N1m report pending.
|
| 128 |
+
- Test conditions: llama-server web UI, normal sampling (temp 1.0),
|
| 129 |
+
no looping observed. First runs only β slight-breakage chance noted.
|
| 130 |
+
|
| 131 |
+
## Published-code duel (different harnesses β theater, same-harness HE decides)
|
| 132 |
+
|
| 133 |
+
| Bench | OxCoder-9B | MiMo-Distill |
|
| 134 |
+
|:------|-----------:|-------------:|
|
| 135 |
+
| SWE Verified | 73.5 | 61.1 (avg@3) |
|
| 136 |
+
| SWE Pro | 49.1 | 44.6 (avg@3) |
|
| 137 |
+
| TerminalBench 2.1 | 49.6β50.8 | 37.1 |
|
| 138 |
+
| GPQA Diamond | 86.9 | ??? (unreported) |
|
| 139 |
+
|
| 140 |
+
Base Qwen3.5-9B itself differs between the two tables (SWE-V 53.2 vs
|
| 141 |
+
60.0) β harness strictness (OpenHands + anti-hacking, no net) vs avg@3
|
| 142 |
+
inflation. Cross-table comparison is void; only same-harness duels count.
|
| 143 |
+
House rule: HE is used for comparison, not for score β if a small hybrid
|
| 144 |
+
scores roughly like stock Q6 on the same harness, most capabilities likely
|
| 145 |
+
survived quantization. The verdict column certifies preservation, not rank.
|
| 146 |
+
|
| 147 |
+
Field temp law (Ox-SMSE-6500, 3D-snake first-try): temp 0.0 = working code
|
| 148 |
+
at once; 0.6/1.0 needed retries. Hypothesis: low temp for codegen,
|
| 149 |
+
high temp for agents/reasoning. (MiMo agentic tasks mandate temp 1.0 β
|
| 150 |
+
same split from the other side.)
|
| 151 |
+
|
| 152 |
+
SRIQ-Q6 (LoRA-child of MiMo, Chinese traces): HE 64.63% (106/164) at
|
| 153 |
+
temp 1.0 vs **78.66% (129/164) at temp 0.0** β greedy moves the VERDICT
|
| 154 |
+
+14pp here (2/3 columns significant), decisiveness flat (6 vs 8 empties).
|
| 155 |
+
CORRECTION 2026-09-28 (manifest duel falsified the old "identical score"
|
| 156 |
+
entry β I never counted the t0 file, scar kept). Below MiMo-Q6 (71.95%)
|
| 157 |
+
at temp 1.0; PPL skipped (wiki canary invalid for Chinese reasoning).
|
| 158 |
+
Own SRIQ imatrix done.
|
| 159 |
+
SRIQ-5100-SMSE (own imatrix): 63.41% (104/164), HE+ 56.7, **1 empty** β
|
| 160 |
+
most decisive build in the lab, but β1.2pp vs Q6: mirror image of MiMo
|
| 161 |
+
(+1.2pp). SMSE buys decisiveness everywhere, verdicts are family-local.
|
| 162 |
+
|
| 163 |
+
## Scars (chain discipline)
|
| 164 |
+
|
| 165 |
+
- 2026-09-22: heavy CPU quant launched parallel to GPU HE battery β RAM
|
| 166 |
+
pressure β server death, 30-min battery progress DISCARDED. Second time
|
| 167 |
+
(same kill took Q6-HE at task 60). Rule (README:118, "GPU is a strict
|
| 168 |
+
queue") now enforced as: one heavy job at a time, sequential chains
|
| 169 |
+
only (chain5: PPLΓ3 then HEΓ3).
|
| 170 |
+
- earlyoom (root, main rig) unkillable without local sudo (root locked
|
| 171 |
+
after wrong attempts) β lives on, chain discipline is the shield.
|
| 172 |
+
- Server deaths mid-long-run (tasks 25/30/101/110) on 8 GB VRAM:
|
| 173 |
+
6.4 GB weights + KV c4096 + fragmentation β random OOM. Mitigations:
|
| 174 |
+
`--cache-reuse 256`, `-c 4096` (c8192 KV doesn't fit: create_context
|
| 175 |
+
fail, proven). c8192 works on 9B/5 GB builds (RINIQ era).
|
| 176 |
+
|
| 177 |
+
## Next
|
| 178 |
+
|
| 179 |
+
- [x] Q6_K stock baseline: quant done, PPL+HE in chain5
|
| 180 |
+
- [ ] Fusions (N1/N1m), weight compass, LCB: see `RINIQ-NEXT.md`
|
| 181 |
+
- [ ] LiveCodeBench v6 (backlog): MiMo's arena (SWE/terminal), Ox's too
|