wepiqx commited on
Commit
3a9800a
Β·
verified Β·
1 Parent(s): 10c54af

Upload MIMO.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. MIMO.md +181 -0
MIMO.md ADDED
@@ -0,0 +1,181 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # MIMO β€” MiMo-V2.6-Distill-Qwen-9B workbench (2026-09-22)
2
+
3
+ Xiaomi MiMo SFT distillate of Qwen3.5-9B (77.4B SFT tokens, 27.2B
4
+ loss-bearing; 27% visual). Agentic generalist, not a code specialist.
5
+ SFT β‰ˆ continued-pretraining scale β€” deltas vs base are huge
6
+ (SWE Pro 32β†’44.6, TerminalBench 27β†’37.1).
7
+
8
+ Donors/files in `/mnt/Vsio/Downloads/`:
9
+ - `MiMo-V2.6-Distill-Qwen-9B-bf16.gguf` (17 GB)
10
+ - `MiMo-V2.6-Distill-Qwen-9B-imatrix.gguf` (own, 4.9 MB)
11
+ - `MiMo-V2.6-Distill-Qwen-9B-calibration-v6.txt` (own cal, 1.1 MB)
12
+ - `MiMo-V2.6-Distill-Qwen-9B-MERNIK-5100.gguf` (SMAPE, own lens)
13
+ - `MiMo-V2.6-Distill-Qwen-9B-Q6_K.gguf` (stock baseline)
14
+ - `RINIQ-N1-BF16.gguf` (Ox base + MiMo 15,16,17, via fuse_layers.py)
15
+ - `RINIQ-N1m-BF16.gguf` (MiMo base + Ox 15,16,17, mirror)
16
+ - `RINIQ-N1-MERNIK-5100.gguf` / `RINIQ-N1m-MERNIK-5100.gguf` (SMAPE-5100, dual-max lens)
17
+
18
+ MiMo ships its own jinja (macro-based, 3.9 KB) β€” NOT the Qwen
19
+ froggeric template (16 KB). `--jinja` works natively on all runs.
20
+
21
+ Sampling (Qwen family protocol): temp 1.0 / top_p 0.95 / top_k 20,
22
+ presence 0.0, max_tokens 2048. No Gemma leftovers (those lived only in
23
+ chain scripts). HE runner now logs battery time (start stamp + min/task).
24
+
25
+ ## Imatrix compass (MiMo vs Ox vs Neo, 248 common tensors)
26
+
27
+ | Pair | Pearson | Rank agree |
28
+ |:-----|--------:|-----------:|
29
+ | mimo–ox | 0.9989 | 0.9975 |
30
+ | mimo–neo | 0.9992 | 0.9977 |
31
+ | ox–neo | 0.9999 | 0.9999 |
32
+
33
+ Same law as FUSION: recipe dominates, finetune nearly invisible. But the
34
+ divergence clusters mid-layer FFN **15,16,17,19,13,27,14,18,23,11** β€” the
35
+ weight-compass "finetune soul" zone. Scale: MiMo cal runs 1.6Γ— hotter
36
+ (mean 1.9e5 vs 1.2e5) β€” quantize on its own imatrix, never mixed.
37
+ (Dual-max(Ox,MiMo) is a fiction: max always picks MiMo's hotter scale.
38
+ N1 dry-run with dual lens == pure-MiMo geometry bit for bit.)
39
+
40
+ ## Verdicts (slow ring: HumanEval pass@1, temp 1.0 / top_p 0.95 / top_k 20, presence 0.0, max 2048)
41
+
42
+ | Build | Size | PPL (ctx1024) | HE pass@1 | HE+ |
43
+ |:------|-----:|:-------------:|:---------:|:---:|
44
+ | MiMo-5100-SMAPE (own imatrix) | 5.0 GB | 8.7435 | 70.12% (115/164) | 66.5% |
45
+ | MiMo-4500-SMAPE (own imatrix, allow-q3) | 4.5 GB | 10.5343 | 29.88% (49/164) | 26.2 β€” lava zone (Neo-3650-Q2: 23.78%), sub-4 at ~4GB |
46
+ | MiMo-4500-SMSE (own imatrix, allow-q3) | 4.5 GB | **9.4092** | **54.88% (90/164)** | **53.0** β€” rescue +25pp over SMAPE, TOX-like save |
47
+ | MiMo-5100-SMSE (own imatrix) | 5.0 GB | 8.7389 | 71.34% (117/164) | 65.9 β€” 3 empties. Duel 0/3 columns (HE p=0.88, HE+ p=0.76 against, empties p=0.13): the +1.2pp is NOISE, crown demoted 2026-09-28. Decisiveness edge (3 vs 7) stands as observation |
48
+ | MiMo-6500-MSE (own imatrix) | 6.5 GB | 8.7655 | **75.00% (123/164)** | **68.9** β€” beats Q6_K by +3pp at βˆ’0.7 GB |
49
+ | MiMo-6500-SMSE (own imatrix) | 6.5 GB | 8.7275 | 72.56% (119/164) | 67.1 β€” LOSES to MSE by 2.4pp: PPL tie, verdict fail. Burn wins; refinement: SMSE had MORE Q8 (125 vs 108) yet lost β€” at big budgets the MIDDLE decides (MSE Q6 67 vs SMSE 41) |
50
+ | MiMo-Q6_K (stock) | 7.2 GB | 9.1362 | 71.95% (118/164) | 66.5% |
51
+ | Ox-SMAPE-5100 (ref) | 5.0 GB | 7.5670 | 88.41% (145/164) | β€” |
52
+ | Neo-SMAPE-5100 (ref) | 5.0 GB | 7.8012 | 82.93% (136/164) | β€” |
53
+
54
+ Fusions (N1/N1m) live in `RINIQ-NEXT.md`.
55
+
56
+ Geometry tells the rescue story (audit_tiers):
57
+
58
+ | Tier | 4500-SMAPE | 4500-SMSE | 5100-SMAPE | 5100-SMSE |
59
+ |:-----|----------:|----------:|----------:|----------:|
60
+ | sub-Q4 | **107** (104Γ—IQ2_XXS!) | **66** (59Γ—IQ2_XXS, rest IQ3/IQ4) | 0 | 0 |
61
+ | Q4 | 4 | 6 | 206 | 197 |
62
+ | Q5 | 2 | 27 | 11 | 29 |
63
+ | Q6 | 0 | 82 | 11 | 0 |
64
+ | Q8 penthouses | **118** | **0** | 0 | 0 |
65
+ | HE | 29.88% | 54.88% | 70.12% | 71.34% |
66
+
67
+ SMSE never affords a penthouse (Q8 = 0 at both budgets), halves the
68
+ dungeon, and rebuilds the middle (Q5+Q6 = 109 @4500). "Penthouses only
69
+ if affordable" β€” in numbers. Bonus catch: SMAPE-4500 evicted even the
70
+ free F16 residents (22β†’0), SMSE-4500 kept 8 β€” total war vs discipline.
71
+ Open risk @6500: MSE's 75% rides 108 Q8 penthouses; if SMSE arrives
72
+ with zero Q8 there and loses, the law extends to "penthouses required
73
+ at big budgets". Dry-run geometry at quant time will tell first.
74
+
75
+ Dry-run geometry @5100 (SMAPE, own lens): F16 177 / Q4 226 / Q5 13 /
76
+ Q6 11 / Q8 0 β€” all floor, zero penthouses (same as Ox-SMAPE-5100 shape).
77
+
78
+ Empties: MiMo 7 (Ox-like decisiveness, not Neo hesitation).
79
+ Speed: MiMo ~20 s/task (vs 40+ Gemma-12B, 60+ RINIQ duels) β€” decisive,
80
+ no thinking-chewing. Timer now logged per battery.
81
+ Battery times: N1 77.1 min (28.2 s/task); MiMo-5100 ~50 min (~18 s/task).
82
+ Fusion decisiveness (N1: 24 empties, seam friction) β†’ `RINIQ-NEXT.md`.
83
+
84
+ HE+ rescore runs after every battery (CPU, EvalPlus 80Γ—) β€” strictness
85
+ column. Layer swaps (15–17) may move other benchmarks (SWE, terminal,
86
+ vision) in either direction β€” those arenas are queued, not claimed.
87
+ If a build behaves weird on your hardware, open a discussion with setup
88
+ + task β€” every scar goes in the ledger.
89
+
90
+ ## Field notes (manual testing, 2026-09-22)
91
+
92
+ - Thinks a lot: long reasoning traces, lower t/s than OxCoder but FEELS
93
+ faster (decisive, no hesitation). Part of the gap was CPU contention
94
+ on the test box β€” clean comparison gives MiMo ~+1 tok/s over RINIQ/Ox.
95
+ - Agent-scaffold leak: in the pi agent framework, a bare "Privet" (hello
96
+ in Russian) makes it confabulate a task (portfolio brief) β€” operator
97
+ aborted manually upon noticing ("Operation aborted" was the operator,
98
+ not the model).
99
+ Other frameworks fine; RINIQ/Ox/Neo never do this. Read: MiMo's agent
100
+ training (TerminalBench/Toolathlon scaffolds) misfires on agent-style
101
+ system prompts β€” training artifact, harness-dependent, not quant damage.
102
+ - Temp sensitivity: temp 0.6 breaks tool-call format adherence (confuses
103
+ file-create vs bash); temp 1.0 works. Vendor temp 1.0 mandatory for
104
+ agentic tasks β€” lower temp commits to the wrong pattern confidently.
105
+ - `--reasoning-effort` WORKS (unique in family): low β†’ <5k thinking
106
+ tokens, xhigh β†’ ~15k. Controllable depth/speed dial out of the box.
107
+ Ox/Neo ignore these flags (thinking stripped).
108
+ - 3D-snake field test, first attempt BROKEN (critical: wall.push coords,
109
+ isHead args, food destructuring [] vs {}, drawObj syntax, undefined
110
+ dir0/tSince; logic: matrices rewritten, WebGL buffers, input).
111
+ OxCoder baseline: wrote 3D models but NEVER got it running.
112
+ Iterations-to-working TBD.
113
+ - low-effort FIRST TRY: WORKING 3D snake (renders, orbit camera, Russian
114
+ UI, score/game-over flow; not quite playable). Total 11k tokens vs
115
+ 17k for the broken xhigh attempt.
116
+ CONFOUND (honest): prompts differed by ONE word (originally in
117
+ Russian) β€” xhigh got "make an HTML 3D snake...", low got "make a
118
+ WORKING HTML 3D snake...".
119
+ The win may belong to the word, not the effort level. Unconfounded
120
+ A/B (same prompt, low vs xhigh) still open.
121
+ - opencode harness: 400-line file, noticed a bug immediately, fixing
122
+ it himself. Harness-independent thinking, second framework confirmed.
123
+ - RINIQ-N1 field test: thinking depth lands BETWEEN parents (Ox shallow
124
+ < N1 medium < MiMo deep) β€” mid-layer blocks 15–17 appear to govern
125
+ reasoning depth. 3D-snake attempts: 1st = 2D game, 2nd = semi-working
126
+ 3D, 3rd = working snake (ultra slow but works), each attempt <4k
127
+ tokens. N1m report pending.
128
+ - Test conditions: llama-server web UI, normal sampling (temp 1.0),
129
+ no looping observed. First runs only β€” slight-breakage chance noted.
130
+
131
+ ## Published-code duel (different harnesses β€” theater, same-harness HE decides)
132
+
133
+ | Bench | OxCoder-9B | MiMo-Distill |
134
+ |:------|-----------:|-------------:|
135
+ | SWE Verified | 73.5 | 61.1 (avg@3) |
136
+ | SWE Pro | 49.1 | 44.6 (avg@3) |
137
+ | TerminalBench 2.1 | 49.6–50.8 | 37.1 |
138
+ | GPQA Diamond | 86.9 | ??? (unreported) |
139
+
140
+ Base Qwen3.5-9B itself differs between the two tables (SWE-V 53.2 vs
141
+ 60.0) β€” harness strictness (OpenHands + anti-hacking, no net) vs avg@3
142
+ inflation. Cross-table comparison is void; only same-harness duels count.
143
+ House rule: HE is used for comparison, not for score β€” if a small hybrid
144
+ scores roughly like stock Q6 on the same harness, most capabilities likely
145
+ survived quantization. The verdict column certifies preservation, not rank.
146
+
147
+ Field temp law (Ox-SMSE-6500, 3D-snake first-try): temp 0.0 = working code
148
+ at once; 0.6/1.0 needed retries. Hypothesis: low temp for codegen,
149
+ high temp for agents/reasoning. (MiMo agentic tasks mandate temp 1.0 β€”
150
+ same split from the other side.)
151
+
152
+ SRIQ-Q6 (LoRA-child of MiMo, Chinese traces): HE 64.63% (106/164) at
153
+ temp 1.0 vs **78.66% (129/164) at temp 0.0** β€” greedy moves the VERDICT
154
+ +14pp here (2/3 columns significant), decisiveness flat (6 vs 8 empties).
155
+ CORRECTION 2026-09-28 (manifest duel falsified the old "identical score"
156
+ entry β€” I never counted the t0 file, scar kept). Below MiMo-Q6 (71.95%)
157
+ at temp 1.0; PPL skipped (wiki canary invalid for Chinese reasoning).
158
+ Own SRIQ imatrix done.
159
+ SRIQ-5100-SMSE (own imatrix): 63.41% (104/164), HE+ 56.7, **1 empty** β€”
160
+ most decisive build in the lab, but βˆ’1.2pp vs Q6: mirror image of MiMo
161
+ (+1.2pp). SMSE buys decisiveness everywhere, verdicts are family-local.
162
+
163
+ ## Scars (chain discipline)
164
+
165
+ - 2026-09-22: heavy CPU quant launched parallel to GPU HE battery β†’ RAM
166
+ pressure β†’ server death, 30-min battery progress DISCARDED. Second time
167
+ (same kill took Q6-HE at task 60). Rule (README:118, "GPU is a strict
168
+ queue") now enforced as: one heavy job at a time, sequential chains
169
+ only (chain5: PPLΓ—3 then HEΓ—3).
170
+ - earlyoom (root, main rig) unkillable without local sudo (root locked
171
+ after wrong attempts) β€” lives on, chain discipline is the shield.
172
+ - Server deaths mid-long-run (tasks 25/30/101/110) on 8 GB VRAM:
173
+ 6.4 GB weights + KV c4096 + fragmentation β†’ random OOM. Mitigations:
174
+ `--cache-reuse 256`, `-c 4096` (c8192 KV doesn't fit: create_context
175
+ fail, proven). c8192 works on 9B/5 GB builds (RINIQ era).
176
+
177
+ ## Next
178
+
179
+ - [x] Q6_K stock baseline: quant done, PPL+HE in chain5
180
+ - [ ] Fusions (N1/N1m), weight compass, LCB: see `RINIQ-NEXT.md`
181
+ - [ ] LiveCodeBench v6 (backlog): MiMo's arena (SWE/terminal), Ox's too