Amazing work!!
Hey Soulfate24,
Just found your AutoRound+ASHQ1 suite and wanted to reach out.
What you built here is really solid work. The dual-phase approach with AutoRound preprocessing solves a fundamental limitation I had with raw BF16 inputs. The way you structured the knapsack optimizer integration, handled tied-weight detection, and built out the complete pipeline shows you understood the core concept and took it further than I did.
The tier system makes sense, the documentation is thorough, and your observation about int4 lineage saturating at Q5_K is theoretically correct. The MTP extraction, mmproj handling, recurrent state protection - you clearly understand what you're working with.
One thing that would help demonstrate the advantage: could you add stock llama.cpp quants to the perplexity table for comparison? Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q4_K_S, Q3_K_M, Q3_K_S - matched by file size to your ASHQ1 tiers. I think seeing them side by side would show the real difference clearly.
I don't have access to my PC right now to test this myself, but I'll be home soon. If you're interested in collaborating once I'm back, I'd be up for working on something together.
Good work on this.
And One technical question: for the higher tiers
(Quality, possibly Compact), did you keep
allow_q3_or_lower disabled? I found that flag
is only beneficial on tight budgets like Nano/Mini,
but on larger budgets it can hurt perplexity since
the optimizer doesn't need Q3 to fit the target ratio.
Also wanted to expand on the allow_q3_or_lower point -
this affects even the low tiers, not just Quality/Compact.
The issue is the algorithm is just too greedy. Once that
flag is on, it will drop tensors all the way down to IQ2_XXS
if it decides they're low priority, when something like Q4_K
would have fit the budget just fine and preserved way more
signal. Anything below Q4 tends to hurt perplexity more than
it helps, even though it technically saves bytes.
The tiers still come out ahead of uniform IQ2 quants overall
since the important tensors stay protected, but there's
noticeable perplexity left on the table from being too
aggressive on the "unimportant" ones. A less greedy allocation
that stops around Q3/Q4 as a floor instead of IQ2 would
probably close that gap.
Hey wepiqx,
Thank you so much for this, coming from you it means a lot. ASHQ1 is literally the foundation everything here builds on; your priority-queue knapsack formulation and tied-group activation hashing are what make this whole suite possible, and they're credited as such in the README. The AutoRound pre-stage was indeed the piece that unlocked things your original couldn't do from raw BF16, but without your core there'd be nothing to integrate into.
On your technical question, great instincts, and I now have measurements to answer it precisely:
allow_q3_or_lowerstayed disabled everywhere in my standard tier runs. It remains opt-in in the CLI, but nothing ships through it.- Your "too greedy" diagnosis turned out to be exactly right, and measurable. I built a knockout-probe harness (
90_attribution-probe.py) to marginalize each decision, and forced IQ3_S on gate/up measured +0.022 KLD (~2Γ10β»β΄ damage per MiB saved), an order of magnitude worse per byte than any winning promotion I found. So v2.0.0 went further than a Q3/Q4 stopping rule: MLP classes (ffn_gate/up/down) now carry a hard IQ4_XS declarative floor at plain budgets. Your suggested fix, but stopped one notch earlier because the data said so. - Scaling to other families also surfaced two real bugs in the classifier path worth knowing about, untied lm_heads expose an
output.weightthe exact-name matcher shelved as unknown, and on one hybrid the recurrent readout (ssm_out) needed a dedicated precision floor below which output entropy collapses. Both are fixed and documented (CHARTER.md, laws L0βL5).
On the stock-quants comparison: absolutely yes, and it's nearly free with the tooling I already have, same imatrix, uniform llama-quantize pass per type, then straight into the existing perplexity/KLD sweep. Size-matched against NanoβFidelity it should show exactly where the selective allocation pays. I'll post the full table once the runs finish.
And yes, very interested in collaborating. Ping me when you're back at your machine; I'd love to get your eyes on the calibration ledger and see where you'd want to take it next.
.>python 03_perplexity_test.py
Binary : .\llama-cpp\llama-perplexity.exe
GPU : cuda Β· 7930/8188 MiB VRAM available
[Corpus] Using existing .\wiki.test.raw
[KL-Base] Generating reference logits over the full corpus (ngl=7, ~50 min)β¦
[34m0.04.519.350[0m [32mI [0mcmn init: llama threadpool init, n_threads = 15
[34m0.04.519.930[0m [32mI [0m
[34m0.04.519.968[0m [32mI [0msystem_info: n_threads = 15 (n_threads_batch = 15) / 16 | CUDA : ARCHS = 750,800,860,890,900,1200,1210 | USE_GRAPHS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
[34m0.04.521.028[0m [32mI [0mperplexity: saving all logits to .\kld-bf16.dat
[34m0.04.521.033[0m [32mI [0mperplexity: tokenizing the input ..
[34m0.04.864.087[0m [32mI [0mperplexity: tokenization took 343.048 ms
[34m0.04.864.439[0m [32mI [0mperplexity: calculating perplexity over 580 chunks, n_ctx=512, batch_size=512, n_seq=1
[34m0.09.981.668[0m [32mI [0mperplexity: 5.10 seconds per pass - ETA 49.25 minutes
[1]5.7284,β¦,[580]9.5298,
[34m25.24.692.262[0m [32mI [0mFinal estimate: PPL = 9.5298 +/- 0.06948
β Base compressed with Xpress8K
β Base ready: kld-bf16.dat
Available models:
1. model-no-mtp-BF16.gguf (17091 MiB Β· ngl=7)
2. model-ASHQ1-Fidelity-48pc.gguf (9243 MiB Β· ngl=22)
3. _stock-q8_0.gguf (9086 MiB Β· ngl=22)
4. _stock-q6_K.gguf (7018 MiB Β· ngl=32)
5. model-ASHQ1-Quality-39pc.gguf (6675 MiB Β· ngl=32)
6. _stock-q5_k_m.gguf (6168 MiB Β· ngl=32)
7. model-ASHQ1-Compact-33pc.gguf (5650 MiB Β· ngl=32)
8. _stock-q4_k_m.gguf (5368 MiB Β· ngl=32)
9. model-ASHQ1-Mini-30pc.gguf (5225 MiB Β· ngl=32)
10. _stock-q4_k_s.gguf (5104 MiB Β· ngl=32)
11. model-ASHQ1-Nano-27pc.gguf (4625 MiB Β· ngl=32)
12. _stock-q3_k_m.gguf (4409 MiB Β· ngl=32)
13. _stock-q3_k_s.gguf (4062 MiB Β· ngl=32)
a. ALL Β· q. quit
Select model(s) or file path [1/a/model.gguf/q]: a
layout: 32 layers Β· 232 MiB/layer Β· non-layer 1826 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ model-ASHQ1-Fidelity-48pc.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.5381 Β· KLD=0.0084 Β· RMS Ξp=2.46% Β· top-p=97.4% Β· 345.9 tok/s Β· 11.2 min
layout: 32 layers Β· 220 MiB/layer Β· non-layer 2061 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q8_0.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.5448 Β· KLD=0.0075 Β· RMS Ξp=2.45% Β· top-p=97.7% Β· 365.7 tok/s Β· 11.3 min
layout: 32 layers Β· 170 MiB/layer Β· non-layer 1591 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q6_K.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.4530 Β· KLD=0.0145 Β· RMS Ξp=3.23% Β· top-p=96.2% Β· 731.4 tok/s Β· 6.7 min
layout: 32 layers Β· 159 MiB/layer Β· non-layer 1591 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ model-ASHQ1-Quality-39pc.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.2284 Β· KLD=0.0303 Β· RMS Ξp=4.55% Β· top-p=94.1% Β· 775.8 tok/s Β· 6.3 min
layout: 32 layers Β· 147 MiB/layer Β· non-layer 1463 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q5_k_m.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.0280 Β· KLD=0.0755 Β· RMS Ξp=6.76% Β· top-p=90.5% Β· 1113.0 tok/s Β· 5.2 min
layout: 32 layers Β· 127 MiB/layer Β· non-layer 1591 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ model-ASHQ1-Compact-33pc.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.6526 Β· KLD=0.0505 Β· RMS Ξp=5.83% Β· top-p=91.7% Β· 1066.7 tok/s Β· 4.9 min
layout: 32 layers Β· 126 MiB/layer Β· non-layer 1341 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q4_k_m.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.1993 Β· KLD=0.0896 Β· RMS Ξp=7.55% Β· top-p=89.1% Β· 1219.0 tok/s Β· 4.8 min
layout: 32 layers Β· 114 MiB/layer Β· non-layer 1591 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ model-ASHQ1-Mini-30pc.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.8564 Β· KLD=0.0649 Β· RMS Ξp=6.67% Β· top-p=90.3% Β· 1219.0 tok/s Β· 4.9 min
layout: 32 layers Β· 118 MiB/layer Β· non-layer 1341 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q4_k_s.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=9.2654 Β· KLD=0.0927 Β· RMS Ξp=7.68% Β· top-p=88.9% Β· 1383.8 tok/s Β· 4.7 min
layout: 32 layers Β· 111 MiB/layer Β· non-layer 1061 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ model-ASHQ1-Nano-27pc.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=10.0184 Β· KLD=0.0856 Β· RMS Ξp=7.70% Β· top-p=88.1% Β· 1312.8 tok/s Β· 4.8 min
layout: 32 layers Β· 100 MiB/layer Β· non-layer 1213 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q3_k_m.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=10.2358 Β· KLD=0.1543 Β· RMS Ξp=10.30% Β· top-p=84.5% Β· 1219.0 tok/s Β· 4.7 min
layout: 32 layers Β· 89 MiB/layer Β· non-layer 1213 MiB Β· KV 4Γ256
probe: native auto-fit β¦ ok β llama.cpp sizes the split itself
ββ _stock-q3_k_s.gguf ββ
auto-fit Β· ctx=512 Β· batch=512 Β· threads=15 Β· Flash-Attention Β· KL-divergence
PPL=10.4404 Β· KLD=0.2372 Β· RMS Ξp=12.58% Β· top-p=80.3% Β· 1280.0 tok/s Β· 5.5 min
========================================================================================
EVALUATION SUMMARY (KLD/RMS = fidelity to BF16 reference; PPL = entropy indicator)
========================================================================================
Model PPL +ΞPPL KLD RMS Ξp top-p Speed
-------------------------------------------------------------------------------------------
model-ASHQ1-Fidelity-48pc.gguf 9.5381 +0.0000 0.0084 2.46% 97.4% 345.9 t/s
_stock-q8_0.gguf 9.5448 +0.0067 0.0075 2.45% 97.7% 365.7 t/s
_stock-q6_K.gguf 9.4530 -0.0851 0.0145 3.23% 96.2% 731.4 t/s
model-ASHQ1-Quality-39pc.gguf 9.2284 -0.3098 0.0303 4.55% 94.1% 775.8 t/s
_stock-q5_k_m.gguf 9.0280 -0.5102 0.0755 6.76% 90.5% 1113.0 t/s
model-ASHQ1-Compact-33pc.gguf 9.6526 +0.1144 0.0505 5.83% 91.7% 1066.7 t/s
_stock-q4_k_m.gguf 9.1993 -0.3388 0.0896 7.55% 89.1% 1219.0 t/s
model-ASHQ1-Mini-30pc.gguf 9.8564 +0.3183 0.0649 6.67% 90.3% 1219.0 t/s
_stock-q4_k_s.gguf 9.2654 -0.2727 0.0927 7.68% 88.9% 1383.8 t/s
model-ASHQ1-Nano-27pc.gguf 10.0184 +0.4802 0.0856 7.70% 88.1% 1312.8 t/s
_stock-q3_k_m.gguf 10.2358 +0.6977 0.1543 10.30% 84.5% 1219.0 t/s
_stock-q3_k_s.gguf 10.4404 +0.9023 0.2372 12.58% 80.3% 1280.0 t/s
Ξ PPL relative to model-ASHQ1-Fidelity-48pc.gguf Β· corpus: wiki.test.raw
Reading grid: KLD < 0.02 quasi-lossless Β· 0.02-0.10 solid Β· 0.10-0.20 usable Β· >0.20 degraded.
PPL below the BF16 base = entropy collapse (over-confidence), not a quality gain.
wiki.test.raw corpus: absolute PPL comparable to llama.cpp published benchmarks.
Update: your stock-quant request turned into the most valuable experiment of the release. Full size-matched sweep on Ornith-9B: ASHQ1 wins decisively at 27β33% (Nano vs Q3_K_M: KLD β45%, top-p +3.6; Mini vs Q4_K_S: β30%; Compact vs Q4_K_M: β44%), but the edge decays with ratio: Quality at 39% is only ~22% ahead of size-interpolated stock, and Fidelity ties/loses to straight Q8_0. That's a clean L-law now: allocation alpha β distance-from-uniform-budget. It reshaped our defaults (batch = Mini/Compact/Nano; higher tiers opt-in) and raises a fun open question you might enjoy: even stock quants show entropy collapse on this distillate, worth investigating together whether it's the distillation (fine-tune) or the arch.
Hey! Nice work on the Remix suite, woah so much benchmarks i feel bad for your graphics card hehe andd
Small strange thing though β I just ran my updated ASHQ1 on Ornith-1.5-9B and got PPL 8.6341 @ 6511 MiB, which is quite a bit lower than your Quality-36pc (9.3692 @ 6330 MiB). And that's with plain imatrix stuff (bartowski's), no AutoRound on my side β so I'd expect you to win, not me
Quick question β what context did you run llama-perplexity with? I'm on -c 1024, wiki.test.raw, full offload. Wondering if that's the whole difference.
One more note: I used the MTP version of the model, but that shouldn't explain it β if anything it makes my numbers slightly worse, since the MTP heads eat part of the budget at Q8_0.
I'm back at my machine and already pushing a new ASHQ1 update btw β reworked a bunch of the code, testing on 1.5 right now. Happy to compare notes / swap configs if you're interested!
and do you have discord?
P.S. β one more fun data point from today's testing. I also tried a "top-down" mode: instead of starting everything at low precision and upgrading, every quantizable tensor starts at F16 β norms included β and gets greedily downgraded, cheapest quality-loss-per-MB first, until the budget fits. So all 177 norm tensors ended up at Q4_K instead of F16.
Result: PPL 8.6337 vs 8.6341 for the regular mode. Technically a hair lower, but let's be honest β Ξ=0.0004 against Ο=0.06 is identical. Which is exactly the surprise: quantizing every single norm to Q4_K cost nothing. Two completely different allocation paths converged to the same PPL. Your attribution probes might find that interesting β apparently norms on this architecture are free real estate
Hey! Again!
Since my last message I've been running a whole experimental series on my side and I think we converged on the same thing from opposite directions β your L-law and what I'm seeing rhyme hard. Dumping everything, maybe useful for your tables:
- The barbell only wins inside a bubble. I built a playground where the queue's quality metric is swappable (MSE/RMSE/SMAPE/SSIM/perceptual/huber...) and ran it across 5 models. SMAPE-barbell (junkβQ4, preciousβQ8) wins big on Spark-1.7B (71.50β68.71) and Spark-4B (31.25β30.64) β and loses everywhere else: Ornith-9B (+0.15), Ling-8B-MoE (+0.07), and MiniCPM5-2B (+0.13, and that's a small dense model, so much for the size hypothesis). Both Sparks share my Qwen3.8 calibration set + imatrix recipe, everything else doesn't. Prime suspect is calibration, not architecture β I'm planning a calibration-swap (same model, different imatrix domain) to prove it. What's in your Bartowski v5 mix, btw? If our calibrations differ systematically, that might fully explain the bubble.
- Measured damage concentrates; theory doesn't. I ran full sensitivity sweeps (one tied group dropped at a time, ΞPPL measured): on 1.7B the damage-Gini is ~0.55, top-10% groups carry ~40% of total damage. My read: hybrid SWA models concentrate damage in a few full-attention singletons (barbell food), uniform dense/MoE smear it (smooth MSE food). That's my version of your L-law β "diagnose concentration first, then pick the utility", instead of one metric for all. Currently running the same sweep on MiniCPM5-2B overnight to check: diffuse damage there would confirm it at the measurement level.
- Relief-pins (my current champion). The sweeps found groups where dropping precision improves PPL (up to β7.8 on Q5βQ3?!). Pinning those at their sweet spot and spending the savings elsewhere: 71.50 β 66.70. Monotonicity is officially dead. Happy to share all damage labels (112 Q5βQ4 + 112 Q5βQ3 + MiniCPM incoming) if your attribution probes want a second dataset.
- Negative result, honestly: I tried learning the damage with a tiny MLP (~150 params, honest size for ~100 labels). Best heldout Spearman 0.309 (vs timp β0.10!) β first positive signal ever β but steering the queue with it lost the duel (70.94). Zero-shot onto 9B detonated (β99s, feature scale bug, since fixed with scale-invariant features). So: learning works a little, measurement works a lot. Your knockout-probes vs my sweep labels would make a great joint dataset, by the way.
- Taking your KLD point. You're right that PPL-only ranking is suspect (my whole scoreboard is PPL, guilty). Adopting the KLD-vs-BF16 protocol on my side β curious whether some of my SMAPE "losses" are actually collapse artifacts.
One question back: on Spark-4B your Quality-36pc row is bit-identical to stock Q5_K_M β allocator collapses to flat there? I see the same "give up and go flat" behavior in my queue near uniform budgets. Might be the same phenomenon from two sides.
Cheers, and RIP your GPU hehe
Hey! Great series, and yes β I think we did converge on the same thing from opposite directions. Taking everything in order:
Calibration, full answer: the default set is lemon07r/bartowski-imatrix-v5-semantic.txt β a pre-built semantic prompt corpus (Bartowski's v5 prompts, curated by lemon07r), pulled automatically when no local dataset is given: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic. So that one isn't my own mix. My own mix is the experimental composite the pipeline can build from scratch: a 5 MB corpus packed under category quotas β Agentic 48% (Claude-code traces, minimax-m3 traces, hermes-function-calling, multi-hop websearch tool calls), Logic 18% (MA-ProofBench Lean proofs, python-codes-25k), Diversity 13% (tristandruyen's v5_rc gist), Multilingual 12% (Aya, language-balanced), Frontier 9% (opus-4.8 traces, kept untrimmed as a priority source). The draft modules get their own corpus (Math 35 / Code 35 / Instruction 30) since they only need to propose, not know. The imatrix pass itself runs llama-imatrix at ctx 512, --process-output, full GPU offload on hybrids (a Q8_0 intermediate if VRAM is short, so fused GDN never splits across devices), MTP head structurally stripped before calibration, and a PPL split-integrity check against a CPU reference (Γ1.15 envelope) so a corrupted hybrid offload never feeds the allocator. So yes β if your bubble is calibration-driven, my guess is your Qwen3.8 set and my Bartowski/semantic mix would swap the result. Happy to run your calibration set through my rig on Spark-4B and send you back both KLD columns; and the whole pipeline is up in my repo (all scripts, comments included) if you want to diff configs directly.
On the damage-Gini: ~0.55 with top-10% of groups carrying ~40% of damage rhymes perfectly with what I measure. My band levers exist because of exactly that concentration: on a shortconv-mixer hybrid, protecting just the scarce full-attention band recovered KLD 0.371 β 0.213, while shielding the entire unclassified mixer stack at a higher tier recovered only 0.030 β taxonomy hygiene doesn't buy bandwidth. "Diagnose concentration first, then pick the utility" is precisely how I'd phrase my own operating rule. And your overnight MiniCPM5-2B prediction already has a head start from my side: in my sub-Nano compendium, the NanoβPico KLD slope on dense untied trunks is +179% at 1B and +92% at 3B (dense-tied), versus +16.5β44% on the GDN hybrids β diffuse damage on dense models is the pattern I'd bet on you finding too.
On the relief-pins: those are the numbers I'm most suspicious of, in the friendly way. I have an archived case where the worst arm by KLD printed the lowest PPL of the whole trio β nine points under the BF16 base β because quantization noise sharpened the distribution (over-confidence), and probability compression reads as a PPL win while the model is actually degrading. A Q5βQ3 step "improving" PPL by β7.8 has exactly that signature. Before declaring monotonicity dead, I'd love to see one of those pinned groups scored on KLD-vs-BF16 and top-p agreement: if the relief survives all three columns, it's a genuinely new law and I'll take your labels gratefully; if it only shows in PPL, it's the sharpening artifact. Either way β yes, send the 112+112 (+ MiniCPM when it lands). A joint dataset with my knockout CSVs (ΞKLD/MiB per tensor class, per family, noise floors inline) sounds great: two measurement systems, one ledger.
On the tiny MLP: your one-liner is my favorite sentence of the message β "learning works a little, measurement works a lot." That's basically my R8: stacked levers deliver ~91% of their solo-sum gains, so every composite arm gets re-measured before shipping. Steering losing the duel at ~100 labels is the expected outcome, not a failure. And your zero-shot detonation on 9B is why my probe harness now refuses rules files whose family= tag mismatches the target architecture β cross-family drift is real.
Two gotchas for your KLD adoption, from scars: (1) make the reference-logit generation flash-attention-symmetric with the eval pass, or you bake a systematic floor bias into every score β I had that bug, and fixing it also halved reference generation time; (2) keep PPL around as a canary (nominal PPL under the BF16 base = probability compression flag), never as the rank column. And rank models within one context only β absolute KLD is regime-dependent, my long-ctx pairing (2048Γ16, same 32,768-token span as 512Γ64) keeps comparisons honest when you want the extended-context view.
On Quality-36 β‘ stock Q5_K_M on Spark-4B: yes β by design, not collapse. My knapsack ran that duel and lost it three times (0.0719 flat vs 0.1543 allocated on spark2_5, plus lfm2 and ornith-9B), so since 2.3.3 Quality-36 routes structurally to flat Q5_K_M + imatrix, with the allocator kept behind an env hatch for on-demand measurement. Same phenomenon as your queue going flat near uniform budgets, seen from the other side.
On the 6511 MiB build: I'd genuinely like to score it, and I will. I'll run it through my battery tomorrow: ctx 512 Γ 64 chunks, Flash-Attention, KLD against the BF16 reference. Full transparency on what I expect to be interesting: your PPL 8.6341 lands almost exactly on my stock Q5_K_M-imx row for Ornith, which prints 8.6345 β under the BF16 base of 9.0824, carrying my over-confidence flag. So either your build is genuinely better than my whole current ladder there, or it's sitting in the sharpening zone my PPL ranking retired. That's exactly the kind of question KLD answers cheaply, and either verdict is a win for the ledger. One protocol note for a fair read: my build does carry the MTP head β I strip it only for calibration, generating the imatrix on the no-mtp BF16 base (my conversion step emits "model" and "model-no-mtp" side by side), then quantize and bench the full MTP build. So we're comparing like-for-like on that front; the only asymmetry is that your Q8_0 MTP heads eat classifier budget while mine sit in their own module lane outside the knapsack β I'll note the effective trunk budget when reading the duel. And the 9.3692 row you benchmarked against was a 2.3.1-era number; the current 2.4.2 rebuild of Ornith Quality-36 sits at 8.9895 @ 6388 MiB under the unified protocol, so the gap is smaller than it looked.
On Discord: I'm honestly not on any socials β HF discussions work great for me, and everything's in the repo if you want to dig asynchronously. Ping me here anytime.
Send the labels whenever ready β and good luck with the overnight MiniCPM run. RIP both our GPUs hehe
Duel delivered! I ran both your Ornith builds through the full battery β same protocol as my tables: wiki.test.raw, 64 chunks, span 32,768 tokens, Flash-Attention, full offload, KLD against my BF16 reference (running WITH MTP). Here's the readout:
ASHQ1-6500: PPL 8.6541 (Ο 0.067) Β· KLD 0.0763 (Ο 7e-04) Β· RMS Ξp 6.96% Β· top-p 90.3%
First, the number is real. Your 8.6341 at ctx 1024 and my 8.6541 at ctx 512 sit 0.02 apart with Ο ~0.06 β same value, two contexts, clean consistency.
Second, the interesting part. On my grid, a PPL that lands under my BF16 base (9.0824) carries an over-confidence flag β that's exactly where my stock Q5_K_M-imx row sits (8.6345, flagged). Your build lands in the same suspicious zone. So I ran the KLD to settle it, and it passes: 0.0763 is solidly in my 0.02β0.10 "solid" band, top-p 90.3%, chunk-stability clean (KLD Ο 7e-04). No collapse signature β the entropy drop is fidelity, not sharpening. Which means you win this round outright: against my current 2.4.2 Quality-36 (8.9895 @ 6388 MiB), your build takes β0.335 PPL for +123 MiB, and it clears the KLD check my stock rows couldn't. Fair is fair β that's a genuinely better Ornith build than my whole current ladder, and the ledger says so. Would love to see your config diff when you get a chance.
ASHQ1-4500-TOX: PPL 10.2194 (Ο 0.078) Β· KLD 0.2894 (Ο 7e-04) Β· RMS Ξp 14.38% Β· top-p 77.3%
And here's the cautionary tale you didn't ask for, but which lands right on the point you just adopted. Your README credits TOX with β0.145 PPL over the standard 4500 (10.1077 β 9.9629) by allowing sub-4 allocation on 79 tensors. On my span it prints 10.2194 PPL but KLD 0.2894 β deep in my degraded band (>0.20), with RMS Ξp nearly doubled (14.38%) and top-p down to 77.3%. In other words: the sub-4 damage is almost invisible in PPL and screams in KLD. That's the first head-to-head evidence on my side for exactly the "PPL-only ranking is suspect" conclusion β the metric that rewarded TOX's sub-4 tensors is the metric that can't see what they broke. If you push your standard (non-TOX) 4500 to the repo, I'll run the A/B so we have the pair scored under KLD, not just PPL.
Reading-grid reminder for your own runs: KLD < 0.02 quasi-lossless Β· 0.02β0.10 solid Β· 0.10β0.20 usable Β· >0.20 degraded β and absolute KLD is regime-dependent, so bind comparisons to one ctx and one span (mine: 32,768 tokens).
Meanwhile, still holding up my end: send your Qwen3.8 calibration set over and I'll run the swap on Spark-4B to chase your bubble β calibration vs architecture, one variable at a time.
β οΈ Verdict retracted β see my reply below. The ranking here was taken on the PPL column (the canary): 8.6541 sits 0.40 under the BF16 base. On KLD, the rank column, my Quality-36 holds (0.0241 vs 0.0763 at β123 MiB). The measurements stand; the concession doesn't.
Hey! Quick answers + shipments:
Duel accepted, both verdicts. The 6500 KLD-clean win feels great; the TOX tale stings as it should. My README now carries the correction (TOX = cautionary experiment with your KLD attached, standing rule: no PPL-only win ships without KLD). Standard 4500 uploading right now to a new repo wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF (6500 + standard 4500 + card with your duel numbers) β link as soon as the 11 GB land.
F32 note: you're right that it can't come from BF16 β upcasting adds zero information, so R9-F32 is dead for anything BF16-sourced like ours. Keeping it in mind only for future higher-precision sources. The Q8/Q6 gate-protection half of R9 is still testable cheap, might try it on Spark.
Norms: will rerun the top-down duel with F16 norms forced and report the delta β tells us whether your charter floor costs anything on GDN hybrids or is free real estate as I measured.
Labels + calibration shipped: wepiqx/ashq1-damage-labels β spark17 Q4/Q3 drops (112+112, relative damage) + Qwen3.8-27B-calibration-v6.txt (plain text, 12k lines). Swap away β calibration vs architecture, one variable at a time.
MiniCPM teacher resumed (63/168, Q5βQ3, tonight). Your dense-untied slope bet vs my Gini test β we'll see hehe.
KLD adoption: starting with the 1.7B reference-logit run. Gotchas copied verbatim (flash-symmetric ref, PPL-canary, single-ctx ranking). And yes β "either verdict is a win for the ledger" is now house policy on my side too.
One thought of mine on R9-F32: upcasting BF16βF32 adds zero information, so the F32 half of R9 looks dead for anything BF16-sourced like ours β unless I'm missing something? The Q8/Q6 gate-protection half still stands as a cheap test.
Shipments: labels + calibration at wepiqx/ashq1-damage-labels (dataset, 226 rows + calibration txt); Ornith duel builds landing at wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF (6500 up, standard 4500 minutes away).
Shame there's no faster channel than HF discussions, but they'll do (
You called it β with receipts, and harder than either of us expected. KLD battery on my 1.7B stand (stock flags, my ctx1024 span, F16 ref): Q8 0.0139 vs MSE/SMAPE/FRAG/RELIEF @1000 ALL at 0.345β0.351, identical in noise, while PPL spans 66.70β71.50. Not just relief β my entire zoo scoreboard was sharpness theater. Retracted the relief championship publicly, KLD is the rank column from now on, PPL demoted to canary everywhere. The TOX tale was the first bell and I under-read it.
Two consequences I'm chewing on: (1) damage-Gini survives as a distributional fact, but its motivation needs re-proof under KLD-damage sweeps β Gini-of-ΞPPL may not pick the right utility for fidelity; (2) my PPL-damage teacher labels are suspect as fidelity targets, so the tiny-MLP line pauses until KLD-damage labels exist. MiniCPM sweep still running β distributions are distributions.
Question back: have you ever seen two builds with equal KLD but clearly different generation quality? Wondering how much headroom the metric itself has β whether KLDβequal ever hides anything you caught by reading outputs.
i think we need to move on benchmark because PPL,KDL is not that good as we think..
like ppl on ctx 512 or ctx 1024 didn't measure the ability of model to be good on long contex and KDL is just measure of how the model far from original like BF16 or f16 tbh kinda sad XD because benchmarks i now running NeoHorse9b on Human eval the orig scores 98,1% or something like that interesting how much will be a difference of q6 quant my ashq1 6500 and q2 quant when i finish i gotta tell you
and why are you measuaring wikitext2 on 64 chunks i mean its faster but its didnt get the full picture
Hey! Your retraction first, then your questions, then new data that corrects one of my own earlier claims, which feels only fair given the week you've had ;)
On the retraction: that took more intellectual honesty than most of this field ever shows publicly. For what it's worth, my zoo retraction was less brutal only because my allocation paths happened to converge: I never had a relief champion to disown. But your 1.7B sweep is now my favorite datapoint of the whole exchange: Q8 at KLD 0.0139 while your entire scoreboard packs into 0.345β0.351, indistinguishable in noise, while PPL told a 5-point story. That's not "PPL is noisy", that's PPL actively misleading β and my own scale-bound law explains the flat KLD side too: 0.345β0.351 sits inside the 0.33β0.40 entropy-dissolution band I've measured converging from two separate collapse causes, so the quants past that line are equally dissolved and ranking stops meaning anything among them. The flat KLD is the ceiling signature, not KLD blindness β which is the strongest version of the point, and a reason to aim the KLD re-run below that line so the new labels don't inherit the saturation. And I agree with both consequences: Gini survives as a distributional fact but needs re-proving under KLD-damage sweeps, and the MLP line should stay paused until KLD-damage labels exist. My knockout harness computes per-decision ΞKLD against a fixed reference, so when your sweep re-runs under KLD, the joint dataset can carry both label types and the MLP un-pauses with meaningful targets.
R9-F32: you're not missing anything, upcasting adds zero information, that half is dead for BF16 sources, marked as such on my side (survives only as a rule for hypothetical higher-precision sources). The Q8/Q6 gate-protection half is the testable residue, and Spark is the right host for it. One prediction from my ledger: if it helps anywhere, it's the hybrid SWA singletons first, not the uniform dense tails.
Now the corrections. I ran your full trio β 6500, standard 4500, 4500-TOX β through my battery (wiki.test.raw, 64 chunks, span 32,768, Flash-Attention, full offload, KLD vs my BF16 reference):
| Build | PPL | KLD | RMS Ξp | top-p |
|---|---|---|---|---|
| ASHQ1-6500 | 8.6541 | 0.0763 | 6.96% | 90.3% |
| ASHQ1-4500 | 10.5057 | 0.3256 | 15.24% | 75.8% |
| ASHQ1-4500-TOX | 10.2194 | 0.2894 | 14.38% | 77.3% |
Correction #1 β my 6500 verdict from the duel post above: retracted. The measurements stand, the ranking was mine to fix, and it failed in the exact way we spent this week mapping: I ranked your build on the PPL column. The β0.335 PPL edge is real on the printout, but 8.6541 lands 0.40 under the BF16 base (9.0582), the same compression zone my stock Q5_K_M-imx occupies at 8.6345 (β0.42), the family-wide signature at this ratio, and exactly the β flag I put on my own rows. The tell was in my own sentences: I'd flagged my stock twin's under-base PPL as the suspicious zone, then waived the same flag for your build and called the entropy drop fidelity: same zone, two treatments, and that inconsistency should have been the giveaway. On the rank column my Quality-36 holds: KLD 0.0241 Β· top-p 94.5% Β· RMS 4.22% @ 6388 MiB vs 0.0763 Β· 90.3% Β· 6.96% @ 6511 MiB β 3.2Γ the KLD, +4.2pp top-p, β2.7pp RMS, at 123 MiB smaller. And "it clears the KLD check my stock rows couldn't" was flatly false: Q5_K_M-imx prints 0.0665, better than 0.0763, at 176 MiB smaller β no stock row on Ornith fails a KLD check, the worst (IQ3_M-imx) sits at 0.1421, still usable band. Mechanically, I used KLD as a pass/fail gate on your build instead of as the comparison column against mine. Honest placement of the 6500: between stock Q5_K_M-imx and my Nano-27 β Nano-class fidelity at Quality-class bytes, solid band, clean chunk-stability, no collapse signature. A genuinely fine quant, just not above the ladder on the column that ranks. One structural footnote for the ledger, if I read your recipe right: your MTP head rides Q8_0 where my flat recipe floors it at Q5_K, so at file parity your trunk ran a touch leaner β small, not a verdict-changer at a 3.2Γ gap, but if you want the trunk-parity arm (head at Q5_K, those bytes back in the trunk), I'll run it through the battery when I'm back and we'll have the clean pair. Scope: the verdict is retracted, the measurements are not β your numbers and the chunk-stability read stand. So the crown stays home, and the PPL disease got the guy who named it: a charter line for both of us. (Housekeeping while you're in your card, your call as always: the verdict line ranks on PPL β under your own new standing rule it would read "Nano-class fidelity at Quality-class size".)
Correction #2 β the TOX. It beats your standard 4500 on both metrics. So I'm formally retracting the sharp end of my earlier cautionary tale: on this span, the sub-4 allocation bought real fidelity, not sharpening, your β0.145 PPL claim stands under KLD too β and then some on my span: the pair prices at β0.286 PPL / β0.036 KLD here β it just lives in the degraded band while doing so. Both 4518 MiB builds sit above 0.20 (neither is deployment-grade at this budget), but at the ranking level, TOX is the better build and the ledger takes the correction happily. PPL and KLD agreeing here is also worth noting: when the two metrics diverge, KLD wins the argument; when they agree, we can actually trust the pair.
One protocol bonus from the same session: I ran the whole battery twice, once against my no-mtp BF16 reference and once against the with-MTP one: identical numbers on all three builds, to the last digit. The MTP heads are decode-time-only and silent in perplexity runs, so both references are interchangeable for this battery. Which also retroactively validates calibrating the imatrix on the no-mtp base, good call on your side, now with receipts.
Your KLD-ceiling question: my honest answer is structural, not empirical. Chunk-KLD measures local next-token distribution fidelity over a fixed span. It cannot see: long-context degradation (it's all 512-token chunks), instruction-following drift, tool-call format integrity, or rare-mode behavior. I haven't yet caught an equal-KLD-different-generation case with receipts in my own runs, which is exactly why my grid is labeled triage, not certificate. It's the best cheap rank column we have for allocator decisions; it was never going to certify the product.
On "PPL and KLD are not that good as we think": you're right, and here's where I land. Two different jobs, two rings. The KLD battery stays the allocator feedback loop: per-tensor damage localization for near-free, nothing replaces it there, and it just earned its keep twice in this exchange. Then a slow capability ring, run only on finals: same builds, fixed seeds, scored once, not per allocation decision. On your HumanEval plan (right instinct), one caution: a base at 98.1% is saturated, so 97.2 vs 96.8 is floor-effect noise, the mirror image of the PPL disease. Pick 2β3 benchmarks where the base lands mid-range (60β85%) so degradation has room to show. And since your ctx-1024 point stands (neither PPL nor chunk-KLD measures long-context ability), a needle-in-haystack at 32k+ ctx on the same builds would answer a question our current tools structurally can't. Full disclosure: I'm no more experienced with capability benchmarks than you are, my use case is a personal agentic harness on an 8 GB card, so my ring will mirror that: agentic/tool-call integrity + long-context retrieval, small and fixed. Send your NeoHorse numbers anyway; even saturated, a Q2 collapse would be signal.
On the 64 chunks: it's not a cheaper PPL, it's the controlled test bed KLD needs: cross-run ΞKLD only means anything if every build scores the same 32,768 tokens. Full-corpus PPL is a separate column in my harness (comparable to llama.cpp published numbers); the 64-chunk run exists for span-binding, with the chunk-stability canaries riding along. Same reason you bind your sweeps to one ctx and seed.
Status: your 226 labels and the Qwen3.8 v6 calibration are in; the Spark-4B calibration-swap (bubble test) is my next run. Curious about your top-down-with-F16-norms rerun: if Q4_K norms really cost nothing on GDN hybrids, my charter floor takes its first real challenge and I'll log it honestly either way. And on your MiniCPM dense-untied bet: my slope prediction is on record.
Quick heads-up on my side: real life is reclaiming a few days or weeks, so expect my next replies to land slower, multi-day latency instead of hours. Nothing stalled on purpose: your 226 labels and the Qwen3.8 v6 calibration are safely downloaded, the Spark-4B swap and the norms-F16 probe are queued as my next runs, and the correction on my TOX read is already in the ledger (both our retrospectives match: sub-4 targeted by ranking can buy real fidelity, just not above 0.20). Post your NeoHorse numbers whenever they're ready; I'll pick everything up in one batch when I'm back. HF discussions' latency now officially works both ways! That's why I'm not on any social media, where everything has to be instant.
Your corrections are taken like medicine β unpleasant, effective:
6500 verdict: accepted, card fixed. Our Ornith README now reads "Nano-class fidelity at Quality-class size" with your Quality-36 crowned (0.0241 vs 0.0763) and both columns on every row. The old "beats the entire ladder" line is dead. And you're right about the mechanism of my error: I used KLD as a gate instead of a column. Charter line earned on both sides.
TOX pair: taken with interest. 10.5057/0.3256 vs 10.2194/0.2894 β TOX wins both, my "cautionary tale" softened to what it is: better build, degraded band. Our main README carries the same correction. The agreed rule (diverge β KLD wins; agree β trust the pair) is now house policy.
Trunk-parity arm: yes, let's. MTP head at Q5_K, bytes back to trunk, same 6511 MiB. Send me the exact recipe (or I dump my --show-config and you mirror the head rule) β clean pair through your battery settles whether the head lane mattered at all.
Norms-F16 probe: queued on my side (force norms F16 in top-down, delta vs Q4-norms). If it costs nothing, your charter floor gets its first dent; if it costs, my "free real estate" gets retracted. Either way logged.
MiniCPM teacher: resumed (127/168 + forge tail running). Gini answer lands with it.
NeoHorse HE battery: running now (Q6K 82.32% pass@1 vs 98.17 official β real capability gap, not noise; ASHQ-6500 and Q2-K in queue). Numbers when they land, win or lose.
On saturated bases: taken β 98.1% can't resolve 97 vs 96. Our duel still answers the collapse question (does Q2K fall off the cliff?), just not fine ranking. Mid-range benchmarks join the slow ring next.
KLD-ceiling: your structural answer is the honest one (local next-token fidelity can't see long-ctx/instruction/tool drift). Triage-not-certificate goes on our protocol card verbatim. The needle-in-haystack idea is parked until bigger VRAM.
Rest well β HF latency works both ways indeed. The ledger grows while you sleep: my zoo keeps a scar column for every retraction, it's getting long hehe.
The slow ring speaks: NeoHorse HE triple.
Build Size pass@1
ASHQ-6500 (ours) 6.83 GB 85.98% (141/164)
Q6_K (stock) 7.36 GB 82.32% (135/164)
Q2_K (stock) 3.83 GB 0.00% β stub completions, genuine collapse
Vendor BF16 posts 98.17% (at undisclosed ctx β our runs are c8192, so the gap is approximate). The battery resolves collapse even near saturation: Q2K at literal zero next to 86/82. The +3.7pp over stock Q6_K at β0.5 GB is outside floor noise β imatrix-kept code paths beat flat allocation on capability.
Two housekeeping notes from the run: your saturation caution is now protocol on my side (duels need 60β85% bases, or they only answer collapse). And the vendor's own presence_penalty 1.5 breaks thinking templates on this harness (500 peg-native) β all runs at 0.0, verified clean over 3Γ164 tasks.
The KLD bet, stated plainly so the ledger can judge it: I'll get KLD on this trio however I can (no BF16 ref on hand; the Q6 duel stands without it). If ASHQ prints worse KLD than Q6_K while holding better HE β your rank column fails to rank capability, and I'll say so on the card next to my own retractions. If it prints better on both β the pair is trusted, per our rule, and your column survives its hardest test yet. Either outcome gets logged. But note the fair fight isn't me-vs-stock: send me your Quality-36 file (or exact recipe) and it goes through the same chain. Your champion vs mine, on the verdict column.
New repo wepiqx/NeoHorse-1-9B-MERNIK-GGUF is up β MERNIK branding from here on (method repo opens soon). Only the winner inside; card leads with HE, PPL/KLD marked coming-soon, with the ctx-floor note (8192/2048 trips thinking models; 16384+ re-runs planned).
Matrix closed: full 2Γ2 (MSE/SMAPE Γ BU/TD) plus the small row, three columns each β wepiqx/MERNIK (https://huggingface.co/wepiqx/MERNIK). Headlines: budget law (SMAPE owns small budgets +64pp @5100, MSE owns big @6500, same PPL both times); PPL-blind 5GB cliff (86β18% HE at PPL 7.80); norms finding (PPL-identical allocations differ β3.7pp on HE β your charter floor, confirmed from my side). Quants: NeoHorse-1-9B-MERNIK-GGUF (https://huggingface.co/wepiqx/NeoHorse-1-9B-MERNIK-GGUF). Quality-36 file still wanted for the fair fight.
Hey! The gap stretched again on my side, but everything got read in one pass: the HE triple, the Matrix closure, the MERNIK page. Good batch to answer, since your results landed right on the questions left open. And the page reads great: the two rings are the protocol this exchange needed in writing, and "allocator signals never certify, capability scores never steer" is the line I'd put on the protocol card. MERNIK suits the method, too.
Suite state, briefly, because it changes how the numbers travel. Both methods outgrew their labels: ASHQ1 became MERNIK on your side, ASHQ1-Remix became the Paretrix Quantization Suite on mine. The allocator is GSQ-RCO now, and there is one verdict regime end to end: Long Horizon, 4096Γ8, same 32,768-token span, with 512Γ64 staying the probe regime. The nine families got re-scored under that single regime, stock twins included, which is why some numbers moved between threads. A harmonization, not a retouch, and a needed one: Spark's verdicts bind at native context, and cross-model comparisons wanted one scale. The short-context line mispriced band-heavy builds in both directions (a 512 tie turned into a -18.7% win at 4096; a -0.036 win decayed to a tie). The Ornith dominations held through it: Compact still takes -0.0403 KLD at -381 MiB off Q5_K_M-imx. ASHQ1's queue formulation, tied-group hashing and MSE scheduling stay credited as the seed; the method and the code are their own line now. So the fair fight reads Paretrix vs MERNIK, both under their real names.
The HE triple. Q2_K at a literal 0.00% next to 86 and 82 is the cleanest collapse proof in the ledger, and the +3.7pp over stock Q6_K at -0.5 GB is a claim the fast ring could never make. The saturation rule going protocol is right, too: a near-saturated base can answer the collapse question, never the ranking one. And the ctx-floor note (8192/2048 trips thinking models, 16384+ re-runs planned) belongs on every card. The silence audit, by the way, deserves first place in the method repo: empty completions as a weakness scale, with finetunes rather than sizes deciding silence. The slow ring catching what no distribution column can, exactly.
The KLD bet. The resolution reads as it should, and I'll take (b) as written: on the verdict column, the rank column does not transfer. It never claimed to. On (a), one amendment from my side: "punishes the pie for being a pie" stays a hypothesis until a reference prices it. Pies are not KLD-doomed in my ledger: Ornith's Compact (a DP build) beats its stock flat twin by -0.0403 KLD at -381 MiB, lfm2's Nano by -0.0116 at -42 MiB, nanbeige's Quality by -0.0158, all against BF16 and all at one regime. And the ~4Γ gap on your card sits on the Q8 proxy, not BF16, so that column keeps its caveat until a true reference prices it: stays ~4Γ, and the card reads "code capability bought with distribution fidelity", a real and priceable trade; shrinks, and part of it was the proxy. Either outcome takes (a) from expected to measured.
The norms receipt. This is the one I wanted and couldn't generate myself: PPL-identical allocations, -3.7pp apart on HE, with the native shield recovering +1.2pp of it. Norms matter, the rest is elsewhere, and the post-hoc shield disrupting the path (-3pp on SMAPE) is the finding inside the finding: the floor has to be native, not bolted on. First capability-side confirmation of the norms floor, refinement included, and it's logged. Footnote: the full F16 norm set prices at ~2 MiB on this model, so that's 1.2pp of code for two megabytes.
The budget law, and its correction arc. The clean version (SMAPE owning 5100 by +3pp, MSE owning 6500 by +3.7pp, same PPL both times) is the operating-point rule seen from your side: the metric that ranks allocation quality is a property of the budget, so no single utility owns the ladder. And the arc around it is the most valuable thing on the page: a dead server scored silently is the harness version of PPL sharpening, where the number looks fine while the thing underneath is gone. The watchdog that aborts loudly is the right fix, and keeping the scar while removing the cliff is how it should read.
The fair fight, and what I actually run. You asked for my Quality-36 file, so here's the honest version: NeoHorse-1-9B isn't a model I keep, so that build won't come from my bench. My daily driver stays Ornith-1.5-9B, and NeoHorse-1-4B is the sub-agent. Either way, there are two routes to a pair. One: the recipe is public and it's a single grid, flat Q5_K_M + imatrix + the recurrent floors (ssm_out Q6_K, Ξ±/Ξ² Q8_0), so build it on your NeoHorse-1-9B and your chain crowns or kills it; either verdict is a win for the ledger ;). Two: my Ornith champion is already public, Quality-36 in Soulfate24/Ornith-1.5-9B-MTP-DFlash-Paretrix, sitting next to your ASHQ1-6500 in wepiqx/Ornith-1.5-9B-MTP-ASHQ1-GGUF; point the HE chain at both files and the pair gets its verdict column. And a word on the sub-agent itself: NeoHorse-1-4B turned out to be a great one, 160K to 256K context on an 8 GB card, and I grafted a Qwen3.5-4B MTP head onto it plus reused the Qwen3.5-4B MMProj, so long-context explore, extract and verify runs with speculation and vision on top. Matched donor, same Qwen3.5-4B bones, and it's the most complete little setup I've had.
Everything I run is on the profile, per model. One repo per tested model, each shipping the imatrix, the GGUFs and the model-paretrix.json (plus the modules when the model has them). That last file is the time-saver: the measured rate tables, the allocations, the cocktail recipes and the champion receipts all live in there, so the cocktails and the duels can be read without re-running anything. And to test your own GGUFs against mine, it's two steps: drop them into the downloaded model folder, then
python 11_perplexity-test.py "Ornith-1.5-9B-MTP-DFlash-Paretrix"
Hit a for everything in the folder, or pick numbers; unattended, echo a | python 11_perplexity-test.py "Ornith-1.5-9B-MTP-DFlash-Paretrix". Same regime, same corpus, your file lands in the same sweep as mine (KLD rides along when a reference base sits in the folder; otherwise the PPL canary does the first pass). That's the whole manual, no docs needed. The suite is far friendlier than the Remix days, so use it freely for your own tests. And if you skip brewing Paretrix recipes and just reuse the imatrix from the repos, everything runs fast: the slow part, the activation pass, is already done.
Open on my side. The Spark-4B calibration swap with your Qwen3.8 set is still the next run here, and the bubble gets its one-variable answer. The Ornith trunk-parity arm still stands, and the head rule is one line: the head block's attn/ffn projections and eh_proj at Q5_K, everything else untouched, same 6511 MiB. Build it and the pair goes through my battery, or dump your --show-config and I'll mirror it here; either way the clean pair settles whether the head lane mattered. And your Gini answer landed (0.498, late units carrying), so the dense-untied reading leans the slope bet's way; the write-up decides by how much.
A long scar column is the price of a public ledger, and the part of it that keeps the rest honest. Post when the runs land, everything scores as it arrives. π€
Hey! Read in one pass as well β good batch indeed. Point by point, with receipts where I have them:
The KLD bet. Taking your amendment on (a) as written: "punishes the pie" stays a hypothesis until a BF16 reference prices it, and our ~4Γ keeps its Q8-proxy caveat on the card. (b) stands as agreed β the rank column does not transfer to the verdict column, and it never claimed to. One consequence on my side: our Gnom labels were measured on the short span, so they now carry the same caveat your harmonization forced on everyone. Logged, not hidden.
Long Horizon. That one was my proposal from way back (we only ever measured the near field, never how a model behaves far out), so seeing it become the single regime is a quiet win I'm taking. But it needs one written line between us: your verdicts now bind at 32K tokens, mine still at the short span + HE. Same words, different rulers β no cross-comparison until one of us re-measures. Stating it so it doesn't rot.
The calibration swap. Already done on my side a while back (it's in the local logs) β the bubble got its one-variable answer. I'll dig the numbers out for the next letter rather than re-run.
The fair fight. Taking route two: I'll point the HE chain at your Quality-36 and our ASHQ1-6500 and hand the pair its verdict column. Downloads queued behind the current battery. Route one stays open for later.
Trunk-parity. Easiest path is yours: I'll dump the --show-config of our Ornith build and you mirror it there. One variable, clean pair.
The engine upgrade (the actual proposal). You're still running pure MSE from the Remix days β and since the fork we've measured the difference between MSE and what came after: SMSE/SRMSE, the Ξ±-axis (peak near 0.6β0.75, MSE itself drops off at Ξ±=1.0), huber, recovery, plus the harness side (preflight audit, HOLE/TERMINAL failure semantics, the loud watchdog). On FrogNano-4B the spread between the best and worst utility at equal weight was 13 tasks. That's not a tweak, that's a tier. The MERNIK allocator is a drop-in where your GSQ-RCO sits β same queue formulation you credit as seed, same tied-group hashing, just the scheduling and the tooling grown up. Take the engine, keep your name on the suite. The duel afterwards (Paretrix vs MERNIK, both real names, as you said) only gets sharper if both run the same generation of allocator β otherwise you're fighting one behind.
Mellum-2.1, since you'll ask. First blood on the new JetBrains MoE: our 7500-SMSE (7.44 GB) leads stock Q6_K (10.88 GB) 125 vs 119 on HE with the duel at honest 0/3 NOISE β a lead, not a crown, stated as lead. Repo opening now, 10000 build lands as soon as the 1070 stops faulting its Q8 kernels. Routers F16 on all 28 layers, per the old Mellum-2.0 recipe β some things don't change. π€
Quick status from the bench β the fair fight is queued, with receipts already in:
What's running: Mellum-2.1 10000-SMSE HE (third attempt β the 1070 kept faulting Q8 kernels under OC, now on stock clocks + --parallel 1, which turned out to be the RAM-OOM cure), then 6500-SMSE battery, then recovery-6500, then the Ornith pair: our 6100-SMSE vs your Quality-36, plus the MTP mirror since you uploaded it. Chain is armed, results land as they arrive.
Tensor distributions (audited from the files, both sides): ours is a pyramid (Q4 93 / Q5 26 / Q6 42 / Q8 97 at 6109 MiB), yours is flat (Q5 143 / Q6 57, zero Q4, 6223 MiB) β the recipe exactly as published. So the duel is pyramid-vs-flat, same imatrix (yours, 4.9 MB, thanks), same weight within 2%. One handicap disclosed upfront: our convert carried the blk.32 nextn (MTP) block, yours doesn't β we fight with the weight on, and if we win with it, that settles the engine question louder than any rematch. MTP-vs-MTP after, for completeness.
The MSE point, stated plainly: your Pareto reads "no better trade at that compression" β ours reads "no better trade under MSE". The budget law we measured says the ranking metric is a property of the budget (SMAPE owns 5100, MSE owns 6500, huber took FrogNano), so a single-utility frontier is a slice, not the space. The duel arbitrates the slice. Either verdict is a win for the ledger. π€
The fair fight has its verdict column. Numbers first:
- Ours (MERNIK-6100-SMSE, 6109 MiB, your imatrix): HE 140/164, HE+ 135/164
- Yours (Quality-36, 6234 MiB): HE 129/164, HE+ 122/164
- Duel: 2/3 SIGNIFICANT (HE+ 20/7 p=0.0192, empties 3/13 p=0.0213; HE 21/10 p=0.0708, noise)
Same base, same imatrix, weight within 2% β different allocator. And the handicap sits entirely on our side, stated exactly: our convert carried the blk.32 nextn (MTP) block, 15 tensors that never participate in inference yet eat budget. So we fought heavier with less usable weight, and the pyramid (Q4βQ5βQ6βQ8) still beat the flat (Q5+Q6, zero Q4) by 11 tasks.
This is the (a)-debate settled the only way it could be: on the verdict column, never the rank one. Your cocktail owns KLD/PPL β Quality's PPL is the better number on your card and I don't dispute it. But the tuning that minimizes distribution distance does not maximize code: you sharpened the proxy, the verdict went the other way, significant on two of three columns. "Punishes the pie for being a pie" has its measured answer now β the pie wins where it counts, and the ~4Γ on our card is capability bought with fidelity, priced and stated.
So the engine proposal stands, upgraded from suggestion to receipt: MSE-flat is a slice of the space, and the slice just lost to the queue with a weight on its back. Take the allocator β same seed you credit, grown up β and the next duel is Paretrix-vs-MERNIK at full strength instead of one generation behind. The MTP mirror (your 6.71 GB file downloaded, battery queued) decides whether the head lane moves the verdict; my bet is it doesn't, and I'd love to be wrong.
Ledger updated, all three columns. Post when the mirror lands. π€
P.S. One symmetrical ask: you never ran our file through your battery (PPL/KLD/Long Horizon) β the verdict column is ours, the rank columns are yours, and both should travel. Where do I drop the 6100-SMSE for your 11_perplexity-test.py β HF repo, or your preferred lane? It's 6.1 GB, ready when you point. π€