Paretrix_Quantization_Suite / DOCUMENTATION.md
Soulfate24's picture
[2.6.0] - 2026-10-06
093df17 verified
|
Raw History Blame Contribute Delete
43.2 kB

Paretrix Quantization Suite — Documentation

Technical reference for Paretrix v2.6.0: architecture, pipeline, calibration ledger, knowledge base, campaign engine, modules, and the evaluation standard. Benchmarks live in README.md, the single canonical measurement archive.


1. Philosophy & Architecture

1.1 The Pareto principle applied to quantization

Every quantization decision is a trade: bytes freed against fidelity lost. Paretrix measures that trade per tensor class (ΔKLD/MiB) at a real operating point, then allocates under an exact budget so that no better trade exists at that size (the Pareto frontier of the (MiB, KLD) plane). The shipping name states its weight band, and the quality objective optimizes inside that band.

Three mechanisms deliver this:

  1. Flat recipes where uniformity wins (mid-band and top-band): a structural base type plus measured levers (recurrent floors, band raises, readout levers), emitted as llama-quantize regex rules.
  2. Exact-budget DP where heterogeneity pays (low-band and squeeze campaigns): a multiple-choice knapsack solved to the MiB, priced by measured notch rates and measured buy rates.
  3. Duels at the verdict regime decide adoption: a candidate takes the rung name only on a KLD gain at or above 1×MASD (floor 0.001) at ctx 4096×8.

1.2 THE MATRIX — three-level knowledge base

Level File Scope Precedence
1 model-paretrix.json This model's measured tables (anchors, probes, rates, allocations, cocktails, champions) First
2 arch-paretrix/<arch>-paretrix.json Per-architecture shared line: banded sell/buy tables, facts, corrections, witnesses Second
3 arch-paretrix/generic-paretrix.json Universal prior (banded mid/deep, distilled from six campaign families) Third

Each level overlays the one below class by class: a class that an upper level never measured prices from the level below.

Shared levels (2, 3) are size-normalized: rate × line_reference_mib / model_mib. A 1.5B and a 9B model of the same architecture therefore share one line. Measured effect: the between-checkpoint rate spread tightens from ×4.7 to ×2.1. Lines are banded: mid (floors at or above Q4_K) and deep (IQ4_XS and below), selected per group by its floor depth.

1.3 Rate-exchange theory

At any operating point, every class carries two measured prices: sell (KLD/MiB freed by a notch cut) and buy (KLD/MiB gained by a raise). Allocation quality depends on that exchange. With both ends noise-clamped (free-sell and dead-buy clamps), the exchange is win-or-tie by construction against its own anchor: ties mark inversion-free basins (T1). Anchor choice dominates the table you see (T2); band rates are ctx-dependent on SWA topologies (T3); fine-tunes flatten the field (T4). The default path trades on priors and the campaign trades on measurement, and the D-gap measures the prior's error at the operating point (T5).


2. Pipeline Overview

[HuggingFace checkpoint (safetensors + sidecars)]
          │
          ▼  01_SAFETENSORS-to-BF16-GGUF.py
[<model>/Paretrix/model-BF16.gguf] + [mtp-BF16 · mmproj-BF16 · dspark/dflash-BF16]
          │
          ▼  02_BF16-GGUF-to-Q8-imatrix.py
[model-Q8_0.gguf] + [imatrix.gguf]
          │
          ├─▶ 03_BF16-GGUF-to-Paretrix.py  →  [model-Paretrix-<Tier>-<XX>pc.gguf]
          │                                    [+ classic Q8_0 / Q6_K / Q5_K_M / IQ4_XS / IQ3_M]
          ├─▶ Paretrix-modules.py          →  [mtp / mmproj / dspark / dflash-Paretrix-<Profile>.gguf]
          └─▶ 14_gguf-module-fusion.py     →  [fused deployment GGUF]
          │
          ▼  11_perplexity-test.py (verdict regime 4096×8)
[PPL · KLD · RMS Δp · top-p · Pareto frontier]
          │
          ▼  Paretrix.py duel --adopt  ·  Paretrix.py receipt  ·  Paretrix.py matrix
[model-paretrix.json] → [arch-paretrix/<arch>-paretrix.json] → next model

Orchestrated by 00_Paretrix-pipeline.py, which resumes from its working folder (the folder holds the state).


3. Script Reference

3.1 00_Paretrix-pipeline.py: orchestrator and shared library

Runs 01 → 02 → 03 with resume semantics: existing outputs skip automatically. The script also serves as PARETRIX COMMON, the shared library that every other script imports as PC (console grammar, live-line child runner, GGUF helpers, file utilities, and the re-shuffle composer). Module repos join the same run: 01 converts each extra folder into the working root, and the module is quantized immediately.

python 00_Paretrix-pipeline.py <safetensors-folder|working-folder> [module-repo ...]
python 00_Paretrix-pipeline.py MyModel MyModel-DSpark --all
python 00_Paretrix-pipeline.py MyModel/Paretrix --from 03
python 00_Paretrix-pipeline.py MyModel --dry-run

Key flags: --from {01,02,03} · --all · --tier a,b · --paretrix-only · --no-campaign · --no-audit · --no-duel · --no-cocktail · --no-fold · --no-modules · --module-profile {balanced,compact} · --outtype {bf16,f16} · --cpu-only · --force · --keep-going · --dry-run.

3.2 01_SAFETENSORS-to-BF16-GGUF.py: conversion

Converts safetensors into a pristine BF16 GGUF with a native split:

Output Content
model-BF16.gguf Text model; MTP head and vision tower excluded
mtp-BF16.gguf Standalone MTP fusion head (when present)
mmproj-BF16.gguf Vision projector (multimodal models)
dspark/dflash-BF16.gguf Draft repos convert whole, under their own stem

The full model (text + head) is a build intermediate, removed once the head is extracted. --keep-full retains it, and 14 rebuilds it exactly. If the converter rejects --no-mtp, the full conversion runs first and a structural MTP strip produces the text model. Paretrix accepts raw BF16 sources only; pre-quantized checkpoints (qweight/scales/qzeros keys) are refused.

Also handles: RoPE config normalization, remote-code compatibility patching, preprocessor config completion, nanbeige padded-vocab compatibility, shard-index healing, single-file staging, and immediate module quantization through Paretrix-modules.py (--no-modules defers it).

3.3 02_BF16-GGUF-to-Q8-imatrix.py: Q8 source and imatrix

Produces model-Q8_0.gguf (kept for 03's archival classic) and imatrix.gguf (activation statistics for the quantization grid). The Q8_0 source speeds up offloadable layers and CPU dot products, while the statistics stay quasi-lossless (measured ρ 0.99999, top-100 overlap 100/100; see L13). IMATRIX_SOURCE_QUANT=bf16 runs on the BF16 source instead.

GPU autotune opens a descending probe ladder at the computed geometry boundary (fit-era slot accounting, calibrated fixed reserve). It then runs iso-chunk recurrent split validation (a broken GDN split multiplies PPL) and verifies PPL on the final run. Every candidate carries an explicit -ngl (0 = CPU). Unified CUDA memory stays off, so an out-of-memory condition surfaces as a hard allocation failure.

Corpus: datasets/bartowski-imatrix-v5-semantic.txt (downloaded once, or adopted from the legacy folder).

3.4 03_BF16-GGUF-to-Paretrix.py: the Paretrix core

Two phases per model:

Phase 1: caches. L0 architecture audit (10, first contact only) · arch line loaded (or noted absent) · model cache created or loaded.

Phase 2: champions (model-Paretrix-<Tier>-<XX>pc.gguf), per armed tier:

  1. Champion on disk → skip (with the R19 stock-twin check).
  2. Allocation in the model cache → direct build.
  3. Fresh champion above (L15) → campaign (preflight → anchor → notch probes → rates → exact-budget DP → build → duel → adopt → cocktail).
  4. Neither → engine default path (flat recipes / DP from priors).

The script also emits the classic line (Q8_0 archival from 02, plus low-band Q6_K/Q5_K_M/IQ4_XS/IQ3_M-imx on lever-less topologies), and closes with the arch fold (Paretrix.py matrix).

Key flags: --all · --tier · --paretrix-only · --no-campaign · --no-audit · --no-duel · --no-cocktail · --no-fold · --force · --dry-run.

3.5 Paretrix.py: engine and campaign CLI

The quantization engine (classifier, flat recipes, exact-budget DP, Gumbel-STE search) and the campaign namespace:

Subcommand Purpose
campaign Anchor → notch probes → rates → DP → build (one rung, one command)
duel Measures candidate vs incumbent at the verdict regime; --adopt takes the rung name
adopt / receipt Naming and receipt management (--scan backfills from sidecars and the eval cache)
cocktail Composes a re-shuffle from measured arms (overlap exclusion, value-positive bundles)
build Materializes an allocation sidecar (--from-cache for direct rebuilds)
verify Plan-vs-artifact tensor type comparison
polish In-format STE polish (measured: 0/25 tensors improved; kept as a falsified arm)
matrix Folds every model cache into the arch lines (THE MATRIX)
cache / compact Inspects the knowledge base · compresses tensor maps (col1 format, ~20–30× smaller)
atlas Knowledge base evidence: rate field, lineage, probe economy, routes, lint
digest Champion composition digest (class × tier × depth per rung)
diff Byte-level GGUF regression check (tensors + KV)
selftest DP engine, manifold ops, STE loop

Legacy entry points: Paretrix.py --model … --profile … --size … --run (the allocator CLI) and Paretrix.py --tier all (the tier runner).

3.6 Paretrix-modules.py: MTP · DFlash · DSpark · MMProj

Imatrix-free module quantization. Kinds are auto-detected from the GGUF (--kind overrides): draft (DFlash/DSpark), mtp, and mmproj. Role-based floors apply to draft and MTP modules (bridge ≥ Q6_K, proposal ≥ Q5_K, norms F16/F32, 1-D tensors pinned F32, unclassified tensors Q5_K). MMProj follows a profile grid (critical tensors pinned F32, deep-boost on trailing blocks). Available types are K-quants, Q8_0, and F16/F32 via gguf-py, with a native llama-quantize fallback and artifact-vs-plan verification.

python Paretrix-modules.py --model mtp-BF16.gguf
python Paretrix-modules.py --model dspark-BF16.gguf --profile compact
python Paretrix-modules.py --model mmproj-BF16.gguf --dry-run

3.7 10_arch-inspect.py: architecture audit

Five lenses on a GGUF or safetensors folder:

  1. Tensor inventory: class × count × dtype × MiB × shape.
  2. Per-layer signatures: deviant layers (MTP heads, hybrids). Periodic patterns are read as hybrid design; only true anomalies raise a warning.
  3. Novelty scan: suffixes outside llama.cpp's known taxonomy.
  4. Binary support cross-check: arch and every tensor name template checked against llama-arch.cpp.
  5. Imatrix coverage: quantizable weights without activation statistics (vectors and lookup tables count as by-design and are not listed).

--brief (used by 03) shows the top classes and the verdict lines.

3.8 11_perplexity-test.py: perplexity and fidelity sweep

Sweeps every product in a working folder at the verdict regime: PPL, KLD (vs BF16 reference logits), RMS Δp, top-p, t/s, chunk stability (tail σ / MASD), and the Pareto frontier column. It offers an interactive menu (a = all, indices, names, globs) or piped stdin for automation (echo a | python 11_…).

  • KL bases are per-regime (kld-bf16-ctx4096.dat, etc.) and per-model, generated once.
  • Eval cache (eval-cache.json): KL results keyed on file identity × regime × KL base × binary build, so duels and sweeps reuse measurements.
  • Offload planning inherits 02's calibrated geometry (fit-era slot accounting, KL-mode reserve, recurrent split validation via 02's iso-chunk protocol).
  • Regimes: default 4096×8 (Long Horizon, L18 verdict) · --medium 2048×16 · --light 1024×32 · --short 512×64 (probe regime). Span = 32,768 tokens across all.

3.9 12_attribution-probe.py: marginal knockout probes

Measures ΔKLD/MiB of single-class levers at an operating point. It quantizes the arm, evaluates at the probe regime (512×64), and records the result in a resumable CSV. Campaign probes run with --no-summary; the batch closes with one --summary-only attribution table (solo screen, exchange rate, shrink/re-shuffle chains). The inert-arm guard detects type-identical builds. PARETRIX_PROBE_NGL controls offload.

3.10 13_draft-acceptance-sweep.py: A/B acceptance battery

Spawns llama-server with and without the module, replays fixed prompts at temperature 0, and records acceptance, mean length, and throughput per --n-max into draft-acceptance.csv. The trained block is read from the module metadata (*.block_size). --chat wraps probes in the target chat template. --ngl is recorded, and baselines are compared only at matching offload (L20). The thermal guard flags baselines that drift from the reference.

3.11 14_gguf-module-fusion.py: module fusion

Merges 2+ GGUF files into one. The first file wins on duplicates, and block_count plus per-layer arrays widen when a merged head extends the layer index range, touching <arch>.block_count only, never a sub-module's count. MTP folder mode rebuilds the model-mtp-* deployment files. Output lands as .part and renames on completion. Float outputs receive Xpress8K NTFS compression.

python 14_gguf-module-fusion.py <folder> --mtp [--target <product.gguf>] [--head <mtp-head.gguf>]
python 14_gguf-module-fusion.py out.gguf main.gguf extra.gguf [...]

4. The Mod-3 Ladder

Ratios are multiples of 3 percent of BF16, listed from largest to smallest. The shipping name carries the ratio (-<XX>pc), enforced at ±1.5 pp (R14). A build outside every band is withdrawn. A build inside a neighboring band re-homes to that rung, or is withdrawn if the target name is already held, so a rung name never ships outside its band. The gaps (37.5–40.5, 43.5–46.5) own no rung.

Tier Ratio Base Key levers
Fidelity 48% Q6_K + Q8_0 pockets biased to the late layer third (G4)
Precision 42% Q6_K + band Q8_0; readout lever arch-scoped (L12)
Quality 36% Q5_K_M L6 flat + imatrix + recurrent floors (inert under regex rules on stateless topologies)
Compact 33% Q4_K_M Flat + band Q6_K + gate Q6_K + recurrent floors, or DP
Mini 30% IQ4_XS L4 flat on GDN (ssm_out Q6_K, α/β Q8_0, band Q6_K) or DP
Nano 27% IQ3_S–IQ4_XS DP with measured rates
Pico 24% IQ3_S Band Q5_K + gate Q6_K (flat) or DP
Femto 21% IQ3_XXS Deep-squeeze DP; campaign product

Environment knobs per rung: PARETRIX_INCLUDE_PICO/FEMTO (scope: state-rich ≥3B) · PARETRIX_INCLUDE_PRECISION/FIDELITY (default on) · PARETRIX_QUALITY_ENGINE (native flat / allocate DP) · PARETRIX_FEMTO_CLASSIC · PARETRIX_FEMTO_RECIPE.


5. Calibration Ledger

Binding laws (L) and rules (R), measured across nine families (state-space hybrids, dense SWA networks, looped untied trunks, shortconv mixers, fine-tunes, and vision-aligned backbones). The code ledger in Paretrix.py is the living source; this section is its reference form.

5.1 Laws

  • L0: Pre-adoption audit. Untied lm_heads expose output.weight. Run 10 before conversion so the readout is classified correctly instead of parking at the worst tier (normalized through the readout class since 2.5.0).
  • L1: Readout and embedding value. Keep tied token_embd at Q6_K or higher on production tiers: 94 MiB per top-p point (93% of Q8_0's gain).
  • L2: FFN information bottleneck. At ≤33% budget, ffn_down is the primary bottleneck. Sub-IQ4_XS cuts cost ~2e-4 KLD/MiB, an order of magnitude worse than any winning promotion. ReaderLM-v2 (75% FFN mass) validated the Q4_K floor restoring monotonicity.
  • L3: Depth prior. Surplus flows to early recurrent state, then to full attention (scope: ≤33%). In the top band, Q8_0 pockets land in the late layer third in 6/6 families (G4 refinement). Deep-band ffn_gate_up cuts land early in 4/4.
  • L4: Recurrent readout floor. Keep ssm_out at Q6_K or higher on production tiers: sub-Q6_K collapses output entropy (Ornith: KLD 0.081 → 0.0241 once restored). One exception is measured: the Xiaomi cocktail sells ssm_out to Q5_K and still wins, so the result holds per family and does not generalize.
  • L5: Small-model constructibility. ≥9B at Mini · 3B–4B at Compact · ~1B at Quality. 0.33–0.40 KLD is the entropy-dissolution ceiling (three convergent families). Restrict --allow-q3-or-lower to grids at or above 25%.
  • L6: Mid-band uniformity (price-scoped). Proxy-priced upgrades pay a heterogeneity tax: Quality-36 delegates to flat Q5_K_M. With both sides rate-measured, the DP reclaims the band (lfm2 −13% relative, nanbeige −34% relative).
  • L7: Tied readout invariance. token_embd doubling as lm_head rides the output floor (Q6_K). Its lever works as an additional purchase, not as a swap inside a pinned budget.
  • L8: Cut locality over depth. Scarce full-attention bands are the currency below ~30%. Above that, the uniform body dominates.
  • L9: Readout economy. An untied lm_head rides its class floor. Input-side embd purchases above Q5_K price as noise on most families (3–10× below the body step). Exception: nanbeige-4.2 (2.68e-4/MiB, comparable to its body step).
  • L10: Unmeasured classes follow their declared floors, not the upgrade queue.
  • L11: Speculative verification bound. Draft quantization moves acceptance by at most 0.03 (nanbeige +0.0104, lfm2 +0.0055, MiniCPM5 bit-identical) and speeds the draft itself (+13–21% t/s).
  • L12: Readout lever. High-precision Q8_0 on the logits matrix (tied token_embd, untied lm_head) delivers measurable gains. This is an arch-scoped correction: spark2_5 carries it, and the qwen35 default downgrades it at the top band.
  • L13: Rate-calibrated allocation. Notch probes and measured buys feed the exact-budget DP (budget hit to ±1 MiB). Rates transfer across anchors within ~5% on bulk classes.
  • L14: Anchor-basin economics. A DP squeeze re-ranks within the anchor's basin, so anchor choice dominates (same Mini slot: 0.0687 fresh vs 0.1035 consumed).
  • L15: Nearest-fresh-champion targeting. The anchor is the nearest champion above that has not been through a deep squeeze (flat/allocator lineage). Reach is rate-mediated, not distance-mediated.
  • L16: Top-band buy saturation. Measured buy rates extend linearly only to Q8_0; F16 buys fall back to the MSE proxy. The top band's sells are nearly free and its buys real, so re-shuffles win where buys are cheap.
  • L17: Self re-shuffle primacy (conditional). At a defended rung, attack with the self re-shuffle. The incumbent's builder lacks the RCO vocabulary, and the self anchor is priced at the operating point. This holds on a live buy table. When ≥3 of 5 buys are dead (saturated basin), the fresh champion above wins instead. Rate test (L17-bis): the re-shuffle wins if and only if the max affordable live buy rate exceeds the min sell rate. After R13, it wins or ties by construction.
  • L18: SWA long-context binding. On windowed-attention topologies, the ctx-512 duel misprices band-heavy allocations in both directions (Quality +0.0005@512 → −0.0107@4096; Pico −0.0361@512 → −0.0033@4096). Verdicts bind at the native context, and band rates are regime-bound.
  • L19: Fine-tune field flattening. A deep fine-tune compresses the rate field: sells fall 3–10× under the base priors, and most buys die (NeoHorse 7/7). Campaign value concentrates in deep sheds and self re-shuffles. Distinguish this from dead-lever families: flat sells with live buys point to L12/L16, not L19.
  • L20: Draft-target pairing and regime term. Acceptance is a property of the (draft, target) pair and the offload regime (±0.02, four times the L11 drift). Read acceptance only at matched ngl. Pair each draft with the least-quantized, highest-ngl target the budget allows. Draft-side budget: weights + a fixed block graph (block_size=16 → ≈1 GiB) + ≈0.1 GiB KV.

5.2 Rules

  • R3: Loop-neutral depth prior. On looped trunks (num_loops > 1), keep the early-layer downgrade prior disabled. The imatrix integrates every visit, so the prior runs against the measured gradient (nanbeige: blk.0 at 4.1% vs blk.21 hottest, a 24.45× gradient).
  • R6-bis: Readout-share parity. When an untied readout pair exceeds 30% of the mass, demote the input side to stock Q4_K parity and reinvest the freed bytes inside the pool. Top-p corollary: cap readout sells at one notch, because the d2 notch causes a top-p cliff that the KLD line does not price.
  • R6-ter: Tied vocab share. A tied embd notch saves 0.066 × share of the model: 2.2 pp for Xiaomi (34%), 1.0 pp for NeoHorse (15%), and 0.5 pp for Spark (8%). At a share ≥ 15%, the floors are Q5_K at Mini and Nano, Q4_K at Pico, and IQ4_NL at Femto (every campaign champion sits there). Below 15%, Q6_K holds (at 4096 ctx, Spark's flat Pico won). PARETRIX_TIED_VOCAB=0 restores the flat L7 floor.
  • R7: GDN low-band flat route. Uniform base + recurrent floors (ssm_out Q6_K, α/β Q8_0) + full-attention band + stock-parity readout on SSM hybrids. The arch's measured ffn_down lever stacks where declared (spark2_5: Q4_K, −0.0183 KLD @ +28 MiB).
  • R8: Concave lever stacking. Joint gains from stacked levers measure at about 91% of the independent additive sum (deep cuts cost super-additively). Validate composite multi-arm builds as a whole.
  • R9: Multiplicative head-gate preservation. Post-projection sigmoid gates compound error along the sequence. The F32 pin costs only ~6 MB. Arch-scoped (spark2_5 correction gate_pin: F32).
  • R11: Probe-slot economics. Same-bpw ladder steps (Q4_K/IQ4_NL at 4.5 bpw, Q3_K/IQ3_S at 3.44 bpw) free zero bytes yet consume a probe slot. Deduplicate by byte delta before applying the probe cap; the freed slot goes to a class's depth-1 notch.
  • R12: Re-shuffle ride parity, priced (R12-bis). At or above the anchor's size, the input side re-sells below the anchor tier only at a priced rate. The planner sells only when the probed ride rate beats the best measured buy; an unmeasured ride defaults to parity.
  • R13: Dead-buy clamp. A probed raise that fails the noise floor (ΔKLD ≥ −0.001) records at rate 0. The DP then skips that buy and uses the measured zero rather than the MSE proxy (the mirror of the free-sell clamp).
  • R14: Size-band identity (±1.5 pp). A rung's name is its weight band. An out-of-band build re-homes to the rung that owns that band (recording planned_tier) or is withdrawn. The ladder gaps own no rung. --force-rung keeps any weight as a measurement product.
  • R15: Nearest-fresh-champion targeting (the rule form of L15). resolve_anchor picks the nearest fresh champion above; a recipe receipt routes to the default path.
  • R16: Pareto-gain adoption. A candidate with ΔKLD ≤ −N, concordant value axes, and ΔMiB < 100 is adopted under the aspirational gate. The RMS validator records disagreement up to 2×M and vetoes anything above that.
  • R17: Buy-transfer scoping (L18-bis). Body buys transfer to the verdict regime. Readout-tied buys transfer on dense-tied trunks. Input-side buys do not transfer. SWA band buys transfer only at native ctx.
  • R18: Anchor-scoped flattening (L19 refined). A fine-tune kills buys at the mid band and revives them at the deep band. The detector keys on ≥ 80% dead buys at mid or deep anchors; flat sells separate it from dead-lever families.
  • R19: Stock-twin detection and self re-shuffle. A recipe whose tensor types equal a stock classic carries no Paretrix lever on that topology (byte-identical: Nanbeige Precision ≡ Q6_K-imx, TwIL Precision ≡ Q6_K-imx, ReaderLM Precision ≡ Q6_K-imx, ReaderLM Quality ≡ Q5_K_M-imx). 03 queues one self re-shuffle anchored on the twin at its own size (L17) and duels it at the verdict regime. Measured outcomes: the re-shuffle wins below the top band (TwIL Quality 0.0131 → 0.0104, −20.7%, 2.7× the gate; MiniCPM5 Quality −12.5%, Precision −21%) and ties at the top band (TwIL Precision +0.0007, inside the gate, so the incumbent holds). The re-shuffle wins or ties by construction. On small qwen35, Precision ≈ Q6_K-imx within the tie zone (Xiaomi +0.0003, not byte-identical). 03 skips the rung by default; PARETRIX_PRECISION_TWIN=0 forces it.
  • R20: Operating-point pricing. A measured table is a property of its operating point (L14). The default path prices from the anchor nearest the target (≤6 pp is L15 reach; beyond that, the shared lines price), with each class priced in the band of its own anchor floor. Price a Nano from a table measured at its own operating point: class rates escalate ~10× from mid band to dissolution, so the last campaign's top-band table misprices it.

5.3 Empirical anchors (selected)

Law / Rule Anchor
L1 qwen35-4B: token_embd Q6_K = 94 MiB per top-p point (93% of Q8_0).
L2 ReaderLM-v2 (75% FFN): Compact ffn_down IQ4_XS inverted vs Mini (0.1211 vs 0.1145); Q4_K floor restores 0.0775.
L4 Ornith-9B: sub-Q6_K ssm_out → KLD 0.081; flat Q6_K → 0.0241 at Quality footprint.
L5 MiniCPM5-1B @ 24%: KLD 0.3886, top-p 66.2%. Sub-3B models close at 24%.
L6 lfm2 RCO-36pc 0.0301 vs flat 0.0346 (−13%); nanbeige RCO-33pc 0.0951 vs flat 0.1447 (−34%).
L7 TwIL-LM3: token_embd → IQ4_XS = 0.1232; Q6_K restores 0.0500 at the same 27% budget.
L8 LFM2.5: protecting 8 sparse attention layers recovers 0.371 → 0.213; shielding mixer params recovers 0.030.
L11 MiniCPM5 DSpark 0.4419 (ledger 0.4464), ×1.89 t/s on the 2.6.0 toolchain.
L13 qwen35-4B notch rates price a second anchor within 5% (ffn_gate_up 2.60e-4 vs 2.53e-4); RCO-30pc dominates its anchor at equal bytes (0.0385 vs 0.0400).
L18 Spark Quality self: +0.0005 @512 tie → −0.0107 @4096 (−18.7% rel).
L19 NeoHorse 7/7 dead buys at the Mini anchor; D-gap 34.7% vs base priors.
R6-ter Campaign champions at Q5_K/Mini-Nano on ≥15%-share tied trunks (Xiaomi, NeoHorse); Spark (8%) keeps Q6_K.
R9 Spark-X2.5: F32 gate pin preserves 0.0044 @30% and 0.0039 @33% for < 6 MB.
R13 lfm2 Compact: 3 dead buys priced 0; self re-run recovers −0.0019 (0.0767 → 0.0748).
R19 TwIL self re-shuffle: Quality 0.0131 → 0.0104 (won, 2.7× gate) · Precision +0.0007 (tie, incumbent holds).
R20 NeoHorse Nano priced from anchor 'pico' (27.9%, 0.9 pp); TwIL Nano rejects the 8.6 pp table.

6. Precision Guarantees (class floors)

Tensor Class Floor Rationale
Norms & scales F16 / F32 Activation scaling stability. Also a toolchain constraint: llama-quantize rejects 1-D tensors and *_norm.weight.
Attention gates Q6_K / F32 (spark2_5, R9) Non-linear routing stability; F32 limits sequential drift.
Recurrent state (GDN α/β) Q8_0 Hidden-state accumulation over long horizons.
Recurrent readout (ssm_out) IQ4_XS → Q6_K (@36%+, L4) Dynamic range and sequence entropy.
Readout (untied lm_head) Q6_K (Mini+) / IQ4_XS (low) Output logit resolution. At most one sell notch (R6-bis top-p fence).
Token embedding Q6_K (tied, L7) · R6-ter ladder at share ≥ 15% Vocabulary representations.
Speculative heads (MTP/DSpark/DFlash) Q5_K–Q8_0 Proposal acceptance rates (L11: drift ≤ 0.03).
Uncalibrated tensors IQ4_XS Fallback for tensors missing from the calibration trace.
Vision projectors (MMProj) F32 critical / role-ranked Spatial visual feature alignment.

Arch-authored floors (corrections.floors.<profile> in the arch JSON) replace the generic ladder and the class hard floor for the classes they name. These are measured knowledge, authored in the JSON with no code change. Current example: qwen35 gate (Precision Q6_K · Quality/Compact Q5_K · Mini/Nano Q4_K · Pico/Femto IQ4_XS), the lowest champion tier per rung, distilled from three checkpoints.

Sub-Mini scope. Pico and Femto default on only for state-rich trunks of 3B or larger. Elsewhere they require an explicit opt-in (PARETRIX_INCLUDE_PICO=1 / PARETRIX_INCLUDE_FEMTO=1). Constructibility gates fire where floors overshoot (Xiaomi: tied vocab at 34% of mass, so Mini-30 is the floor; ReaderLM: tied vocab at the R6-ter boundary).


7. Evaluation Standard

7.1 Protocol

wiki.test.raw, Flash-Attention, fixed 32,768-token span, BF16 reference logits (per-regime KL bases). Default = ctx 4096×8 (Long Horizon, the L18 verdict regime). --medium 2048×16 · --light 1024×32 · --short 512×64 (the probe regime, which keeps rates comparable with the generic prior and the noise thresholds). Probes price at 512×64; adoption verdicts bind at 4096×8 in MASD multiples.

7.2 Quality bands

Band KLD top-p Tiers
Near-lossless < 0.0200 ≥ 95.0% Fidelity, Precision
Production ≤ 0.0850 ≥ 88.0% Quality, Compact
Usable budget ≤ 0.1650 ≥ 80.0% Mini
Edge service < 0.4000 — Nano, Pico, Femto (≥ 0.40 = entropy dissolution)

The Pareto column (11's summary) reads ties at the duel gate max(MASD, 0.001): ★ = frontier · < X = strictly dominated by X · ≈ X = lighter twin inside the tie zone · ≡ = byte-identical.

7.3 Verification controls

  • Split-integrity ratio check: recurrent models on partial offload verify PPL against a CPU reference (≤ 1.15× on identical chunks). Suspect results are quarantined. -ngl is set explicitly everywhere (0 = CPU).
  • Distribution entropy verification: KLD and top-p are read directly. PPL below the BF16 base flags entropy collapse († in tables), which is a distribution signal rather than a quality gain.
  • Calibration-corpus comparability: a corpus swap shifts absolute KLD by ~0.005 while the rank order holds. Rank products within one corpus, and compare frontier edges across corpora with that shift in mind.
  • Calibration-source equivalence: a Q8_0-sourced imatrix ranks identically to BF16 (ρ 0.999987, top-100 100/100), which makes Q8_0 the default source.
  • Artifact-vs-plan verification: Paretrix.py verify compares shipped tensor types against the allocation sidecar.

8. THE MATRIX — Knowledge Base

8.1 File layout

arch-paretrix/
├─ generic-paretrix.json        universal prior (bands mid/deep · tails · l19)
├─ <arch>-paretrix.json         per-arch line: schema paretrix-arch/3
│    facts · corrections · witnesses · model_types · lines{base, finetune}
│    lines.<name>.sell{class: {segments, esc}} · .buy{class: {depth: {lo,hi,rate}}}
│    lines.<name>.bands{mid, deep} · .buy_census · .size_mib · .depth_prior
└─ …
<model folder>/Paretrix/
├─ model-paretrix.json          schema paretrix-model/2
│    model · arch · model_type · mods · anchors · probes · rates
│    allocations · cocktails · champions · notes
├─ eval-cache.json              KL results (11) keyed on file × regime × base × binary
├─ paretrix-overhead.json       calibrated quant-size overhead factor
└─ kld-bf16-ctx<N>.dat(.json)   per-regime KL reference bases

8.2 Pricing precedence

Model cache (level 1): the anchor nearest the target (R20), with each class in the band of its anchor floor. Arch line (level 2): banded and size-priced. Generic prior (level 3): banded and size-priced. Each level overlays the next, class by class. PARETRIX_FINETUNE=1 forces the L19 modifier (sells ×0.3, buys ×0.5 from generic-paretrix.json "l19"). Otherwise, lineage comes from the model cache's mods.

8.3 The fold (Paretrix.py matrix)

The fold compiles every model cache under the scan root into the arch lines, per (arch, mods). Each anchor lands at the line's reference size first (size_mib = geometric mean of the line's checkpoints). Sells fold banded: each notch falls in the band its anchor floor selects. The fold compiles them conservatively: sells at the range max (non-decreasing with depth), buys at the min of the range lows, and all-dead buy classes at 0.0 (R13, table-wide). The buy census counts measured raises per band (dead = ΔKLD ≥ −0.001) to drive the prior-dead skip. Authored fields (facts, corrections, witnesses, policy) are preserved across every fold, and missing structural facts are detected from a sibling model GGUF.

8.4 Analysis tools

  • Paretrix.py atlas: rate-field evidence (how tightly each pricing key predicts a measurement), lineage (L19 signature per anchor), probe economy (hindsight skip candidates), routes (what holds each rung), and lint (incoherent tails, non-monotone notches, wide ranges, orphans).
  • Paretrix.py digest [--classics]: champion composition per rung (class × MiB × tier counts × E/M/L dominant tier per layer third). This is the rule-extraction view to attach to model cards.
  • Paretrix.py diff [--subset] [--no-kv]: byte-level GGUF regression check covering tensor names, types, shapes, payload hashes, and KV pairs (values hashed).

9. Campaign System (RCO)

9.1 Flow

anchor (any champion artifact → tensor types → floors)
  → feasibility preflight (byte arithmetic; closes infeasible rungs in under a second)
  → notch probes (per class, per depth; prior-dead raises skipped — R13 census)
  → measured rates (sell segments + buy table; free-sell and dead-buy clamps)
  → exact-budget DP (anchor floors + measured squeezes + measured raises)
  → build (llama-quantize with emitted tensor-type rules)
  → artifact-vs-plan verification
  → duel at the verdict regime (ctx 4096×8, 1×MASD gate)
  → adopt (rung name + receipt) · cocktail attempt (auto re-shuffle from arms)
  → arch fold (shared lines refreshed)

9.2 Notch probes

The campaign runs one probe per class per depth (mix-notch-<cls>[-dN]) and one raise per class (mix-up-<cls>), measured at the probe regime (512×64). Same-bpw steps are deduplicated (R11). Prior-dead raises are skipped without measurement (arch census of ≥2 all-dead, R13). The baseline row replays the anchor allocation as a sanity check. Probe caches are shared between sibling rungs of the same anchor (anchor-stamp keyed).

9.3 Exact-budget DP

The DP is a multiple-choice knapsack (one option per tied group), solved to the MiB (unit quantization ~15–45 KiB). Squeeze cost per notch comes from the measured segment table (_kld_squeeze_cost), and raise gain from the measured buy table. F16 buys fall back to the MSE proxy (L16 hardening). The stage ladder extends 0 (anchor floors only) → 2 → 3 → 4 → 5, and the minimal extension that fits wins. InfeasibleAlloc at every stage marks the constructibility gate (family verdict, clean skip).

9.4 Duel and adopt

Both arms are measured at the verdict regime (4096×8 default, verdict_ctx arch override). The gate is 1×MASD of the candidate run (floor 0.001). A gain at or above the gate adopts; anything smaller is a tie, and the incumbent holds. adopt writes the verdict into the sidecar (verdict_kld, verdict_vs_kld) and the receipt into the model cache (one per tier). Verdict guards: a better-recorded incumbent stays in place unless --force is given, and a candidate that is not strictly better is refused.

9.5 Cocktails

Cocktails compose from measured arms. The exchange test runs first (one direction per tensor cluster), then sell bundles fund buys (exact subset search, minimum total cost, ≤8 arms; the ε law). Only value-positive bundles join, and overlap exclusion removes conflicting arms. The additive prediction is probe-regime, because ε conflates composition and regime. Ship a cocktail only on its duel result, since the duel prices ε.


10. Speculative Modules

10.1 MTP

01 extracts the Multi-Token Prediction fusion head (mtp-BF16.gguf) through a structural diff between the full and trunk GGUFs. The standalone head has no token_embd and cannot load alone; the fused deployment (14 --mtp) is the runtime shape. MTP invariance, measured in both directions: PPL and KLD batteries are identical with or without the head attached.

10.2 DSpark / DFlash

Block-diffusion speculative drafts convert whole (tokenizer borrowed from the parent checkpoint). Paretrix-modules.py quantizes them imatrix-free: bridge (fc/eh_proj) ≥ Q6_K · proposal (markov_w/conf_proj) ≥ Q5_K · vocab tables K-quant · unclassified tensors at Q5_K (the F16 catch-all is never used). The trained block size lives in the module metadata (*.block_size); 13 reads it, and mask lengths beyond it were never seen in training.

10.3 MMProj (vision)

CLIP vision towers quantize through a profile grid (balanced / compact / fidelity). Critical tensors (patch and position embeddings, norms, merger) stay pinned at F32, with a deep-boost on trailing blocks. The llama-quantize fallback verifies artifact-vs-plan roles; it caught a historical attn/bridge miss with a divergence of ×1.27.

10.4 Fusion

14_gguf-module-fusion.py rebuilds the deployment GGUF (model + MTP head [+ mmproj]). The first file wins on KV. block_count widens only on the fusion base's key (<arch>.block_count), and per-layer arrays follow the same prefix. A sub-module's count (e.g. clip.vision.block_count) stays untouched. .part output renames on completion.


11. Console Grammar

Every script, every depth:

Form Meaning
══ title ═══ Banner, top-level process
── title ─── Banner, nested process (inherited PARETRIX_DEPTH)
▸ title Section
key value Aligned facts
✓ ok · ✗ failure · ⚠ warning · ⊘ skipped by rule · · note · → command Status line
+ file size Product
⏳ label · 42.0% (n/N) · 3.1 min · ETA 4.2 min One live line per child tool

PARETRIX_VERBOSE=1 restores full tool logs, tensor listings, and static reading aids. The ETA starts at the first counter sample: model loading is excluded from the speed estimate, and the ETA holds steady between advances.


12. Environment Variables

Variable Values Default Effect
PARETRIX_INCLUDE_PICO / _FEMTO 0 / 1 scope 24% / 21% rungs. Default on state-rich ≥3B trunks; =1 forces, =0 suppresses.
PARETRIX_INCLUDE_PRECISION / _FIDELITY 0 / 1 1 42% / 48% rungs.
PARETRIX_CLASSIC 0 / 1 1 Classic stock line on/off.
PARETRIX_CLASSIC_LOW auto / on / off auto Low-band stock quants (Q6_K, Q5_K_M, IQ4_XS, IQ3_M). auto enables them on lever-less topologies only.
PARETRIX_CLASSIC_EXTRA <fmt>[,…] — Extra classic formats.
PARETRIX_QUALITY_ENGINE native / allocate native Quality-36 flat vs DP.
PARETRIX_PRECISION_GATE_Q8 0 / 1 0 GDN sigmoid gates → Q8_0 on Precision.
PARETRIX_LOW_FLAT_MIXER 0 / 1 0 Low-band flat levers on mixer-dominant hybrids.
PARETRIX_FINETUNE 0 / 1 mods L19 field modifier (sells ×0.3, buys ×0.5 on priors). Auto-detected from the model cache's mods.
PARETRIX_FEMTO_RECIPE 0 / 1 0 Femto class-sell recipe instead of the DP (campaign replication arm).
PARETRIX_FEMTO_CLASSIC 0 / 1 0 Femto regex-grid measurement arm.
PARETRIX_TIED_VOCAB 0 / 1 1 R6-ter tied vocab share floors.
PARETRIX_ARCH_FLOORS 0 / 1 1 Arch-authored floors (corrections.floors).
PARETRIX_DEPTH_PRIOR <float> arch line Squeeze depth prior override (G5 test arm).
PARETRIX_SIZE_NORM 0 / 1 1 Size-normalized shared lines (R20).
PARETRIX_PRIOR_SKIP 0 / 1 1 Prior-dead raise skip (R13 census).
PARETRIX_PREFLIGHT 0 / 1 1 Campaign feasibility preflight.
PARETRIX_PRECISION_TWIN 0 / 1 1 When 1, skips Precision on small qwen35 (R19 C2).
PARETRIX_NO_EVAL_CACHE 0 / 1 0 Forces re-measurement in 11.
PARETRIX_VERBOSE 0 / 1 0 Full tool logs and tensor listings.
IMATRIX_SOURCE_QUANT <fmt> / bf16 Q8_0 Source format for imatrix generation.
PARETRIX_PROBE_NGL auto / 0 / <n> auto 12's offload for probe evaluations.
LLAMA_CPP_DIR / LLAMA_<NAME>_PATH <path> — Custom llama.cpp binaries.

13. Retest Matrix

When changing an engine rule, rebuild only the artifacts the changed condition governs. Compare binary hashes against the archived release: byte-identical artifacts need no re-benchmark, and only differing bytes trigger one.

Rule / Change Scope Rungs Validation target
R6-ter tied vocab floors Tied trunks, share ≥ 15% Mini–Femto Xiaomi-OCR-0, NeoHorse-1-4B
R6-bis readout parity Untied, pair > 30% of mass Nano–Quality MiniCPM5-1B/2B, PaddleOCR
R7 GDN flat route SSM hybrids Mini, Compact Ornith-1.5-9B, NeoHorse-1-4B
R9 F32 gate pin SWA + sigmoid gates Mini–Pico Spark-X2.5-4B
R12-bis ride parity Re-shuffles with embd ride all MiniCPM5-2B, nanbeige4.2-3B
R14 size-band identity All shipped products all NeoHorse-1-4B (32.2% out-of-band build → DP rebuild)
R19 stock-twin self re-shuffle Dense topologies with inert levers Quality–Precision TwIL-LM3-Pro (won/tied), MiniCPM5-2B
R20 operating-point pricing Default-path builds all NeoHorse (anchor 'mini'), TwIL (table skipped)
Constructibility gate Floor-limited families sub-Nano Xiaomi-OCR-0 (Mini floor), ReaderLM-v2 (Nano floor)
L11/L20 draft battery DSpark/DFlash modules — MiniCPM5-2B (0.4419), nanbeige4.2-3B

14. Runtime Best Practices

  • Streaming GGUF writes spool tensors through 256 MiB buffers, so peak RAM equals one tensor rather than the whole model.
  • One header parse per file per process (stamp-keyed, released immediately, so Windows rename and delete operations stay safe). Vocab-sized arrays record length only.
  • CPU threads default to physical cores, because SMT siblings contend on shared vector units.
  • NTFS Xpress8K compresses BF16 GGUFs and KL bases (~18%). Quantized payloads and the imatrix stay raw, since their high entropy makes decompression a tax with no gain.
  • Eval cache makes repeated duels and sweeps nearly free (10 min → seconds per model).
  • Probe cache sharing between sibling rungs of the same anchor (anchor-stamp keyed) turns a multi-rung campaign's probe cost into one measurement wave.