test1111111 / docs /VALIDATION.md
spitfire4794's picture
CISM remote autobench: full source + fleet runner, serve results on 7860
28a1a01
|
Raw History Blame Contribute Delete
47.2 kB

Development Snapshot Validation

Date: 2026-09-08. This is not beta certification or a llama.cpp comparison.

Environment

  • Windows x64, AMD Ryzen 5 5600, 6 physical cores / 12 logical processors.
  • CPython 3.13, MSVC Visual Studio Build Tools 2022 17.14.7.
  • One native inference thread, runtime-dispatched AVX2.
  • Installed wheel: cism 0.1.0.dev0.
  • NumPy 2.5.3, Transformers 5.16.1, PyTorch 2.14.0+cpu for reference tests only.
  • FastAPI 0.141.1, Uvicorn 0.52.4, OpenAI Python SDK 2.54.0.

Tests

python -m pytest -q: 555 passed, two upstream TestClient deprecation warnings. Includes independent NumPy numerical references, Transformers tiny model parity, packed quantization, invalid checkpoint/config handling, generation metadata, cancellation, queue limits, SSE, official SDK, and a live Uvicorn TCP roundtrip using real native inference.

python -m pip check: no broken requirements.

Both scalar-only and AVX2 native builds were also tested separately during kernel development. Linux, sanitizers, long-running load tests, and full-context real checkpoint certification remain outstanding.

Real Checkpoints

Compared default FP32 eager HF computation with native last-token logits and four-token greedy continuations at prompt lengths 4, 8, and 16. All tested continuations agreed; maximum absolute logit error was below 0.0001.

  • SupraLabs/Supra-Mini-v5-8M: FP32, INT8, hybrid INT4.
  • veyra-ai/Veyra2-Blueberry-10M-Base: FP32.

INT8/INT4 references used reconstructed quantized weights, not original FP32 weights. These results establish implementation parity only, not a PPL quality budget or task-quality equivalence.

Local Throughput

cism bench SupraLabs/Supra-Mini-v5-8M --precision int8 --runs 5 --max-tokens 128

Model revision: bb98e3566a5ab3f24be16ee8db7f06f0cb884eb5. Packed weight storage including scales/norms: 8,022,784 bytes.

  • Prompt: 6 tokens; generation: 128 tokens; EOS ignored for fixed-length timing.
  • One warmup, then five runs; native generation through coarse Python bindings.
  • Median decode throughput excluding prefill/first token: 1,622.86 tokens/s.
  • Runs: 1,595.72 / 1,632.20 / 1,624.86 / 1,603.32 / 1,622.86 tokens/s.
  • Median prefill plus first token: 3.744 ms.

No tokenization, detokenization, HTTP, or queue latency is included in decode timing. These are growing-context token-generation results, not weight sweeps. CPU affinity and clocks were not locked. No comparison with a tuned llama.cpp baseline has been run; the required speed advantage is not established.

Why These Numbers Differ From The Original Experiments

The earlier session results (hundreds of thousands of tok/s) are not comparable to these, for verifiable reasons on both sides:

  • The supplied C harness (decode_int8/decode_hybrid) looped over flat weight buffers with no layers, no logits, no sampling, and no token feedback. Its "tok/s" measured buffer-processing rate, not decoding.
  • Its Hugging Face baseline ran use_cache=False, which undercounts real HF decode speed; its DDR4 column in the showcase table was computed from 34500 / model_MB, not measured.
  • Kernel improvements have raised this snapshot's decode rate several-fold from the first working build, and it remains compute-bound (see the probe results above), so neither the old synthetic numbers nor the current 1,974 tok/s represent a final ceiling.

Kernel Optimization History

Baseline (per-row indirect dot calls, no FMA, serial INT4 block reduces), Supra-Mini-8M, one thread, 128-token decode: FP32 757, INT8 1,623, hybrid INT4 1,152 tokens/s.

After the first kernel pass (explicit FMA with 4 independent accumulator chains, AVX2 matvec kernels that absorb the row loop so one native call serves a whole matrix, FMA required in the CPU dispatch check):

Precision Kernel-only rate Full-model decode Before Gain
FP32 30.8 GB/s 936 tok/s (31.5 MB weights) 757 +24%
INT8 24.2 GB/s 1,974 tok/s (8.0 MB weights) 1,623 +22%
Hybrid INT4 4.8 GB/s per byte 1,382 tok/s (6.6 MB weights) 1,152 +20%
SmolLM2-135M INT8 - 127 tok/s (135.4 MB weights) 120 +6%

Cache-residency probe (unchanged conclusion): decode effective bandwidth (15.8 GB/s at INT8) remains below even the cold DRAM scan rate (25.1 GB/s), so decode is still compute-bound. The INT4 kernel is the current kernel bottleneck (nibble unpacking without VNNI); full-model non-dot work (attention, norms, 16k-logit sampling, epilogue) accounts for roughly a third of decode time at INT8.

Next levers, by expected impact: multithreaded matvec (row-parallel), prompt-prefill matrix-matrix batching, and further INT4 unpack improvements.

Activation-quantized kernel follow-up (2026-09-12)

The experimental act_precision="int8" path is now implemented for quantized matrix products. It quantizes every activation row in 32-element blocks using 32767 / absmax, uses int16 pmaddwd MACs, and retains FP32 block scales for the epilogue. It is numerically sound but is not enabled by default: on this Ryzen 5 5600 (Zen 3, AVX2/FMA, no VNNI), the quantization work dominates small layer GEMMs. Supra-Mini-8M, one thread, measured about 1,753 tok/s with INT8 weights and FP32 activations versus 603 tok/s with the experimental path; hybrid-INT4 measured 1,297 versus 607 tok/s, and hybrid-FP4 1,238 versus 574. Custom PPL comparisons put the activation-mode delta near +0.01 on Supra-Mini and at or below +0.1 on SmolLM2-135M. Keep it as an explicit experimental API until a fused/coarser quantizer or a CPU with suitable integer-dot instructions shows a net benefit.

Unpack-tax reduction, phase 1 (2026-09-12)

Shipped: skip the nibble interleave (even/odd streams dot a permuted activation scratch built once per call from plain local buffers, no TLS), halving shuffle-port work per 32 weights; dead matvec_* family deleted (~350 lines, verified uncalled); 5-arg scalar adapters keep non-AVX2 builds working. Tried and reverted the same session: folding per-8 scale multiplies into a per-block scalar epilogue — the serial per-block reduce loses to four parallel mulps on Zen 3 out-of-order execution (measured slower, not faster).

Measured on Ryzen 5 5600: full suite green, Mini PPL identical to 3 decimals, production AutoBench once showed fp4-2T beating int4-2T 1254 to 1015 (+24%) then losing the rematch — direction flips are inside this box's ±15-20% multithread noise, so neither is claimed. Drift-controlled A/B (alternating int8/fp4 in one process): 0.87 vs 0.90 pre-change. Verdict: the remaining ~26 µops/32 (converts, scale muls, act loads) are structural on AVX2 without VNNI; 4-bit cannot beat int8 kernel efficiency on this chip. Realistic next levers: int8-container upcast at load (int4 quality, near-int8 speed, int8 size — wins wherever cache residency makes bytes free), the q8 integer-MAC track, stronger ISA (VNNI/AMX/tensor cores), or DRAM-bound regimes where the 2x byte saving dominates.

Fused epilogues tried and reverted (2026-09-12): accumulating o/down dots directly onto x_ and SiLU into the up-gemm epilogue measured ~7% SLOWER on Supra2-Medium int8-1T in a back-to-back stash A/B (580 vs 623, tight clusters, wrong order for a thermal explanation). Likely cause: branches plus a temp in the FMA loop beat clean separate vector passes on wide rows, and an unconditional permute scratch taxed int8/fp32 paths (since made lazy). Lesson: on Zen 3 streaming hardware, an extra cache-resident vector pass can be cheaper than a branch in the hot loop — fuse only with A/B proof.

Session scratch arena (2026-09-12): per-call permute/row mallocs replaced by Session-persistent grow-only buffers (caller-owned, pool-safe; q8 TLS left alone). PPL identical; 4-bit/int8 tok/s ratios up ~0.1 at 2T on both models (Medium int4 0.83 -> 0.98, fp4 0.83 -> 0.96) — consistent with removing malloc-lock contention across pool threads rather than raw bytes. Model now reports info["compiled"] = {key, hit: False, mode: "generic", canonical} as the Phase-0 key plumbing goes live (execution still generic).

Policy: AutoBench now defaults to 512 decode tokens and compiles nothing unless asked (inductor warmup taxed every default run); thread guidance is 2T on a loaded server, never 12T (oversubscription convoy collapses pooled paths while serial paths sail through).

Supra2-Medium-Base production run (2026-09-12, live server, 512 tokens)

First model that spills past L3 (fp32 96.78 MB, DRAM-streaming):

Precision tok/s (2T) PPL vs fp32
int8 (7.65 MB L3) 540 59.40 +0.62%
hybrid-fp4 (20.32 MB L3) 533 60.28 +2.12%
hybrid-int4 (20.89 MB L3) 475 62.61 +6.07%
fp32 (96.78 MB DRAM) 278 59.03 —
HF eager fp32 93 59.03 —
HF dynamic int8 72 95.45 +61.7%

hybrid-fp4 beats hybrid-int4 on speed (+12%) and quality (+2.1% vs +6.1%) where matrices are large enough to amortize the unpack. int8 still leads overall at near-zero quality cost. Speculation accepts at 96-99% and is the largest single lever here (fp32 +65% at spec 8, int4 +17%); HF dynamic int8 is outclassed on both axes. Threads peak at 2-4T; 12T convulses as usual.

Spec K curve (2026-09-12, Medium-Base int8, 2T, 64 tokens)

Cap raised 8 -> 16 to test whether near-100% acceptance favors longer drafts. Measured: K=0: 844, K=2: 974 (84% acc), K=4: 1016 (76%), K=8: 954 (71%), K=12: 915 (69%), K=16: 834 (71%). Peak at K=4: acceptance decays with draft length while block-verify cost grows linearly, so dilution loses past the peak. Optimal K is prompt-dependent (longer, more repetitive generations accept more); the sweep now covers 0-16 with per-thread accept% tables so each setup can read its own peak. Thread sweep preset trimmed to 1,2,4 (above 4T never pays on these sizes).

SmolLM2-135M precision x thread (2026-09-13, single-harness, act fp32, rev 93efa2f0)

Weight bytes (engine.info): int8 135.4 MB (4.2x L3), hybrid-int4 105.1 MB, hybrid-fp4 100.2 MB; all spill, fits_l3=false. 512-token decode, greedy, spec_k=0, one live engine at a time:

Threads int8 tok/s (GB/s) hybrid-int4 tok/s (GB/s) hybrid-fp4 tok/s (GB/s)
1T 88 (11.9) 72 (7.5, 0.82x) 73 (7.3, 0.83x)
2T 120 (16.2) 107 (11.3, 0.89x) 106 (10.6, 0.88x)
4T 124 (16.8) 122 (12.8, 0.98x) 122 (12.2, 0.99x)

PPL (514 scored tokens, window 128): int8 34.74, hybrid-int4 42.07 (+21.1% vs int8), hybrid-fp4 41.32 (+18.9% vs int8; fp4 beats int4 on quality at fewer bytes, speed tied within noise). Central 64-token 1T check: int8 120.9, int4 90.4, fp4 89.0 tok/s with identical weight bytes — same ordering. 128-token bracket int8-1T 114.5 tok/s / 15.5 GB/s matches the old 120-127 table; 512-window context growth costs ~21%, so ratios are intra-run only.

Verdict: int8 is bandwidth-bound (saturates at 2T, +3.6% to 4T), 4-bit paths are kernel-bound at every thread count (1T->4T ~1.7x scaling, DRAM idle, never cash the 1.29x byte edge). Kernel bet unchanged on Zen 3 / no-VNNI: int8 default everywhere, fp4 the 4-bit pick. 360M not fetched (724 MB BF16 shard); expect the amplified same pattern. Next levers: spec on DRAM-bound int8-135M, int8-container upcast at load for cache-resident sizes (test-gated), VNNI-class ISA. Compile Phase-0 stays hit:False/mode:generic (plumbing only, execution generic) — correct until a PPL-gated execution change ships.

Threading Results

Persistent native pool, row-sliced matvecs, pause-spin barriers (yield/sleep only after long idle). Supra-Mini-8M INT8, 256-token decode, medians:

Threads Tokens/s vs 1T
1 1,548 1.00x
2 2,148 1.39x
4 2,036 1.31x (regresses)

FP32 (31.5 MB, DRAM-bound) shows no thread scaling: 998 -> 1,004 tok/s at 4T. The Ryzen 5600 sweet spot is 2 threads for cache-sized quantized models; beyond that, shared L3/LSU bandwidth saturates. gVisor-class sandboxes showed ~19 us sync round trips, so threading there may need larger models to pay off. cism bench/serve --threads N sets the pool size; default remains 1.

Speculative Decoding (Prompt Lookup, Greedy)

Implemented: n-gram drafter over full token history (up to 4-gram tails), block verification via per-matrix GEMM (all T candidates share one weight stream), greedy argmax acceptance with KV rollback, budget/EOS boundaries. Speculative greedy output is bitwise identical to plain greedy decode (tested across architectures, precisions, rejection paths, and budgets).

SmolLM2-135M (135 MB INT8 / 540 MB FP32 - exceeds the 32 MB L3), one thread, 256 tokens, repetitive prompt for high draft acceptance:

Configuration Tokens/s Gain
FP32 plain 51 1.00x
FP32 spec_k=4 83 1.63x
INT8 plain 106 1.00x
INT8 spec_k=4 130 1.23x

Reason: speculation amortizes weight-bandwidth (one stream per verify block), so gains appear when decoding is DRAM-bound - the FP32 regime. The INT8 kernel is compute-bound (conversion-heavy SIMD), so amortizing bandwidth buys little until integer kernels land. On the AVX-512 VNNI cloud target, integer kernels should make INT8 bandwidth-bound and transfer the full speculation multiplier. Two structural costs: the bonus token costs one extra single-token pass (ideal (K+1)/2, measured 1.6x at K=4), and low-acceptance text can make speculation slightly negative - disable it when drafts do not match the domain.

Memory Control Verification

Real-model page locking and touch on this machine (Supra-Mini-8M INT8):

  • lock_pages(): 8,022,880 of 8,022,784 weight bytes locked (page rounding).
  • touch(): 8,022,880 bytes swept.
  • Generation while keep-warm active and pages locked: normal.
  • unlock_pages(): 6,268,768 bytes reported. Smaller than the lock count because adjacent weight regions share 4 KiB pages; a page released by one region cannot be released again by its neighbor. Counts are approximate byte sums of successful OS calls, not exact page accounting.

x86 cannot pin cache lines; locking prevents OS paging, and sweeping only refreshes residency while running. Suite at time of writing: 571 tests passed.

Cache-Residency Probe Results

cism cache-report measures a warm full-weight scan, a cold scan after evicting caches with a 96 MiB scrub buffer, and decode effective bandwidth (weight bytes x decode tok/s). AMD Ryzen 5 5600, INT8, one thread:

Model Weights Warm scan Cold scan Decode Effective Verdict
Supra-Mini-8M 8.0 MB 28.5 GB/s 25.1 GB/s 1,646 tok/s 13.2 GB/s Compute-bound
SmolLM2-135M 135.4 MB 21.4 GB/s 23.1 GB/s 120 tok/s 16.3 GB/s Compute-bound

Interpretation:

  • Decode bandwidth is well below even the cold (post-eviction) scan rate in both cases. The decoder does not saturate DRAM, so cache residency is not the current bottleneck; kernel efficiency is (scalar conversion-heavy quantized kernels, sequential prefill, single thread).
  • The warm/cold scan ratio is small for these linear probes; hardware prefetchers hide much of DRAM latency on sequential streams, so this probe bounds bandwidth rather than proving residency. Definitive per-event counts (LLC miss) require hardware PMU tooling such as WindowsPerf or AMD uProf.
  • The earlier session's implied "SRAM 3x-6.5x multiplier" was not observed by this probe on this hardware and remains unproven.

Speed work priorities implied by this data: quantized kernel quality, prefill matrix-matrix, and multithreading before further cache-tuning.

Measured Comparison Against PyTorch 2.14 And torch.compile

Supra-Mini-v5-8M, revision bb98e356, FP32 checkpoint, greedy, one torch/CISM thread, 6-token prompt, 256-token fixed-length decode (EOS ignored both sides), warmup absorbed, Dynamo counters verified zero recompilations during measurement. Torch baselines use a KV cache (never use_cache=False). Script: scripts/bench_vs_torch.py; CISM: cism bench --max-tokens 256.

Decoder Tokens/s CISM FP32 (998.5) CISM INT8 (1,716.4)
HF generate(use_cache=True) eager 158.9 6.28x 10.80x
Manual eager loop + StaticCache 195.8 5.10x 8.77x
torch.compile default, dynamic=False 260.1 3.84x 6.60x
torch.compile default, dynamic=True 224.9 4.44x 7.63x

The fair comparison is against torch.compile dynamic=False, its best configuration here: 3.84x at matched FP32 precision, 6.60x for CISM INT8.

Reasons, with evidence:

  1. Per-token dispatch. torch.compile lowers eager's per-op Python/ATen dispatch (159 -> 260 tok/s, +64%) but still launches dozens of generated kernels through a Python-level decode loop each token. CISM executes the whole token in one C++ call. Eager's effective bandwidth is 31.5 MB x 159 = 5.0 GB/s and compile's is 8.2 GB/s - both overhead-bound, while CISM FP32 runs memory-bound at ~30 GB/s kernel rate.
  2. Memory traffic. CISM INT8 reads 8.0 MB/token versus 31.5 MB for the FP32 checkpoint, cutting the memory-bound floor ~4x on top of the dispatch elimination.
  3. Quality. CISM FP32 matches eager within 1e-5 max logit error (established above). INT8 quality reporting remains a beta gate.

Caveats: compile mode is "default" (max-autotune not used); one compile configuration initially crashed with an Inductor codegen bug on this model (index-out-of-bounds in a generated kernel) and only ran after cache-reset handling; multithreaded torch was not measured; prefill/TTFT is excluded from decode rates. This is one cache-sized model on one machine - not yet the agreed multi-model llama.cpp gate.

4-bit quantization scheme study (2026-09-12)

NumPy simulation mirroring the native hybrid-int4 quantizer (round-half-away, blocks of 32 along the row, per-block scale; validated against the native engine to 8 decimal places on all checkpoints). PPL over the fixed AutoBench corpus (514 tokens, window 128), MLP matrices 4-bit, all other 2D matrices per-row INT8, vectors FP32.

Machine: Ryzen 5 5600, local checkpoints, scripts/quant4_study.py.

scheme (MLP element + scale) bits/w 8M ΔPPL 100M ΔPPL 135M ΔPPL
int4 + FP32 b32 (hybrid-int4) 5.00 +15.4% +6.1% +23.2%
int4 + E4M3 b32 4.25 +15.5% +6.0% +26.7%
int4 + E8M0 b32 4.25 +308% +80.8% +537%
E2M1 + E8M0 b32 (OCP MXFP4) 4.25 +318% +92.4% +323%
E2M1 + E4M3 b32 4.25 +18.7% +3.5% +23.3%
E2M1 + E4M3 b16 (NVFP4 granularity) 4.50 +15.0% +2.1% +21.0%
int4 + E4M3 b16 4.50 n/a +8.1% n/a
E2M1 + FP32 b32 5.00 +13.7% +4.4% +19.9%
NF4 + FP32 b32 (QLoRA) 5.00 +44.1% +8.8% +32.0%
int4 + error feedback b32 5.00 +36.5% +16.9% +80.9%

Findings:

  • Scale precision dominates element format; power-of-two (E8M0) scales are catastrophic for both uniform and FP elements.
  • E2M1 elements beat uniform INT4 at equal scale precision and block size (100M model, b16+e4m3: +2.1% vs +8.1%).
  • E2M1 + E4M3 b16 is the only scheme that improves on hybrid-int4 at every model size while using fewer bits (4.50 vs 5.00): tied on 8M, 3x better on 100M, better on 135M.
  • NF4 (QLoRA) loses badly on these trained small models; naive per-block error feedback with a fixed scale inflates residuals and hurts everywhere.

Why hybrid-int4 decodes slower than int8 (measured, 1 thread, same checkpoints): the nibble unpack is a ~10-instruction dependent chain per block, so hybrid-int4 delivers only 7.6-8.1 GB/s of weight-equivalents vs 12.5-14.0 for int8 and 26.6 for fp32, despite moving 2.5x fewer bytes than fp32. The E2M1 LUT kernel replaces that chain with a 16-entry table lookup and is the reason to build the mode.

Compile one-shot canonical parity (2026-09-14)

Phase-0 key plumbing is now cross-language exact. Python compile_cache.canonical_model_config() builds the native 12-field pipe model_type|hidden|intermediate|layers|heads|kv_heads|head_dim|vocab| context|hexfloat(eps)|hexfloat(rope_theta)|tied with FP32-round-trip struct.pack('f') hexfloat, matching Model::refresh_compile_key bit-for-bit. Verified: tiny Llama 16/32/1/4/2/4/32/32 probe Python canonical equals Model.info["compiled"]["canonical"], and native.compile_key(python_canonical, precision, act, threads) equals Model.info["compiled"]["key"] (d4547092c73a7aac on the probe). Legacy canonical_config() sorted-JSON remains for early tests only. Kept as a stable model fingerprint (zero runtime cost, tested); the packed-artifact load idea is dropped with the project (see below).

SmolLM2-135M full AutoBench matrix (2026-09-13, 512 tokens, act fp32)

Local file autobench_135M_512.json (untracked measurement, torch baselines disabled by policy include_compile=False): 4 precisions x threads (1,2,4) x spec_k (0,2,4,8,12,16) = 72 decode cells, PPL once per precision (514 scored tokens, window 128). Revision 93efa2f0 (report does not yet stamp revision; next run should copy engine.info["revision"]). All fits_l3=false (L3 cache size 32.0 MB).

Weight storage in megabytes: int8 129.16 MB (135,439,104 bytes), hybrid-int4 100.27 MB (105,141,504 bytes), hybrid-fp4 95.52 MB (100,164,864 bytes), fp32 513.13 MB.

Decode rate in tokens/s at spec_k=0 (effective bandwidth in GB/s = weight bytes x tok/s):

Precision 1 thread tok/s (GB/s) 2 threads tok/s (GB/s) 4 threads tok/s (GB/s) Best tok/s (threads, spec_k) PPL (+vs int8)
fp32 43.78 (23.6) 57.44 (30.9) 55.34 (29.8) 106.41 (4T, K=8) 34.15
int8 91.53 (12.4) 124.82 (16.9) 130.03 (17.6) 147.19 (4T, K=2) 34.74 (--)
hybrid-int4 71.93 (7.6) 106.54 (11.2) 115.69 (12.2) 140.59 (4T, K=16) 42.07 (+21.1%)
hybrid-fp4 71.85 (7.2) 102.10 (10.2) 117.51 (11.8) 132.90 (4T, K=2) 41.32 (+18.9%)

Ratios int8 vs 4-bit at spec 0: 1T 1.27x, 2T 1.17-1.22x, 4T 1.11-1.12x. Scaling 1T to 4T: int8 1.42x but 2T to 4T only +4.2% (saturates, bandwidth-bound); int4 1.61x, fp4 1.64x, still climbing with DRAM idle (kernel-bound nibble unpack, no VNNI). Verdict unchanged: int8 default everywhere, fp4 the 4-bit pick (better quality at fewer bytes, speed tied within ±15-20% multithread noise; ratios single-harness only).

Spec K-curve on 135M DRAM-bound int8 (512 tokens, greedy)

Accept rate in percent is deterministic across threads per K (88.0% at K=2 decaying to 72.4% at K=16).

Threads K=0 tok/s K=2 tok/s K=4 tok/s K=8 tok/s K=12 tok/s K=16 tok/s Peak gain vs K=0
1T 91.53 98.42 100.47 99.45 96.14 93.67 K=4, +9.8%
2T 124.82 136.05 132.62 131.34 129.72 126.94 K=2, +9.0%
4T 130.03 147.19 146.55 135.13 136.81 143.53 K=2, +13.2%

Global best: int8 147.19 tok/s at 4 threads, spec_k=2. Contrast fp32 +73-92% peaking K=8 (bandwidth-bound, block verify amortizes DRAM) vs 4-bit +5-21% erratic (kernel-bound). Peak at small K because acceptance decays with draft length while GEMM verify cost grows linearly, plus the bonus-token extra pass. Prompt-lookup drafter (up to 4-gram, greedy-only, bitwise identical) helps repetitive prompts only; it does not fix the 135M slow complaint (still ~147 tok/s vs ~1,600 tok/s cache-resident 8M). Default 135M int8: 4 threads, spec_k=2 (2-thread servers also K=2).

trust_remote_code custom-arch assessment (2026-09-14)

trust_remote_code=True authorizes tokenizer Python only, never custom modeling operators; trusted code itself is not sandboxed. Enforcement: loader _ARCHITECTURES={"llama","qwen3"}, validate_config rejects non-llama/qwen3 model_type and any auto_map beyond AutoTokenizer even with trust enabled, cross-repo refs rejected, _expected_shapes plus _load_weights reject extra/missing tensors, native Config::validate/parse_config reject bias/MoE/sliding-window/ non-silu/non-default-RoPE. Custom Negative/CMA/Ember2/Spark2A/Blaze targets (ROADMAP gate 4: recurrence, routing, lane mixing, Engram state) need new native Config fields, operators, KV layouts, quant definitions, and PPL/parity gates per arch — large work, not a flag flip. Keep explicit UnsupportedModelError; server stays opt-in --allow-remote-code-load (403 otherwise).

Compile project dropped (2026-09-14, verdict: correctly built, not worth it)

Target was one-time load cost for faster batch-1 CPU decode (generic AVX2, PPL-gated). Verdict: the plumbing is correct and tested, but the goal is unreachable on this CPU, so the project is dropped and its speedup fiction removed:

  • What was real: cross-language canonical parity (bit-exact, tested), hoisted dispatch loops (same kernels/order, bitwise identical, kept as plain code), manifest helpers (kept, unused by load path).
  • What was fiction: manifest-hit as "compile hit" (bookkeeping, not a faster layout), compiled.hit/mode in Engine info and tester cards (removed), any decode tok/s attributed to compile (zero: mode was generic everywhere; the +13.2% is spec K=2).
  • Why not worth it: remaining ~26 µops/32 are structural on AVX2/no-VNNI; int8 is bandwidth-bound (2T saturates), 4-bit kernel-bound; fresh 512 runs show hoisting moved 4-bit ~+23% but int8 K0 flat (-0.9%), confounded by idle-machine effect — no clean decode win to claim; inductor costs 107 s warmup for 32 tok/s against CISM's 190 (6-8x lead from the pre-existing single-call C++ path + quant, not compile). VNNI/VBMI or integer-MAC kernels would reopen the question; until then the line is closed. Fingerprint key + hoisted loops stay; hit/UI/ future-speedup prose goes.

Fresh 512-token 4T confirmation (SmolLM2-135M, all precisions, spec sweep)

Re-ran 24 cells (4 precisions x 4T x spec 0,2,4,8,12,16), 512 tokens, 127.3 s wall, torch off. Accept% deterministic and identical to the old file; PPL/weights identical (fp32 34.15, int8 34.74, int4 42.07, fp4 41.32; 513.13/129.16/100.27/95.52 MB, all fits_l3=false). Absolute tok/s ran +11-30% hotter than the old file (idle machine); cells >25% off are marked SUSPECT — use this table, do not mix files.

Precision K=0 K=2 K=4 K=8 K=12 K=16 Best
fp32 58.77 72.89 101.92 125.09 122.38 121.18 K=8 125.09
int8 128.90 159.99 162.64 171.76 171.37 177.03 (SUSPECT vs old) K=16 177.03
hybrid-int4 142.63 155.69 151.97 159.86 160.29 156.40 K=12 160.29
hybrid-fp4 144.58 150.73 153.96 160.41 155.88 153.82 K=8 160.41

Accept%: fp32 96.5/94.0/89.9/88.1/85.5; int8 88.0/83.4/77.1/73.0/72.4; int4 90.1/84.4/79.1/75.6/75.3; fp4 84.1/76.6/70.3/66.7/67.6 (K=2..16). Note 4-bit now leads int8 at K=0 (142.63/144.58 vs 128.90) in this run — confounded by machine state (old file had int8 leading), so no compile attribution claimed without a same-state A/B. fp4 vs int4: tied within noise at every K (max gap 3.7%), fp4 keeps the quality edge (+18.9% vs +21.1% PPL), so fp4 stays the 4-bit pick.

4T K=2 vs regular baselines (SmolLM2-135M int8, 512 tokens)

Global best cell is int8 4 threads spec_k=2 at 147.19 tok/s (accept 87.96%). Regular baselines from the same autobench_135M_512.json run (include_torch=False, so no HF eager/compile in this file):

Baseline (4T) Tok/s vs 4T K=2 (147.19)
int8 K=0 (same weights, no spec) 130.03 +13.2% with spec
int8 2T K=0 124.82 +17.9% (threads+spec)
int8 1T K=0 91.53 +60.8%
fp32 K=0 55.34 2.66x
fp32 best (K=8) 106.41 1.38x
hybrid-int4 K=0 115.69 1.27x
hybrid-int4 best (K=16) 140.59 1.05x
hybrid-fp4 K=0 117.51 1.25x
hybrid-fp4 best (K=2) 132.90 1.11x

PPL in same run: int8 34.74, fp32 34.15, int4 42.07, fp4 41.32. Spec pays (+13.2%) because 135M int8 spills (129.16 MB vs 32 MB L3): block verify streams weights once per block via GEMM. It does not make 135M fast in absolute terms (still ~147 vs ~1,600 cache-resident 8M).

HF eager/compile baselines (SmolLM2-135M, 512 tokens, 4T, same harness)

One run_benchmark (src/cism/autobench.py:223) with include_torch=True, include_compile=True, max_new_tokens=512, threads 4, spec sweep 0-16, rev 2c4d40c. Whole sweep ran ~4-15% under the fresh reference (mild thermal throttle); HF rows are baselines, not throttle-sensitive conclusions.

Decoder (512 tok, 4T) Tok/s PPL (514 scored tokens)
CISM int8 K=0 131.88 34.74 (nll 3.5480, +1.73% vs eager)
CISM int8 best K=16 167.23 same
CISM fp32 K=0 / best K=8 53.32 / 114.67 34.15 (+0.002%)
CISM hybrid-int4 K=0 / best K=8 137.20 / 153.07 42.07 (+23.18%)
CISM hybrid-fp4 K=0 / best K=16 137.38 / 150.25 41.32 (+21.00%)
HF eager fp32 23.79 (21.52 s) 34.15 (nll 3.5308)
HF dynamic int8 19.74 (25.94 s) 81.12 (broken, +137.5%)
HF torch.compile fp32 25.18 (20.33 s, warmup 147.3 s) not scored
HF compile int8 16.86 (30.38 s) not scored
torchao int4 skipped (not installed) --

Speedups (same harness): CISM int8 K=0 5.54x eager / 5.24x compile; best K=16 7.03x / 6.64x; fp32 K=0 2.24x/2.12x, best K=8 4.82x/4.55x. Compile barely beats eager (+5.9%) after a 147 s warmup; dynamic int8 is slower than fp32 AND destroyed (PPL 81). Weights: 513.13/129.16/100.27/ 95.52 MB, all fits_l3=false (L3 32 MB).

Surjo hybrid native Phase1 (loader+fp32 decode)

Scope: SurjoLabs Surjo hybrid import plus native fp32 decode runs. No Surjo tok/s, no Surjo PPL, no speedup claim in this section. Reference dense numbers below are context only, single-harness, ±15-20% multithread noise (see 135M matrix notes); do not mix files or attribute cross-run deltas to Surjo.

Config fingerprint (Surjo-50m, tests/test_surjo.py:20): model_type surjo / SurjoForCausalLM, vocab 32768, hidden 512, intermediate 1536, layers 10, context 2048, prelude 1 / recurrent 8 / coda 1, groups 2 / passes 2 / gdn_per_xsa 3, heads 8 / kv 4 / head_dim 64, gdn_v_heads 8 / gdn_k_dim 64 / gdn_v_dim 64 / kernel 4, gdn_allow_neg_eigval false, xsa_projection true, bias false, dropout 0.0, hidden_act silu, eps 1e-5, theta 10000.0, default RoPE (partial_rotary_factor 1.0), tied true, layer_types [F,L,L,L,F,L,L,L,F,F], auto_map pinned to configuration_surjo/modeling_surjo (never executed). Base checkpoints ship use_cache=false; import_model (src/cism/loader.py:666) preserves the flag verbatim and it does not disable native caching (XSA KV slots + GDN recurrent/conv state always allocate).

Checkpoint header: 178xF32 tensors in model.safetensors, tied lm_head omitted, no rotary inv_freq buffers (RoPE is computed in forward()). _expected_shapes_surjo length is 179 including the tied-head alias entry; spot-checked in test_surjo_50m_expected_shapes_match_header (embed (32768,512), XSA q (512,512) / k (256,512) / q_norm (64,), GDN q/k (512,512) / v (512,512) / conv (512,1,4) / f0 (64,512) / f1 (512,64) / A_log (8,) / dt_bias (512,) / g1.bias (512,) / o_norm (64,)).

Loader (src/cism/loader.py): validate_surjo_config, surjo_exec_plan (18 steps: 1 prelude + 2 passes x 2 groups x (3 GDN + 1 XSA) + 1 coda; XSA slots [0..5], 12 GDN (layer,pass) steps), surjo_num_xsa_slots (=6), _expected_shapes_surjo:393, _load_weights Surjo branch (no buffers, strict shape/dtype/finite checks, tied-head alias by identity), import_model dispatch (model_type=="surjo" -> Surjo path, else dense). _SURJO_SUPPORT string still reads "next phase" in this snapshot; Phase1 below is what landed natively.

Native (native/runtime.{hpp,cpp}, native/bindings.cpp:468, src/cism/engine.py:80 dispatch to _native.SurjoModel): SurjoModel/SurjoSession fp32 decode. XSA: 6 slots, KV [slots,capacity,kv_width] with kv_width=256; at max context 2048 that is 4 MiB/slot (K+V), 24 MiB total (session-sized by prompt+max_new, so 24 MiB is the upper bound). GDN: 12 states (6 GDN layers x 2 passes), each S is H*K*V=8*64*64=32768 floats = 128 KiB (12x128 KiB = 1.5 MiB), plus depthwise-conv raw FIFO (kernel-1)=3: q/k/v FIFOs ~72 KiB each, ~0.21 MiB total. MLP is ClampedMLP (gate clamped to [-15,15] then silu*up, runtime.cpp:2010). spec_k is rejected for Surjo (next_tokens throws unless spec_k==0; Engine.stream(spec_k=2) raises ValueError, covered by test_surjo_engine_decode_runs). Prefill is sequential per-token (forward_tokens loop); logits / nll teacher-forcing paths exist but are unscored against HF so far.

Quant mapping reuse, no new format: protected_storage (embed, head, all XSA/GDN projections incl. f0/f1/g0/g1/o_proj) is fp32 or int8; MLP gate/up/down follow the dense hybrid rule (hybrid-int4 -> int4, hybrid-fp4 -> fp4). Same Matrix kernels/scratch/act_q8 path as dense; only norms/vectors/conv/A_log/dt_bias stay fp32.

Parity gate: dense bar is max abs logit error <1e-4 vs HF eager. Surjo fp32 decode/smoke runs (tiny shapes fixture, tied-head identity, Engine.logits/generate finite, spec rejection). HF modeling_surjo.py parity script and Surjo PPL are pending — no parity number claimed here.

Batch/paging: Phase0 files unchanged. Single-request serialized Engine generation lock, create_session/next_tokens, cancel/finish, lock_pages/touch/scan weight-memory controls are reused as-is. No continuous batching, no paged-KV layout, no new server path for Surjo.

FLA CPU verdict: decode needs no Flash-Linear-Attention (single-token GDN recurrence is O(1) state + per-slot causal XSA). Chunkwise/FLA prefill is an optional long-prompt throughput optimization only, not correctness; current sequential prefill stands.

Dense context (same autobench_135M_512.json, include_torch=False, NOT Surjo results): 4T K=0 fp32 55.34 / int8 130.03 / hybrid-int4 115.69 / hybrid-fp4 117.51 tok/s; best cells fp32 106.41 (K=8), int8 147.19 (K=2, accept 87.96%), hybrid-int4 140.59 (K=16), hybrid-fp4 132.90 (K=2); PPL fp32 34.15 / int8 34.74 / int4 42.07 / fp4 41.32. Ratios single-harness only; 4-bit vs int8 gaps sit inside ±15-20% noise. Surjo inherits none of these numbers.

Surjo-50m full AutoBench (Phase2, 2026-09-15, 4T K0)

File autobench_surjo50m_512.json (local snapshot 35c9fa6e, 512-token HF window, CISM greedy K0 only, act fp32, threads 1/2/4). Weights (engine.info): fp32 205.21 MB, int8 51.84 MB, hybrid-int4 43.27 MB, hybrid-fp4 41.86 MB; L3 32.0 MB, all fits_l3=false. PPL over the fixed AutoBench corpus (527 scored tokens): fp32 59.29 / int8 59.46 (+0.30% vs HF eager) / hybrid-int4 62.90 (+6.10%) / hybrid-fp4 61.80 (+4.25%). CISM int8/fp32 PPL matches HF eager fp32 59.28 to 0.3% — quant quality gate holds at int8, 4-bit pays +4-6%.

CISM decode at 4T K0 (greedy; CISM stopped on EOS at 43-48 tokens, HF forced 512 via min_new_tokens — lengths differ, ratios indicative):

Precision 1T tok/s 2T tok/s 4T tok/s PPL
fp32 62.18 81.03 81.06 59.29
int8 129.86 158.39 207.80 59.46
hybrid-int4 109.05 159.04 175.49 62.90
hybrid-fp4 105.29 153.13 171.31 61.80

HF baselines (same harness, FP32 CPU, same thread count):

Decoder (4T) Tok/s PPL
HF eager fp32 26.27 (1T 28.08 / 2T 27.74 — inverse scaling, MT noise) 59.28
HF dynamic int8 18.18 (main cell 19.93; sweep 17.42/18.95/18.18) 66.35 destroyed (+11.9%)
HF torch.compile fp32 31.72 INVALID — StaticCache path, see below not scored
HF compile int8 16.73 (same StaticCache invalidity) not scored

Speedups at 4T K0 vs HF eager 26.27: int8 7.91x, hybrid-int4 6.68x, hybrid-fp4 6.52x, fp32 3.09x. speedup_vs_compile in the JSON divides by the invalid 31.72 and is void — do not cite.

Compile correction: modeling_surjo.py:645 discards non-SurjoCache past (if not isinstance(past, SurjoCache): past = None), so the autobench StaticCache loop allocated a fresh cache per forward call and state never carried — 31.72 tok/s is not a Surjo compile number. src/cism/autobench.py now detects surjo_mode and threads returned past_key_values (SurjoCache) under torch.compile(dynamic=True), recording {"skipped": "StaticCache invalid for Surjo"} if the threaded path cannot run — never a StaticCache tok/s. Correct path proven by scripts/bench_compiler.py: threaded SurjoCache + dynamic=True 41 tok/s (compile warmup 43-119 s across runs; 189.3 s in the old JSON was the invalid-path warmup); dynamic=False unusable (0.06 tok/s, recompiles per token as XSA KV length grows). torchao int4 blocked on Windows (no mslk wheel — Requires mslk >= 1.0.0); torchao int8wo is the CPU-working control only. FLA CPU dead here (Triton needs GPU; fla.ops.gdn2 chunk/fused probe does not run on CPU) — decode needs no FLA anyway (single-token GDN recurrence is O(1) state + per-slot causal XSA).

Spec: unsupported for Surjo hybrid (GDN state destructive) — K0 only (spec_note in JSON); K2≈K0 where probed, no speculation gain claimed. Native SurjoSession rejects spec_k!=0; Engine.stream(spec_k=2) raises.

Rigor: this JSON predates the harness (matrix rows lack rep_rates/thermal_throttled/cpu_snapshot). Current src/cism/autobench.py:_timed reports 1 warmup + median-of-reps (--reps/--runs wired through cism bench; default 3) with min/max/ std, thermal_throttled when max-min spread exceeds 15% (wider than the ±15-20% MT noise band), and best-effort cpu_snapshot (clock/temp; {} when unavailable, never fails the run). Next Surjo rerun picks up medians + thermal flags automatically; do not mix old single-sample cells with new median cells.

Surjo-50m rerun on current code (2026-09-15 night, 72 cells, reps=3)

autobench_surjo50m_512.json overwritten with current native (GDN-OPT + XSA-blocked + forward_block + spec verify). Threads 1/2/4 x spec 0/2/4/8/12/16, 512tok, act fp32, median-of-3, all thermal_throttled=false except two fp32-4T spec cells. CISM EOS-stops 43-48tok vs HF forced-512 — same caveat as before.

K0 medians: fp32 60.32/84.73/88.86, int8 141.76/197.34/234.87, int4 114.59/169.65/197.28, fp4 113.54/169.98/214.61. Best cells: fp32 K2 92.66, int8 K4 236.91, int4 K4 201.33, fp4 K0 214.61. Spec accept on bench prompt is 0.0 for fp32/int8 (no gain, no loss beyond noise); int4 6.25%->1.7%, fp4 8.3%->2.7% decaying with K, and spec K>0 is slower than K0 there (verify cost, no amortization win on this prompt). PPL unchanged: 59.29/59.46/62.90/61.80.

HF equal-opportunity, same harness, threads swept, median-of-3: eager 31.22/30.43/29.50 (4T 29.50 PPL 59.28), dynamic-int8 20.27/22.59/22.37 (4T 22.69 PPL 66.35), compile threaded-SurjoCache dynamic=True 38.77 warmup 100.6s, compile-int8 21.10, torchao-int4 failed Requires mslk>=1.0.0. Speedups vs eager 29.50: int8 K0 7.96x, best K4 8.03x; vs compile 38.77: K0 6.06x. Old invalid StaticCache 31.72 is superseded — do not cite it.

Supra-50M-Base full AutoBench (dense Llama, 2026-09-15 night)

File autobench_supra50M_512.json, Hub rev 521bfd3d3901fbeefa943a2f3d461a9d193c52ec, dense Llama (hidden 512, layers 12, heads 8/kv 4/head_dim 64, intermediate 1408, vocab 32000, ctx 1024, tied, use_cache=false preserved). Threads 1/2/4 x spec 0/2/4/8/12/16 = 72 cells, 512tok, act fp32, median-of-3. Weights: fp32 ~209MB / int8 ~53MB class (all fits_l3=false). PPL (527 scored tokens): fp32 63.32 / int8 63.79 (+0.7%) / int4 65.36 (+3.2%) / fp4 68.90 (+8.8%).

CISM K0 medians: fp32 108.29/131.54/123.55, int8 214.81/284.36/292.58, int4 178.04/247.49/250.87, fp4 171.67/234.54/251.38. Best cells: fp32 K8 268.19 (2T), int8 K12 349.10 (4T), int4 K2 257.43, fp4 K8 274.25. Accept on bench prompt is 100% for fp32/int8/fp4 at every K (hence spec pays: int8 4T K0 292.58 -> K12 349.10, +19.3%; fp32 K0 123.55 -> K8 268.19, +117%); int4 86%->70%.

HF same harness, threads swept, median-of-3: eager 48.59/47.90/43.61 (4T 43.61 PPL 63.32), dynamic-int8 37.70/40.49/37.78 (4T 34.91 PPL 101.68 destroyed), compile StaticCache dynamic=False 55.23 warmup 143.0s, compile-int8 33.14, torchao-int4 Requires mslk>=1.0.0. Speedups vs eager 43.61: int8 K0 6.71x, best K12 8.00x; vs compile 55.23: K0 5.30x, best 6.32x.

FWKV Myosotis-1-base full support + AutoBench (2026-09-15 night)

Files inspected in snapshot 031fe84a34f667bc8bfff7dc3f8a77ef3c17efb2: config.json (model_type fwkv, d_model 768, d_emb 192, n_layers 13, ffn_mult 4, vocab 50257, wkv_floor 0.1, tied), configuration_fwkv.py (FWKVConfig + HF aliases), modeling_fwkv.py (FactorizedTiedHead, FWKVBlock with proj_k/v/r/out + W recurrence + GELU FFN + 2x LayerNorm, _supports_cache_class=False, tuple-state threading), model.safetensors (147 F32 tensors, no rotary buffers).

Native support landed, mirroring the Surjo vertical slice: loader.validate_fwkv_config + _expected_shapes_fwkv (147 shapes) + _load_weights fwkv branch (also fixed an eager config["head_dim"] default crash affecting any arch without head_dim) + import_model dispatch; native FwkvConfig / FwkvModel / FwkvSession (WKV recurrence, exact GELU, LayerNorm eps 1e-5, factorized head, fp32/int8/int4/fp4 reuse via Matrix, forward_block batched projections + sequential recurrence, spec verify with ~40KB state checkpoint/restore); bindings FwkvModel/FwkvSession; engine dispatch; tests/test_fwkv.py 11 passed; scripts/parity_fwkv.py: tiny 1.19e-07, live vs HF fp32 5.72e-06 (<1e-4 OK). Autobench: trust_remote_code=True for fwkv HF loads, threaded-FWKV-state compile path (StaticCache invalid, same class of bug as Surjo), _hf_context fallback to seq_len (FWKVConfig has no max_position_embeddings).

File autobench_fwkv_512.json, 72 cells, reps=3 (v4 drafter rerun, 2026-09-15 late night; supersedes the v1-drafter run below). Weights: fp32 425.38MB / int8 107.64MB. PPL: fp32 156.68 (= HF eager 156.68, gate holds) / int8 156.80 (+0.07%) / int4 156.63 (-0.04%) / fp4 158.53 (+1.18%). High absolute PPL is base model quality, faithfully reproduced.

CISM K0: fp32 64.65/93.73/87.24, int8 169.01/268.22/342.27, int4 121.54/215.32/284.66, fp4 119.21/205.16/303.52. Best: fp32 K16 177.34, int8 K4 382.74, int4 K2 290.19, fp4 K2 305.98. Accept (bench prompt): fp32 85%->61%, int8 85%->62%, int4 77%->45%, fp4 74%->44% (K2->K16). Spec pays on fp32/int8: int8 4T K0 342.27 -> K4 382.74 (+11.8%); fp32 K0 87.24 -> K16 177.34 (+103%); 4-bit int4 still peaks near K0-K2, fp4 peaks K2.

HF same harness, threads swept, median-of-3: eager 42.72/45.59/43.42 (PPL 156.68), blocks-only dynamic-int8 34.74/43.88/51.26 (4T 51.22 PPL 158.42 +1.1% ok — fixed by quantizing blocks Linears only; full-model quantize fails on FactorizedTiedHead.proj.weight.t()), compile threaded-FWKV-state dynamic=True 49.73 warmup 45.9s, compile-int8 same path 46.93 warmup 36.2s, torchao-int4 Requires mslk>=1.0.0 (no Windows wheel). Speedups vs eager 43.42: int8 K0 7.88x, best K4 8.81x; vs compile 49.73: K0 6.88x, best 7.70x. Previous run (v1 drafter): K0 59.82/83.70/87.54 etc., best int8 K8 359.97 — v4 gains come from restored accept (19%->85% @K2).

Spec drafter v4 + doom-loop check (2026-09-15 late night)

v1 (most-recent single match, verbatim tail truncated at history end): Surjo bench 0% (correct text, nothing to match), Supra bench 100% (degenerate loop flatters it). v2 (pure frequency vote) collapsed on drifting text (FWKV bench int8 K2 85%->19%, K8 69%->9%). v3 (recency-decayed vote) same failure. v4 (landed): most-recent strictly-inside occurrence per position + iterative extension (fixes v1 truncation) + rejected-block latch REMOVED (the 32-propose/<1/3-accept latch fired during the short-history warmup and blinded the session to later 85%-accept regions; rejected blocks stream weights once per block via forward_block, so worst case on 0%-accept text is ~K0 speed). Shared by dense/Surjo/FWKV; tests/test_native.py::test_speculative_draft_hits_and_misses reimplemented against v4; suite 759 passed; K0==Kx exact everywhere probed.

Doom-loop probe (greedy, 128 tok, fp32+int8, 3 prompts): Surjo stops coherently on bench/code prompts (48/15 tok, uniq 0.73/0.93, loop 0) but loops story prompts (128 tok, uniq 0.09-0.16, 5/11-period loops); Supra loops ALL prompts (128 tok, uniq 0.07-0.11, 7/9/12-period loops — bench, story, AND code); FWKV never hard-loops bench/story (uniq 0.28-0.54, loop 0; int8 bench drifts into a 10-period body/brain cycle) but degenerates to 128 spaces on code (uniq 0.02 both precisions). So Supra's 100% spec accept is a looping artifact, Surjo's 0% is healthy text, FWKV sits between. No repetition penalty exists anywhere (greedy only); spec accept must always be read next to uniqueness.

Portable kernels: Intel/AMD tiers + ARM64 phones (2026-09-16)

Tier ladder (runtime-dispatched, scalar fallback always works): scalar (any CPU) < avx2+FMA (x86-64 baseline, all current numbers) < avx-vnni (Zen4/Zen5, Alder Lake+: int8 VNNI fast path) < avx512-vnni (server AVX512) / neon (ARM64 phones, Apple Silicon, Raspberry Pi). New files: native/kernels_vnni.cpp (dpbusd int8xint8, int8-activation ABI via quantize_row_i8), native/kernels_neon.cpp (fp32/int8/int4/fp4 + uzp permute), CPUID detection in native/kernels.cpp (AVX-VNNI = leaf7:1 EAX[4], AVX512-VNNI = leaf7:0 EBX[11] + foundation + ZMM state; MSVC VNNI TU is EVEX so it dispatches only on AVX512-VNNI, GCC VEX on AVX-VNNI), per-file codegen flags in CMakeLists (dispatcher/runtime stay generic). kernel_name() keeps the fp32 family; new kernel_variant() + has_avx_vnni / has_avx512_vnni / has_neon bindings; cpu_features gains avx_vnni; Model.info["kernels"] gains variant. tests/test_kernels_portable.py (5 passed): tier reporting, flag consistency, pre-VNNI fallback untouched, i8 quantizer round-trip bounds, dpbusd bias-identity in numpy.

Proven on Ryzen 5 5600: builds clean (MSVC /arch:AVX512 VNNI TU accepted), kernel avx2 / variant avx2 / vnni False, suite 764 passed (759 + 5 new) — zero behavior change, as designed (VNNI selector is null here so Matrix keeps the int16 Q8 / fp32-activation paths byte-for-byte). Pending hardware CI: VNNI tok/s + int8-activation PPL re-gate on Zen4/ADL+ (expect prefill/spec-verify wins ~2-3x integer MACs; decode stays bandwidth-bound so batch-1 gains concentrate in K>0), NEON tok/s + parity on ARM64 hardware (expects the usual Silicon/Android numbers; tails fall back to scalar by construction). No speedup is claimed for either until measured.