Spaces:
Sleeping
Download docs/VALIDATION.md from spitfire4794/test1111111: direct link, hf CLI and curl.
- Browser
- Download file 47.2 kB
-
https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/VALIDATION.md
- Command line
-
hf download hf://spaces/spitfire4794/test1111111/docs/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/VALIDATION.md
Development Snapshot Validation
Date: 2026-09-08. This is not beta certification or a llama.cpp comparison.
Environment
- Windows x64, AMD Ryzen 5 5600, 6 physical cores / 12 logical processors.
- CPython 3.13, MSVC Visual Studio Build Tools 2022 17.14.7.
- One native inference thread, runtime-dispatched AVX2.
- Installed wheel: cism 0.1.0.dev0.
- NumPy 2.5.3, Transformers 5.16.1, PyTorch 2.14.0+cpu for reference tests only.
- FastAPI 0.141.1, Uvicorn 0.52.4, OpenAI Python SDK 2.54.0.
Tests
python -m pytest -q: 555 passed, two upstream TestClient deprecation
warnings. Includes independent NumPy numerical references, Transformers tiny
model parity, packed quantization, invalid checkpoint/config handling, generation
metadata, cancellation, queue limits, SSE, official SDK, and a live Uvicorn TCP
roundtrip using real native inference.
python -m pip check: no broken requirements.
Both scalar-only and AVX2 native builds were also tested separately during kernel development. Linux, sanitizers, long-running load tests, and full-context real checkpoint certification remain outstanding.
Real Checkpoints
Compared default FP32 eager HF computation with native last-token logits and four-token greedy continuations at prompt lengths 4, 8, and 16. All tested continuations agreed; maximum absolute logit error was below 0.0001.
- SupraLabs/Supra-Mini-v5-8M: FP32, INT8, hybrid INT4.
- veyra-ai/Veyra2-Blueberry-10M-Base: FP32.
INT8/INT4 references used reconstructed quantized weights, not original FP32 weights. These results establish implementation parity only, not a PPL quality budget or task-quality equivalence.
Local Throughput
cism bench SupraLabs/Supra-Mini-v5-8M --precision int8 --runs 5 --max-tokens 128
Model revision: bb98e3566a5ab3f24be16ee8db7f06f0cb884eb5.
Packed weight storage including scales/norms: 8,022,784 bytes.
- Prompt: 6 tokens; generation: 128 tokens; EOS ignored for fixed-length timing.
- One warmup, then five runs; native generation through coarse Python bindings.
- Median decode throughput excluding prefill/first token: 1,622.86 tokens/s.
- Runs: 1,595.72 / 1,632.20 / 1,624.86 / 1,603.32 / 1,622.86 tokens/s.
- Median prefill plus first token: 3.744 ms.
No tokenization, detokenization, HTTP, or queue latency is included in decode timing. These are growing-context token-generation results, not weight sweeps. CPU affinity and clocks were not locked. No comparison with a tuned llama.cpp baseline has been run; the required speed advantage is not established.
Why These Numbers Differ From The Original Experiments
The earlier session results (hundreds of thousands of tok/s) are not comparable to these, for verifiable reasons on both sides:
- The supplied C harness (
decode_int8/decode_hybrid) looped over flat weight buffers with no layers, no logits, no sampling, and no token feedback. Its "tok/s" measured buffer-processing rate, not decoding. - Its Hugging Face baseline ran
use_cache=False, which undercounts real HF decode speed; its DDR4 column in the showcase table was computed from34500 / model_MB, not measured. - Kernel improvements have raised this snapshot's decode rate several-fold from the first working build, and it remains compute-bound (see the probe results above), so neither the old synthetic numbers nor the current 1,974 tok/s represent a final ceiling.
Kernel Optimization History
Baseline (per-row indirect dot calls, no FMA, serial INT4 block reduces), Supra-Mini-8M, one thread, 128-token decode: FP32 757, INT8 1,623, hybrid INT4 1,152 tokens/s.
After the first kernel pass (explicit FMA with 4 independent accumulator chains, AVX2 matvec kernels that absorb the row loop so one native call serves a whole matrix, FMA required in the CPU dispatch check):
| Precision | Kernel-only rate | Full-model decode | Before | Gain |
|---|---|---|---|---|
| FP32 | 30.8 GB/s | 936 tok/s (31.5 MB weights) | 757 | +24% |
| INT8 | 24.2 GB/s | 1,974 tok/s (8.0 MB weights) | 1,623 | +22% |
| Hybrid INT4 | 4.8 GB/s per byte | 1,382 tok/s (6.6 MB weights) | 1,152 | +20% |
| SmolLM2-135M INT8 | - | 127 tok/s (135.4 MB weights) | 120 | +6% |
Cache-residency probe (unchanged conclusion): decode effective bandwidth (15.8 GB/s at INT8) remains below even the cold DRAM scan rate (25.1 GB/s), so decode is still compute-bound. The INT4 kernel is the current kernel bottleneck (nibble unpacking without VNNI); full-model non-dot work (attention, norms, 16k-logit sampling, epilogue) accounts for roughly a third of decode time at INT8.
Next levers, by expected impact: multithreaded matvec (row-parallel), prompt-prefill matrix-matrix batching, and further INT4 unpack improvements.
Activation-quantized kernel follow-up (2026-09-12)
The experimental act_precision="int8" path is now implemented for quantized
matrix products. It quantizes every activation row in 32-element blocks using
32767 / absmax, uses int16 pmaddwd MACs, and retains FP32 block scales for
the epilogue. It is numerically sound but is not enabled by default: on this
Ryzen 5 5600 (Zen 3, AVX2/FMA, no VNNI), the quantization work dominates small
layer GEMMs. Supra-Mini-8M, one thread, measured about 1,753 tok/s with INT8
weights and FP32 activations versus 603 tok/s with the experimental path;
hybrid-INT4 measured 1,297 versus 607 tok/s, and hybrid-FP4 1,238 versus 574.
Custom PPL comparisons put the activation-mode delta near +0.01 on Supra-Mini
and at or below +0.1 on SmolLM2-135M. Keep it as an explicit experimental API
until a fused/coarser quantizer or a CPU with suitable integer-dot instructions
shows a net benefit.
Unpack-tax reduction, phase 1 (2026-09-12)
Shipped: skip the nibble interleave (even/odd streams dot a permuted
activation scratch built once per call from plain local buffers, no TLS),
halving shuffle-port work per 32 weights; dead matvec_* family deleted
(~350 lines, verified uncalled); 5-arg scalar adapters keep non-AVX2 builds
working. Tried and reverted the same session: folding per-8 scale multiplies
into a per-block scalar epilogue — the serial per-block reduce loses to four
parallel mulps on Zen 3 out-of-order execution (measured slower, not faster).
Measured on Ryzen 5 5600: full suite green, Mini PPL identical to 3 decimals, production AutoBench once showed fp4-2T beating int4-2T 1254 to 1015 (+24%) then losing the rematch — direction flips are inside this box's ±15-20% multithread noise, so neither is claimed. Drift-controlled A/B (alternating int8/fp4 in one process): 0.87 vs 0.90 pre-change. Verdict: the remaining ~26 µops/32 (converts, scale muls, act loads) are structural on AVX2 without VNNI; 4-bit cannot beat int8 kernel efficiency on this chip. Realistic next levers: int8-container upcast at load (int4 quality, near-int8 speed, int8 size — wins wherever cache residency makes bytes free), the q8 integer-MAC track, stronger ISA (VNNI/AMX/tensor cores), or DRAM-bound regimes where the 2x byte saving dominates.
Fused epilogues tried and reverted (2026-09-12): accumulating o/down dots directly onto x_ and SiLU into the up-gemm epilogue measured ~7% SLOWER on Supra2-Medium int8-1T in a back-to-back stash A/B (580 vs 623, tight clusters, wrong order for a thermal explanation). Likely cause: branches plus a temp in the FMA loop beat clean separate vector passes on wide rows, and an unconditional permute scratch taxed int8/fp32 paths (since made lazy). Lesson: on Zen 3 streaming hardware, an extra cache-resident vector pass can be cheaper than a branch in the hot loop — fuse only with A/B proof.
Session scratch arena (2026-09-12): per-call permute/row mallocs replaced by Session-persistent grow-only buffers (caller-owned, pool-safe; q8 TLS left alone). PPL identical; 4-bit/int8 tok/s ratios up ~0.1 at 2T on both models (Medium int4 0.83 -> 0.98, fp4 0.83 -> 0.96) — consistent with removing malloc-lock contention across pool threads rather than raw bytes. Model now reports info["compiled"] = {key, hit: False, mode: "generic", canonical} as the Phase-0 key plumbing goes live (execution still generic).
Policy: AutoBench now defaults to 512 decode tokens and compiles nothing unless asked (inductor warmup taxed every default run); thread guidance is 2T on a loaded server, never 12T (oversubscription convoy collapses pooled paths while serial paths sail through).
Supra2-Medium-Base production run (2026-09-12, live server, 512 tokens)
First model that spills past L3 (fp32 96.78 MB, DRAM-streaming):
| Precision | tok/s (2T) | PPL | vs fp32 |
|---|---|---|---|
| int8 (7.65 MB L3) | 540 | 59.40 | +0.62% |
| hybrid-fp4 (20.32 MB L3) | 533 | 60.28 | +2.12% |
| hybrid-int4 (20.89 MB L3) | 475 | 62.61 | +6.07% |
| fp32 (96.78 MB DRAM) | 278 | 59.03 | — |
| HF eager fp32 | 93 | 59.03 | — |
| HF dynamic int8 | 72 | 95.45 | +61.7% |
hybrid-fp4 beats hybrid-int4 on speed (+12%) and quality (+2.1% vs +6.1%) where matrices are large enough to amortize the unpack. int8 still leads overall at near-zero quality cost. Speculation accepts at 96-99% and is the largest single lever here (fp32 +65% at spec 8, int4 +17%); HF dynamic int8 is outclassed on both axes. Threads peak at 2-4T; 12T convulses as usual.
Spec K curve (2026-09-12, Medium-Base int8, 2T, 64 tokens)
Cap raised 8 -> 16 to test whether near-100% acceptance favors longer drafts. Measured: K=0: 844, K=2: 974 (84% acc), K=4: 1016 (76%), K=8: 954 (71%), K=12: 915 (69%), K=16: 834 (71%). Peak at K=4: acceptance decays with draft length while block-verify cost grows linearly, so dilution loses past the peak. Optimal K is prompt-dependent (longer, more repetitive generations accept more); the sweep now covers 0-16 with per-thread accept% tables so each setup can read its own peak. Thread sweep preset trimmed to 1,2,4 (above 4T never pays on these sizes).
SmolLM2-135M precision x thread (2026-09-13, single-harness, act fp32, rev 93efa2f0)
Weight bytes (engine.info): int8 135.4 MB (4.2x L3), hybrid-int4 105.1 MB, hybrid-fp4 100.2 MB; all spill, fits_l3=false. 512-token decode, greedy, spec_k=0, one live engine at a time:
| Threads | int8 tok/s (GB/s) | hybrid-int4 tok/s (GB/s) | hybrid-fp4 tok/s (GB/s) |
|---|---|---|---|
| 1T | 88 (11.9) | 72 (7.5, 0.82x) | 73 (7.3, 0.83x) |
| 2T | 120 (16.2) | 107 (11.3, 0.89x) | 106 (10.6, 0.88x) |
| 4T | 124 (16.8) | 122 (12.8, 0.98x) | 122 (12.2, 0.99x) |
PPL (514 scored tokens, window 128): int8 34.74, hybrid-int4 42.07 (+21.1% vs int8), hybrid-fp4 41.32 (+18.9% vs int8; fp4 beats int4 on quality at fewer bytes, speed tied within noise). Central 64-token 1T check: int8 120.9, int4 90.4, fp4 89.0 tok/s with identical weight bytes — same ordering. 128-token bracket int8-1T 114.5 tok/s / 15.5 GB/s matches the old 120-127 table; 512-window context growth costs ~21%, so ratios are intra-run only.
Verdict: int8 is bandwidth-bound (saturates at 2T, +3.6% to 4T), 4-bit paths are kernel-bound at every thread count (1T->4T ~1.7x scaling, DRAM idle, never cash the 1.29x byte edge). Kernel bet unchanged on Zen 3 / no-VNNI: int8 default everywhere, fp4 the 4-bit pick. 360M not fetched (724 MB BF16 shard); expect the amplified same pattern. Next levers: spec on DRAM-bound int8-135M, int8-container upcast at load for cache-resident sizes (test-gated), VNNI-class ISA. Compile Phase-0 stays hit:False/mode:generic (plumbing only, execution generic) — correct until a PPL-gated execution change ships.
Threading Results
Persistent native pool, row-sliced matvecs, pause-spin barriers (yield/sleep only after long idle). Supra-Mini-8M INT8, 256-token decode, medians:
| Threads | Tokens/s | vs 1T |
|---|---|---|
| 1 | 1,548 | 1.00x |
| 2 | 2,148 | 1.39x |
| 4 | 2,036 | 1.31x (regresses) |
FP32 (31.5 MB, DRAM-bound) shows no thread scaling: 998 -> 1,004 tok/s at 4T.
The Ryzen 5600 sweet spot is 2 threads for cache-sized quantized models;
beyond that, shared L3/LSU bandwidth saturates. gVisor-class sandboxes showed
~19 us sync round trips, so threading there may need larger models to pay off.
cism bench/serve --threads N sets the pool size; default remains 1.
Speculative Decoding (Prompt Lookup, Greedy)
Implemented: n-gram drafter over full token history (up to 4-gram tails), block verification via per-matrix GEMM (all T candidates share one weight stream), greedy argmax acceptance with KV rollback, budget/EOS boundaries. Speculative greedy output is bitwise identical to plain greedy decode (tested across architectures, precisions, rejection paths, and budgets).
SmolLM2-135M (135 MB INT8 / 540 MB FP32 - exceeds the 32 MB L3), one thread, 256 tokens, repetitive prompt for high draft acceptance:
| Configuration | Tokens/s | Gain |
|---|---|---|
| FP32 plain | 51 | 1.00x |
| FP32 spec_k=4 | 83 | 1.63x |
| INT8 plain | 106 | 1.00x |
| INT8 spec_k=4 | 130 | 1.23x |
Reason: speculation amortizes weight-bandwidth (one stream per verify block), so gains appear when decoding is DRAM-bound - the FP32 regime. The INT8 kernel is compute-bound (conversion-heavy SIMD), so amortizing bandwidth buys little until integer kernels land. On the AVX-512 VNNI cloud target, integer kernels should make INT8 bandwidth-bound and transfer the full speculation multiplier. Two structural costs: the bonus token costs one extra single-token pass (ideal (K+1)/2, measured 1.6x at K=4), and low-acceptance text can make speculation slightly negative - disable it when drafts do not match the domain.
Memory Control Verification
Real-model page locking and touch on this machine (Supra-Mini-8M INT8):
lock_pages(): 8,022,880 of 8,022,784 weight bytes locked (page rounding).touch(): 8,022,880 bytes swept.- Generation while keep-warm active and pages locked: normal.
unlock_pages(): 6,268,768 bytes reported. Smaller than the lock count because adjacent weight regions share 4 KiB pages; a page released by one region cannot be released again by its neighbor. Counts are approximate byte sums of successful OS calls, not exact page accounting.
x86 cannot pin cache lines; locking prevents OS paging, and sweeping only refreshes residency while running. Suite at time of writing: 571 tests passed.
Cache-Residency Probe Results
cism cache-report measures a warm full-weight scan, a cold scan after
evicting caches with a 96 MiB scrub buffer, and decode effective bandwidth
(weight bytes x decode tok/s). AMD Ryzen 5 5600, INT8, one thread:
| Model | Weights | Warm scan | Cold scan | Decode | Effective | Verdict |
|---|---|---|---|---|---|---|
| Supra-Mini-8M | 8.0 MB | 28.5 GB/s | 25.1 GB/s | 1,646 tok/s | 13.2 GB/s | Compute-bound |
| SmolLM2-135M | 135.4 MB | 21.4 GB/s | 23.1 GB/s | 120 tok/s | 16.3 GB/s | Compute-bound |
Interpretation:
- Decode bandwidth is well below even the cold (post-eviction) scan rate in both cases. The decoder does not saturate DRAM, so cache residency is not the current bottleneck; kernel efficiency is (scalar conversion-heavy quantized kernels, sequential prefill, single thread).
- The warm/cold scan ratio is small for these linear probes; hardware prefetchers hide much of DRAM latency on sequential streams, so this probe bounds bandwidth rather than proving residency. Definitive per-event counts (LLC miss) require hardware PMU tooling such as WindowsPerf or AMD uProf.
- The earlier session's implied "SRAM 3x-6.5x multiplier" was not observed by this probe on this hardware and remains unproven.
Speed work priorities implied by this data: quantized kernel quality, prefill matrix-matrix, and multithreading before further cache-tuning.
Measured Comparison Against PyTorch 2.14 And torch.compile
Supra-Mini-v5-8M, revision bb98e356, FP32 checkpoint, greedy, one torch/CISM
thread, 6-token prompt, 256-token fixed-length decode (EOS ignored both sides),
warmup absorbed, Dynamo counters verified zero recompilations during
measurement. Torch baselines use a KV cache (never use_cache=False).
Script: scripts/bench_vs_torch.py; CISM: cism bench --max-tokens 256.
| Decoder | Tokens/s | CISM FP32 (998.5) | CISM INT8 (1,716.4) |
|---|---|---|---|
| HF generate(use_cache=True) eager | 158.9 | 6.28x | 10.80x |
| Manual eager loop + StaticCache | 195.8 | 5.10x | 8.77x |
| torch.compile default, dynamic=False | 260.1 | 3.84x | 6.60x |
| torch.compile default, dynamic=True | 224.9 | 4.44x | 7.63x |
The fair comparison is against torch.compile dynamic=False, its best configuration here: 3.84x at matched FP32 precision, 6.60x for CISM INT8.
Reasons, with evidence:
- Per-token dispatch. torch.compile lowers eager's per-op Python/ATen dispatch (159 -> 260 tok/s, +64%) but still launches dozens of generated kernels through a Python-level decode loop each token. CISM executes the whole token in one C++ call. Eager's effective bandwidth is 31.5 MB x 159 = 5.0 GB/s and compile's is 8.2 GB/s - both overhead-bound, while CISM FP32 runs memory-bound at ~30 GB/s kernel rate.
- Memory traffic. CISM INT8 reads 8.0 MB/token versus 31.5 MB for the FP32 checkpoint, cutting the memory-bound floor ~4x on top of the dispatch elimination.
- Quality. CISM FP32 matches eager within 1e-5 max logit error (established above). INT8 quality reporting remains a beta gate.
Caveats: compile mode is "default" (max-autotune not used); one compile configuration initially crashed with an Inductor codegen bug on this model (index-out-of-bounds in a generated kernel) and only ran after cache-reset handling; multithreaded torch was not measured; prefill/TTFT is excluded from decode rates. This is one cache-sized model on one machine - not yet the agreed multi-model llama.cpp gate.
4-bit quantization scheme study (2026-09-12)
NumPy simulation mirroring the native hybrid-int4 quantizer (round-half-away, blocks of 32 along the row, per-block scale; validated against the native engine to 8 decimal places on all checkpoints). PPL over the fixed AutoBench corpus (514 tokens, window 128), MLP matrices 4-bit, all other 2D matrices per-row INT8, vectors FP32.
Machine: Ryzen 5 5600, local checkpoints, scripts/quant4_study.py.
| scheme (MLP element + scale) | bits/w | 8M ΔPPL | 100M ΔPPL | 135M ΔPPL |
|---|---|---|---|---|
| int4 + FP32 b32 (hybrid-int4) | 5.00 | +15.4% | +6.1% | +23.2% |
| int4 + E4M3 b32 | 4.25 | +15.5% | +6.0% | +26.7% |
| int4 + E8M0 b32 | 4.25 | +308% | +80.8% | +537% |
| E2M1 + E8M0 b32 (OCP MXFP4) | 4.25 | +318% | +92.4% | +323% |
| E2M1 + E4M3 b32 | 4.25 | +18.7% | +3.5% | +23.3% |
| E2M1 + E4M3 b16 (NVFP4 granularity) | 4.50 | +15.0% | +2.1% | +21.0% |
| int4 + E4M3 b16 | 4.50 | n/a | +8.1% | n/a |
| E2M1 + FP32 b32 | 5.00 | +13.7% | +4.4% | +19.9% |
| NF4 + FP32 b32 (QLoRA) | 5.00 | +44.1% | +8.8% | +32.0% |
| int4 + error feedback b32 | 5.00 | +36.5% | +16.9% | +80.9% |
Findings:
- Scale precision dominates element format; power-of-two (E8M0) scales are catastrophic for both uniform and FP elements.
- E2M1 elements beat uniform INT4 at equal scale precision and block size (100M model, b16+e4m3: +2.1% vs +8.1%).
- E2M1 + E4M3 b16 is the only scheme that improves on hybrid-int4 at every model size while using fewer bits (4.50 vs 5.00): tied on 8M, 3x better on 100M, better on 135M.
- NF4 (QLoRA) loses badly on these trained small models; naive per-block error feedback with a fixed scale inflates residuals and hurts everywhere.
Why hybrid-int4 decodes slower than int8 (measured, 1 thread, same checkpoints): the nibble unpack is a ~10-instruction dependent chain per block, so hybrid-int4 delivers only 7.6-8.1 GB/s of weight-equivalents vs 12.5-14.0 for int8 and 26.6 for fp32, despite moving 2.5x fewer bytes than fp32. The E2M1 LUT kernel replaces that chain with a 16-entry table lookup and is the reason to build the mode.
Compile one-shot canonical parity (2026-09-14)
Phase-0 key plumbing is now cross-language exact. Python
compile_cache.canonical_model_config() builds the native 12-field pipe
model_type|hidden|intermediate|layers|heads|kv_heads|head_dim|vocab| context|hexfloat(eps)|hexfloat(rope_theta)|tied with FP32-round-trip
struct.pack('f') hexfloat, matching Model::refresh_compile_key
bit-for-bit. Verified: tiny Llama 16/32/1/4/2/4/32/32 probe Python
canonical equals Model.info["compiled"]["canonical"], and
native.compile_key(python_canonical, precision, act, threads) equals
Model.info["compiled"]["key"] (d4547092c73a7aac on the probe).
Legacy canonical_config() sorted-JSON remains for early tests only.
Kept as a stable model fingerprint (zero runtime cost, tested); the
packed-artifact load idea is dropped with the project (see below).
SmolLM2-135M full AutoBench matrix (2026-09-13, 512 tokens, act fp32)
Local file autobench_135M_512.json (untracked measurement, torch
baselines disabled by policy include_compile=False): 4 precisions x
threads (1,2,4) x spec_k (0,2,4,8,12,16) = 72 decode cells, PPL once per
precision (514 scored tokens, window 128). Revision 93efa2f0 (report
does not yet stamp revision; next run should copy
engine.info["revision"]). All fits_l3=false (L3 cache size 32.0 MB).
Weight storage in megabytes: int8 129.16 MB (135,439,104 bytes), hybrid-int4 100.27 MB (105,141,504 bytes), hybrid-fp4 95.52 MB (100,164,864 bytes), fp32 513.13 MB.
Decode rate in tokens/s at spec_k=0 (effective bandwidth in GB/s = weight bytes x tok/s):
| Precision | 1 thread tok/s (GB/s) | 2 threads tok/s (GB/s) | 4 threads tok/s (GB/s) | Best tok/s (threads, spec_k) | PPL (+vs int8) |
|---|---|---|---|---|---|
| fp32 | 43.78 (23.6) | 57.44 (30.9) | 55.34 (29.8) | 106.41 (4T, K=8) | 34.15 |
| int8 | 91.53 (12.4) | 124.82 (16.9) | 130.03 (17.6) | 147.19 (4T, K=2) | 34.74 (--) |
| hybrid-int4 | 71.93 (7.6) | 106.54 (11.2) | 115.69 (12.2) | 140.59 (4T, K=16) | 42.07 (+21.1%) |
| hybrid-fp4 | 71.85 (7.2) | 102.10 (10.2) | 117.51 (11.8) | 132.90 (4T, K=2) | 41.32 (+18.9%) |
Ratios int8 vs 4-bit at spec 0: 1T 1.27x, 2T 1.17-1.22x, 4T 1.11-1.12x. Scaling 1T to 4T: int8 1.42x but 2T to 4T only +4.2% (saturates, bandwidth-bound); int4 1.61x, fp4 1.64x, still climbing with DRAM idle (kernel-bound nibble unpack, no VNNI). Verdict unchanged: int8 default everywhere, fp4 the 4-bit pick (better quality at fewer bytes, speed tied within ±15-20% multithread noise; ratios single-harness only).
Spec K-curve on 135M DRAM-bound int8 (512 tokens, greedy)
Accept rate in percent is deterministic across threads per K (88.0% at K=2 decaying to 72.4% at K=16).
| Threads | K=0 tok/s | K=2 tok/s | K=4 tok/s | K=8 tok/s | K=12 tok/s | K=16 tok/s | Peak gain vs K=0 |
|---|---|---|---|---|---|---|---|
| 1T | 91.53 | 98.42 | 100.47 | 99.45 | 96.14 | 93.67 | K=4, +9.8% |
| 2T | 124.82 | 136.05 | 132.62 | 131.34 | 129.72 | 126.94 | K=2, +9.0% |
| 4T | 130.03 | 147.19 | 146.55 | 135.13 | 136.81 | 143.53 | K=2, +13.2% |
Global best: int8 147.19 tok/s at 4 threads, spec_k=2. Contrast fp32 +73-92% peaking K=8 (bandwidth-bound, block verify amortizes DRAM) vs 4-bit +5-21% erratic (kernel-bound). Peak at small K because acceptance decays with draft length while GEMM verify cost grows linearly, plus the bonus-token extra pass. Prompt-lookup drafter (up to 4-gram, greedy-only, bitwise identical) helps repetitive prompts only; it does not fix the 135M slow complaint (still ~147 tok/s vs ~1,600 tok/s cache-resident 8M). Default 135M int8: 4 threads, spec_k=2 (2-thread servers also K=2).
trust_remote_code custom-arch assessment (2026-09-14)
trust_remote_code=True authorizes tokenizer Python only, never custom
modeling operators; trusted code itself is not sandboxed. Enforcement:
loader _ARCHITECTURES={"llama","qwen3"}, validate_config rejects
non-llama/qwen3 model_type and any auto_map beyond AutoTokenizer
even with trust enabled, cross-repo refs rejected, _expected_shapes
plus _load_weights reject extra/missing tensors, native
Config::validate/parse_config reject bias/MoE/sliding-window/
non-silu/non-default-RoPE. Custom Negative/CMA/Ember2/Spark2A/Blaze
targets (ROADMAP gate 4: recurrence, routing, lane mixing, Engram state)
need new native Config fields, operators, KV layouts, quant definitions,
and PPL/parity gates per arch — large work, not a flag flip. Keep
explicit UnsupportedModelError; server stays opt-in
--allow-remote-code-load (403 otherwise).
Compile project dropped (2026-09-14, verdict: correctly built, not worth it)
Target was one-time load cost for faster batch-1 CPU decode (generic AVX2, PPL-gated). Verdict: the plumbing is correct and tested, but the goal is unreachable on this CPU, so the project is dropped and its speedup fiction removed:
- What was real: cross-language canonical parity (bit-exact, tested), hoisted dispatch loops (same kernels/order, bitwise identical, kept as plain code), manifest helpers (kept, unused by load path).
- What was fiction: manifest-hit as "compile hit" (bookkeeping, not a
faster layout),
compiled.hit/modein Engine info and tester cards (removed), any decode tok/s attributed to compile (zero: mode was generic everywhere; the +13.2% is spec K=2). - Why not worth it: remaining ~26 µops/32 are structural on AVX2/no-VNNI; int8 is bandwidth-bound (2T saturates), 4-bit kernel-bound; fresh 512 runs show hoisting moved 4-bit ~+23% but int8 K0 flat (-0.9%), confounded by idle-machine effect — no clean decode win to claim; inductor costs 107 s warmup for 32 tok/s against CISM's 190 (6-8x lead from the pre-existing single-call C++ path + quant, not compile). VNNI/VBMI or integer-MAC kernels would reopen the question; until then the line is closed. Fingerprint key + hoisted loops stay; hit/UI/ future-speedup prose goes.
Fresh 512-token 4T confirmation (SmolLM2-135M, all precisions, spec sweep)
Re-ran 24 cells (4 precisions x 4T x spec 0,2,4,8,12,16), 512 tokens, 127.3 s wall, torch off. Accept% deterministic and identical to the old file; PPL/weights identical (fp32 34.15, int8 34.74, int4 42.07, fp4 41.32; 513.13/129.16/100.27/95.52 MB, all fits_l3=false). Absolute tok/s ran +11-30% hotter than the old file (idle machine); cells >25% off are marked SUSPECT — use this table, do not mix files.
| Precision | K=0 | K=2 | K=4 | K=8 | K=12 | K=16 | Best |
|---|---|---|---|---|---|---|---|
| fp32 | 58.77 | 72.89 | 101.92 | 125.09 | 122.38 | 121.18 | K=8 125.09 |
| int8 | 128.90 | 159.99 | 162.64 | 171.76 | 171.37 | 177.03 (SUSPECT vs old) | K=16 177.03 |
| hybrid-int4 | 142.63 | 155.69 | 151.97 | 159.86 | 160.29 | 156.40 | K=12 160.29 |
| hybrid-fp4 | 144.58 | 150.73 | 153.96 | 160.41 | 155.88 | 153.82 | K=8 160.41 |
Accept%: fp32 96.5/94.0/89.9/88.1/85.5; int8 88.0/83.4/77.1/73.0/72.4; int4 90.1/84.4/79.1/75.6/75.3; fp4 84.1/76.6/70.3/66.7/67.6 (K=2..16). Note 4-bit now leads int8 at K=0 (142.63/144.58 vs 128.90) in this run — confounded by machine state (old file had int8 leading), so no compile attribution claimed without a same-state A/B. fp4 vs int4: tied within noise at every K (max gap 3.7%), fp4 keeps the quality edge (+18.9% vs +21.1% PPL), so fp4 stays the 4-bit pick.
4T K=2 vs regular baselines (SmolLM2-135M int8, 512 tokens)
Global best cell is int8 4 threads spec_k=2 at 147.19 tok/s (accept
87.96%). Regular baselines from the same autobench_135M_512.json run
(include_torch=False, so no HF eager/compile in this file):
| Baseline (4T) | Tok/s | vs 4T K=2 (147.19) |
|---|---|---|
| int8 K=0 (same weights, no spec) | 130.03 | +13.2% with spec |
| int8 2T K=0 | 124.82 | +17.9% (threads+spec) |
| int8 1T K=0 | 91.53 | +60.8% |
| fp32 K=0 | 55.34 | 2.66x |
| fp32 best (K=8) | 106.41 | 1.38x |
| hybrid-int4 K=0 | 115.69 | 1.27x |
| hybrid-int4 best (K=16) | 140.59 | 1.05x |
| hybrid-fp4 K=0 | 117.51 | 1.25x |
| hybrid-fp4 best (K=2) | 132.90 | 1.11x |
PPL in same run: int8 34.74, fp32 34.15, int4 42.07, fp4 41.32. Spec pays (+13.2%) because 135M int8 spills (129.16 MB vs 32 MB L3): block verify streams weights once per block via GEMM. It does not make 135M fast in absolute terms (still ~147 vs ~1,600 cache-resident 8M).
HF eager/compile baselines (SmolLM2-135M, 512 tokens, 4T, same harness)
One run_benchmark (src/cism/autobench.py:223) with
include_torch=True, include_compile=True, max_new_tokens=512,
threads 4, spec sweep 0-16, rev 2c4d40c. Whole sweep ran ~4-15% under
the fresh reference (mild thermal throttle); HF rows are baselines, not
throttle-sensitive conclusions.
| Decoder (512 tok, 4T) | Tok/s | PPL (514 scored tokens) |
|---|---|---|
| CISM int8 K=0 | 131.88 | 34.74 (nll 3.5480, +1.73% vs eager) |
| CISM int8 best K=16 | 167.23 | same |
| CISM fp32 K=0 / best K=8 | 53.32 / 114.67 | 34.15 (+0.002%) |
| CISM hybrid-int4 K=0 / best K=8 | 137.20 / 153.07 | 42.07 (+23.18%) |
| CISM hybrid-fp4 K=0 / best K=16 | 137.38 / 150.25 | 41.32 (+21.00%) |
| HF eager fp32 | 23.79 (21.52 s) | 34.15 (nll 3.5308) |
| HF dynamic int8 | 19.74 (25.94 s) | 81.12 (broken, +137.5%) |
| HF torch.compile fp32 | 25.18 (20.33 s, warmup 147.3 s) | not scored |
| HF compile int8 | 16.86 (30.38 s) | not scored |
| torchao int4 | skipped (not installed) | -- |
Speedups (same harness): CISM int8 K=0 5.54x eager / 5.24x compile; best K=16 7.03x / 6.64x; fp32 K=0 2.24x/2.12x, best K=8 4.82x/4.55x. Compile barely beats eager (+5.9%) after a 147 s warmup; dynamic int8 is slower than fp32 AND destroyed (PPL 81). Weights: 513.13/129.16/100.27/ 95.52 MB, all fits_l3=false (L3 32 MB).
Surjo hybrid native Phase1 (loader+fp32 decode)
Scope: SurjoLabs Surjo hybrid import plus native fp32 decode runs. No Surjo tok/s, no Surjo PPL, no speedup claim in this section. Reference dense numbers below are context only, single-harness, ±15-20% multithread noise (see 135M matrix notes); do not mix files or attribute cross-run deltas to Surjo.
Config fingerprint (Surjo-50m, tests/test_surjo.py:20): model_type
surjo / SurjoForCausalLM, vocab 32768, hidden 512, intermediate
1536, layers 10, context 2048, prelude 1 / recurrent 8 / coda 1,
groups 2 / passes 2 / gdn_per_xsa 3, heads 8 / kv 4 / head_dim 64,
gdn_v_heads 8 / gdn_k_dim 64 / gdn_v_dim 64 / kernel 4,
gdn_allow_neg_eigval false, xsa_projection true, bias false,
dropout 0.0, hidden_act silu, eps 1e-5, theta 10000.0, default RoPE
(partial_rotary_factor 1.0), tied true, layer_types
[F,L,L,L,F,L,L,L,F,F], auto_map pinned to
configuration_surjo/modeling_surjo (never executed). Base
checkpoints ship use_cache=false; import_model
(src/cism/loader.py:666) preserves the flag verbatim and it does
not disable native caching (XSA KV slots + GDN recurrent/conv state
always allocate).
Checkpoint header: 178xF32 tensors in model.safetensors, tied
lm_head omitted, no rotary inv_freq buffers (RoPE is computed in
forward()). _expected_shapes_surjo length is 179 including the
tied-head alias entry; spot-checked in
test_surjo_50m_expected_shapes_match_header (embed (32768,512),
XSA q (512,512) / k (256,512) / q_norm (64,), GDN q/k (512,512)
/ v (512,512) / conv (512,1,4) / f0 (64,512) / f1 (512,64) /
A_log (8,) / dt_bias (512,) / g1.bias (512,) / o_norm (64,)).
Loader (src/cism/loader.py): validate_surjo_config,
surjo_exec_plan (18 steps: 1 prelude + 2 passes x 2 groups x
(3 GDN + 1 XSA) + 1 coda; XSA slots [0..5], 12 GDN (layer,pass)
steps), surjo_num_xsa_slots (=6), _expected_shapes_surjo:393,
_load_weights Surjo branch (no buffers, strict shape/dtype/finite
checks, tied-head alias by identity), import_model dispatch
(model_type=="surjo" -> Surjo path, else dense). _SURJO_SUPPORT
string still reads "next phase" in this snapshot; Phase1 below is
what landed natively.
Native (native/runtime.{hpp,cpp}, native/bindings.cpp:468,
src/cism/engine.py:80 dispatch to _native.SurjoModel):
SurjoModel/SurjoSession fp32 decode. XSA: 6 slots, KV
[slots,capacity,kv_width] with kv_width=256; at max context 2048
that is 4 MiB/slot (K+V), 24 MiB total (session-sized by
prompt+max_new, so 24 MiB is the upper bound). GDN: 12 states
(6 GDN layers x 2 passes), each S is H*K*V=8*64*64=32768 floats =
128 KiB (12x128 KiB = 1.5 MiB), plus depthwise-conv raw FIFO
(kernel-1)=3: q/k/v FIFOs ~72 KiB each, ~0.21 MiB total. MLP is
ClampedMLP (gate clamped to [-15,15] then silu*up,
runtime.cpp:2010). spec_k is rejected for Surjo
(next_tokens throws unless spec_k==0; Engine.stream(spec_k=2)
raises ValueError, covered by test_surjo_engine_decode_runs).
Prefill is sequential per-token (forward_tokens loop); logits /
nll teacher-forcing paths exist but are unscored against HF so far.
Quant mapping reuse, no new format: protected_storage (embed, head,
all XSA/GDN projections incl. f0/f1/g0/g1/o_proj) is fp32 or int8;
MLP gate/up/down follow the dense hybrid rule (hybrid-int4 -> int4,
hybrid-fp4 -> fp4). Same Matrix kernels/scratch/act_q8 path as
dense; only norms/vectors/conv/A_log/dt_bias stay fp32.
Parity gate: dense bar is max abs logit error <1e-4 vs HF eager.
Surjo fp32 decode/smoke runs (tiny shapes fixture, tied-head
identity, Engine.logits/generate finite, spec rejection). HF
modeling_surjo.py parity script and Surjo PPL are pending — no
parity number claimed here.
Batch/paging: Phase0 files unchanged. Single-request serialized
Engine generation lock, create_session/next_tokens,
cancel/finish, lock_pages/touch/scan weight-memory controls
are reused as-is. No continuous batching, no paged-KV layout, no new
server path for Surjo.
FLA CPU verdict: decode needs no Flash-Linear-Attention (single-token GDN recurrence is O(1) state + per-slot causal XSA). Chunkwise/FLA prefill is an optional long-prompt throughput optimization only, not correctness; current sequential prefill stands.
Dense context (same autobench_135M_512.json, include_torch=False,
NOT Surjo results): 4T K=0 fp32 55.34 / int8 130.03 / hybrid-int4
115.69 / hybrid-fp4 117.51 tok/s; best cells fp32 106.41 (K=8), int8
147.19 (K=2, accept 87.96%), hybrid-int4 140.59 (K=16), hybrid-fp4
132.90 (K=2); PPL fp32 34.15 / int8 34.74 / int4 42.07 / fp4 41.32.
Ratios single-harness only; 4-bit vs int8 gaps sit inside ±15-20%
noise. Surjo inherits none of these numbers.
Surjo-50m full AutoBench (Phase2, 2026-09-15, 4T K0)
File autobench_surjo50m_512.json (local snapshot
35c9fa6e, 512-token HF window, CISM greedy K0 only, act fp32,
threads 1/2/4). Weights (engine.info): fp32 205.21 MB, int8
51.84 MB, hybrid-int4 43.27 MB, hybrid-fp4 41.86 MB; L3 32.0 MB,
all fits_l3=false. PPL over the fixed AutoBench corpus (527 scored
tokens): fp32 59.29 / int8 59.46 (+0.30% vs HF eager) / hybrid-int4
62.90 (+6.10%) / hybrid-fp4 61.80 (+4.25%). CISM int8/fp32 PPL matches
HF eager fp32 59.28 to 0.3% — quant quality gate holds at int8, 4-bit
pays +4-6%.
CISM decode at 4T K0 (greedy; CISM stopped on EOS at 43-48 tokens, HF forced 512 via min_new_tokens — lengths differ, ratios indicative):
| Precision | 1T tok/s | 2T tok/s | 4T tok/s | PPL |
|---|---|---|---|---|
| fp32 | 62.18 | 81.03 | 81.06 | 59.29 |
| int8 | 129.86 | 158.39 | 207.80 | 59.46 |
| hybrid-int4 | 109.05 | 159.04 | 175.49 | 62.90 |
| hybrid-fp4 | 105.29 | 153.13 | 171.31 | 61.80 |
HF baselines (same harness, FP32 CPU, same thread count):
| Decoder (4T) | Tok/s | PPL |
|---|---|---|
| HF eager fp32 | 26.27 (1T 28.08 / 2T 27.74 — inverse scaling, MT noise) | 59.28 |
| HF dynamic int8 | 18.18 (main cell 19.93; sweep 17.42/18.95/18.18) | 66.35 destroyed (+11.9%) |
| HF torch.compile fp32 | not scored | |
| HF compile int8 | 16.73 (same StaticCache invalidity) | not scored |
Speedups at 4T K0 vs HF eager 26.27: int8 7.91x, hybrid-int4 6.68x,
hybrid-fp4 6.52x, fp32 3.09x. speedup_vs_compile in the JSON divides
by the invalid 31.72 and is void — do not cite.
Compile correction: modeling_surjo.py:645 discards non-SurjoCache
past (if not isinstance(past, SurjoCache): past = None), so the
autobench StaticCache loop allocated a fresh cache per forward call and
state never carried — 31.72 tok/s is not a Surjo compile number.
src/cism/autobench.py now detects surjo_mode and threads returned
past_key_values (SurjoCache) under torch.compile(dynamic=True),
recording {"skipped": "StaticCache invalid for Surjo"} if the
threaded path cannot run — never a StaticCache tok/s. Correct path
proven by scripts/bench_compiler.py: threaded SurjoCache +
dynamic=True 41 tok/s (compile warmup 43-119 s across runs;
189.3 s in the old JSON was the invalid-path warmup);
0.06 tok/s, recompiles per token as XSA KV
length grows). torchao int4 blocked on Windows (no mslk wheel —
dynamic=False unusable (Requires mslk >= 1.0.0); torchao int8wo is the CPU-working control
only. FLA CPU dead here (Triton needs GPU; fla.ops.gdn2 chunk/fused
probe does not run on CPU) — decode needs no FLA anyway (single-token
GDN recurrence is O(1) state + per-slot causal XSA).
Spec: unsupported for Surjo hybrid (GDN state destructive) — K0 only
(spec_note in JSON); K2≈K0 where probed, no speculation gain claimed.
Native SurjoSession rejects spec_k!=0; Engine.stream(spec_k=2)
raises.
Rigor: this JSON predates the harness (matrix rows lack
rep_rates/thermal_throttled/cpu_snapshot). Current
src/cism/autobench.py:_timed reports 1 warmup + median-of-reps
(--reps/--runs wired through cism bench; default 3) with min/max/
std, thermal_throttled when max-min spread exceeds 15% (wider than the
±15-20% MT noise band), and best-effort cpu_snapshot (clock/temp;
{} when unavailable, never fails the run). Next Surjo rerun picks up
medians + thermal flags automatically; do not mix old single-sample
cells with new median cells.
Surjo-50m rerun on current code (2026-09-15 night, 72 cells, reps=3)
autobench_surjo50m_512.json overwritten with current native
(GDN-OPT + XSA-blocked + forward_block + spec verify). Threads
1/2/4 x spec 0/2/4/8/12/16, 512tok, act fp32, median-of-3, all
thermal_throttled=false except two fp32-4T spec cells. CISM EOS-stops
43-48tok vs HF forced-512 — same caveat as before.
K0 medians: fp32 60.32/84.73/88.86, int8 141.76/197.34/234.87, int4 114.59/169.65/197.28, fp4 113.54/169.98/214.61. Best cells: fp32 K2 92.66, int8 K4 236.91, int4 K4 201.33, fp4 K0 214.61. Spec accept on bench prompt is 0.0 for fp32/int8 (no gain, no loss beyond noise); int4 6.25%->1.7%, fp4 8.3%->2.7% decaying with K, and spec K>0 is slower than K0 there (verify cost, no amortization win on this prompt). PPL unchanged: 59.29/59.46/62.90/61.80.
HF equal-opportunity, same harness, threads swept, median-of-3:
eager 31.22/30.43/29.50 (4T 29.50 PPL 59.28),
dynamic-int8 20.27/22.59/22.37 (4T 22.69 PPL 66.35),
compile threaded-SurjoCache dynamic=True 38.77 warmup 100.6s,
compile-int8 21.10, torchao-int4 failed Requires mslk>=1.0.0.
Speedups vs eager 29.50: int8 K0 7.96x, best K4 8.03x;
vs compile 38.77: K0 6.06x. Old invalid StaticCache 31.72 is
superseded — do not cite it.
Supra-50M-Base full AutoBench (dense Llama, 2026-09-15 night)
File autobench_supra50M_512.json, Hub rev
521bfd3d3901fbeefa943a2f3d461a9d193c52ec, dense Llama
(hidden 512, layers 12, heads 8/kv 4/head_dim 64, intermediate
1408, vocab 32000, ctx 1024, tied, use_cache=false preserved).
Threads 1/2/4 x spec 0/2/4/8/12/16 = 72 cells, 512tok, act
fp32, median-of-3. Weights: fp32 ~209MB / int8 ~53MB class
(all fits_l3=false). PPL (527 scored tokens): fp32 63.32 /
int8 63.79 (+0.7%) / int4 65.36 (+3.2%) / fp4 68.90 (+8.8%).
CISM K0 medians: fp32 108.29/131.54/123.55, int8 214.81/284.36/292.58, int4 178.04/247.49/250.87, fp4 171.67/234.54/251.38. Best cells: fp32 K8 268.19 (2T), int8 K12 349.10 (4T), int4 K2 257.43, fp4 K8 274.25. Accept on bench prompt is 100% for fp32/int8/fp4 at every K (hence spec pays: int8 4T K0 292.58 -> K12 349.10, +19.3%; fp32 K0 123.55 -> K8 268.19, +117%); int4 86%->70%.
HF same harness, threads swept, median-of-3: eager
48.59/47.90/43.61 (4T 43.61 PPL 63.32), dynamic-int8
37.70/40.49/37.78 (4T 34.91 PPL 101.68 destroyed),
compile StaticCache dynamic=False 55.23 warmup 143.0s,
compile-int8 33.14, torchao-int4 Requires mslk>=1.0.0.
Speedups vs eager 43.61: int8 K0 6.71x, best K12 8.00x;
vs compile 55.23: K0 5.30x, best 6.32x.
FWKV Myosotis-1-base full support + AutoBench (2026-09-15 night)
Files inspected in snapshot 031fe84a34f667bc8bfff7dc3f8a77ef3c17efb2:
config.json (model_type fwkv, d_model 768, d_emb 192,
n_layers 13, ffn_mult 4, vocab 50257, wkv_floor 0.1, tied),
configuration_fwkv.py (FWKVConfig + HF aliases),
modeling_fwkv.py (FactorizedTiedHead, FWKVBlock with
proj_k/v/r/out + W recurrence + GELU FFN + 2x LayerNorm,
_supports_cache_class=False, tuple-state threading),
model.safetensors (147 F32 tensors, no rotary buffers).
Native support landed, mirroring the Surjo vertical slice:
loader.validate_fwkv_config + _expected_shapes_fwkv (147
shapes) + _load_weights fwkv branch (also fixed an eager
config["head_dim"] default crash affecting any arch without
head_dim) + import_model dispatch; native FwkvConfig /
FwkvModel / FwkvSession (WKV recurrence, exact GELU,
LayerNorm eps 1e-5, factorized head, fp32/int8/int4/fp4 reuse
via Matrix, forward_block batched projections + sequential
recurrence, spec verify with ~40KB state checkpoint/restore);
bindings FwkvModel/FwkvSession; engine dispatch;
tests/test_fwkv.py 11 passed; scripts/parity_fwkv.py:
tiny 1.19e-07, live vs HF fp32 5.72e-06 (<1e-4 OK).
Autobench: trust_remote_code=True for fwkv HF loads,
threaded-FWKV-state compile path (StaticCache invalid, same
class of bug as Surjo), _hf_context fallback to seq_len
(FWKVConfig has no max_position_embeddings).
File autobench_fwkv_512.json, 72 cells, reps=3 (v4 drafter rerun,
2026-09-15 late night; supersedes the v1-drafter run below). Weights:
fp32 425.38MB / int8 107.64MB. PPL: fp32 156.68 (= HF eager
156.68, gate holds) / int8 156.80 (+0.07%) / int4 156.63
(-0.04%) / fp4 158.53 (+1.18%). High absolute PPL is base
model quality, faithfully reproduced.
CISM K0: fp32 64.65/93.73/87.24, int8 169.01/268.22/342.27, int4 121.54/215.32/284.66, fp4 119.21/205.16/303.52. Best: fp32 K16 177.34, int8 K4 382.74, int4 K2 290.19, fp4 K2 305.98. Accept (bench prompt): fp32 85%->61%, int8 85%->62%, int4 77%->45%, fp4 74%->44% (K2->K16). Spec pays on fp32/int8: int8 4T K0 342.27 -> K4 382.74 (+11.8%); fp32 K0 87.24 -> K16 177.34 (+103%); 4-bit int4 still peaks near K0-K2, fp4 peaks K2.
HF same harness, threads swept, median-of-3: eager
42.72/45.59/43.42 (PPL 156.68), blocks-only dynamic-int8
34.74/43.88/51.26 (4T 51.22 PPL 158.42 +1.1% ok — fixed by
quantizing blocks Linears only; full-model quantize fails
on FactorizedTiedHead.proj.weight.t()), compile
threaded-FWKV-state dynamic=True 49.73 warmup 45.9s,
compile-int8 same path 46.93 warmup 36.2s, torchao-int4
Requires mslk>=1.0.0 (no Windows wheel). Speedups vs
eager 43.42: int8 K0 7.88x, best K4 8.81x; vs compile
49.73: K0 6.88x, best 7.70x. Previous run (v1 drafter):
K0 59.82/83.70/87.54 etc., best int8 K8 359.97 — v4 gains
come from restored accept (19%->85% @K2).
Spec drafter v4 + doom-loop check (2026-09-15 late night)
v1 (most-recent single match, verbatim tail truncated at history
end): Surjo bench 0% (correct text, nothing to match), Supra
bench 100% (degenerate loop flatters it). v2 (pure frequency
vote) collapsed on drifting text (FWKV bench int8 K2
85%->19%, K8 69%->9%). v3 (recency-decayed vote) same failure.
v4 (landed): most-recent strictly-inside occurrence per
position + iterative extension (fixes v1 truncation) +
rejected-block latch REMOVED (the 32-propose/<1/3-accept
latch fired during the short-history warmup and blinded the
session to later 85%-accept regions; rejected blocks stream
weights once per block via forward_block, so worst case on
0%-accept text is ~K0 speed). Shared by dense/Surjo/FWKV;
tests/test_native.py::test_speculative_draft_hits_and_misses
reimplemented against v4; suite 759 passed; K0==Kx exact
everywhere probed.
Doom-loop probe (greedy, 128 tok, fp32+int8, 3 prompts): Surjo stops coherently on bench/code prompts (48/15 tok, uniq 0.73/0.93, loop 0) but loops story prompts (128 tok, uniq 0.09-0.16, 5/11-period loops); Supra loops ALL prompts (128 tok, uniq 0.07-0.11, 7/9/12-period loops — bench, story, AND code); FWKV never hard-loops bench/story (uniq 0.28-0.54, loop 0; int8 bench drifts into a 10-period body/brain cycle) but degenerates to 128 spaces on code (uniq 0.02 both precisions). So Supra's 100% spec accept is a looping artifact, Surjo's 0% is healthy text, FWKV sits between. No repetition penalty exists anywhere (greedy only); spec accept must always be read next to uniqueness.
Portable kernels: Intel/AMD tiers + ARM64 phones (2026-09-16)
Tier ladder (runtime-dispatched, scalar fallback always works):
scalar (any CPU) < avx2+FMA (x86-64 baseline, all current
numbers) < avx-vnni (Zen4/Zen5, Alder Lake+: int8 VNNI fast
path) < avx512-vnni (server AVX512) / neon (ARM64 phones,
Apple Silicon, Raspberry Pi). New files: native/kernels_vnni.cpp
(dpbusd int8xint8, int8-activation ABI via quantize_row_i8),
native/kernels_neon.cpp (fp32/int8/int4/fp4 + uzp permute),
CPUID detection in native/kernels.cpp (AVX-VNNI = leaf7:1
EAX[4], AVX512-VNNI = leaf7:0 EBX[11] + foundation + ZMM state;
MSVC VNNI TU is EVEX so it dispatches only on AVX512-VNNI,
GCC VEX on AVX-VNNI), per-file codegen flags in CMakeLists
(dispatcher/runtime stay generic). kernel_name() keeps the
fp32 family; new kernel_variant() + has_avx_vnni /
has_avx512_vnni / has_neon bindings; cpu_features gains
avx_vnni; Model.info["kernels"] gains variant.
tests/test_kernels_portable.py (5 passed): tier reporting,
flag consistency, pre-VNNI fallback untouched, i8 quantizer
round-trip bounds, dpbusd bias-identity in numpy.
Proven on Ryzen 5 5600: builds clean (MSVC /arch:AVX512 VNNI TU
accepted), kernel avx2 / variant avx2 / vnni False,
suite 764 passed (759 + 5 new) — zero behavior change, as
designed (VNNI selector is null here so Matrix keeps the int16
Q8 / fp32-activation paths byte-for-byte). Pending hardware CI:
VNNI tok/s + int8-activation PPL re-gate on Zen4/ADL+ (expect
prefill/spec-verify wins ~2-3x integer MACs; decode stays
bandwidth-bound so batch-1 gains concentrate in K>0), NEON
tok/s + parity on ARM64 hardware (expects the usual
Silicon/Android numbers; tails fall back to scalar by
construction). No speedup is claimed for either until measured.