test1111111 / docs /VALIDATION.md
spitfire4794's picture
CISM remote autobench: full source + fleet runner, serve results on 7860
28a1a01
|
Raw History Blame Contribute Delete
47.2 kB
# Development Snapshot Validation
Date: 2026-09-08. This is not beta certification or a llama.cpp comparison.
## Environment
- Windows x64, AMD Ryzen 5 5600, 6 physical cores / 12 logical processors.
- CPython 3.13, MSVC Visual Studio Build Tools 2022 17.14.7.
- One native inference thread, runtime-dispatched AVX2.
- Installed wheel: cism 0.1.0.dev0.
- NumPy 2.5.3, Transformers 5.16.1, PyTorch 2.14.0+cpu for reference tests only.
- FastAPI 0.141.1, Uvicorn 0.52.4, OpenAI Python SDK 2.54.0.
## Tests
`python -m pytest -q`: **555 passed**, two upstream TestClient deprecation
warnings. Includes independent NumPy numerical references, Transformers tiny
model parity, packed quantization, invalid checkpoint/config handling, generation
metadata, cancellation, queue limits, SSE, official SDK, and a live Uvicorn TCP
roundtrip using real native inference.
`python -m pip check`: no broken requirements.
Both scalar-only and AVX2 native builds were also tested separately during kernel
development. Linux, sanitizers, long-running load tests, and full-context real
checkpoint certification remain outstanding.
## Real Checkpoints
Compared default FP32 eager HF computation with native last-token logits and
four-token greedy continuations at prompt lengths 4, 8, and 16. All tested
continuations agreed; maximum absolute logit error was below 0.0001.
- SupraLabs/Supra-Mini-v5-8M: FP32, INT8, hybrid INT4.
- veyra-ai/Veyra2-Blueberry-10M-Base: FP32.
INT8/INT4 references used reconstructed quantized weights, not original FP32
weights. These results establish implementation parity only, not a PPL quality
budget or task-quality equivalence.
## Local Throughput
```text
cism bench SupraLabs/Supra-Mini-v5-8M --precision int8 --runs 5 --max-tokens 128
```
Model revision: `bb98e3566a5ab3f24be16ee8db7f06f0cb884eb5`.
Packed weight storage including scales/norms: 8,022,784 bytes.
- Prompt: 6 tokens; generation: 128 tokens; EOS ignored for fixed-length timing.
- One warmup, then five runs; native generation through coarse Python bindings.
- Median decode throughput excluding prefill/first token: **1,622.86 tokens/s**.
- Runs: 1,595.72 / 1,632.20 / 1,624.86 / 1,603.32 / 1,622.86 tokens/s.
- Median prefill plus first token: **3.744 ms**.
No tokenization, detokenization, HTTP, or queue latency is included in decode
timing. These are growing-context token-generation results, not weight sweeps.
CPU affinity and clocks were not locked. No comparison with a tuned llama.cpp
baseline has been run; the required speed advantage is not established.
## Why These Numbers Differ From The Original Experiments
The earlier session results (hundreds of thousands of tok/s) are not comparable
to these, for verifiable reasons on both sides:
- The supplied C harness (`decode_int8`/`decode_hybrid`) looped over flat weight
buffers with no layers, no logits, no sampling, and no token feedback. Its
"tok/s" measured buffer-processing rate, not decoding.
- Its Hugging Face baseline ran `use_cache=False`, which undercounts real HF
decode speed; its DDR4 column in the showcase table was computed from
`34500 / model_MB`, not measured.
- Kernel improvements have raised this snapshot's decode rate several-fold from
the first working build, and it remains compute-bound (see the probe results
above), so neither the old synthetic numbers nor the current 1,974 tok/s
represent a final ceiling.
## Kernel Optimization History
Baseline (per-row indirect dot calls, no FMA, serial INT4 block reduces),
Supra-Mini-8M, one thread, 128-token decode: FP32 757, INT8 1,623,
hybrid INT4 1,152 tokens/s.
After the first kernel pass (explicit FMA with 4 independent accumulator
chains, AVX2 matvec kernels that absorb the row loop so one native call serves
a whole matrix, FMA required in the CPU dispatch check):
| Precision | Kernel-only rate | Full-model decode | Before | Gain |
| --- | --- | --- | --- | --- |
| FP32 | 30.8 GB/s | 936 tok/s (31.5 MB weights) | 757 | +24% |
| INT8 | 24.2 GB/s | 1,974 tok/s (8.0 MB weights) | 1,623 | +22% |
| Hybrid INT4 | 4.8 GB/s per byte | 1,382 tok/s (6.6 MB weights) | 1,152 | +20% |
| SmolLM2-135M INT8 | - | 127 tok/s (135.4 MB weights) | 120 | +6% |
Cache-residency probe (unchanged conclusion): decode effective bandwidth
(15.8 GB/s at INT8) remains below even the cold DRAM scan rate (25.1 GB/s),
so decode is still compute-bound. The INT4 kernel is the current kernel
bottleneck (nibble unpacking without VNNI); full-model non-dot work (attention,
norms, 16k-logit sampling, epilogue) accounts for roughly a third of decode
time at INT8.
Next levers, by expected impact: multithreaded matvec (row-parallel),
prompt-prefill matrix-matrix batching, and further INT4 unpack improvements.
### Activation-quantized kernel follow-up (2026-09-12)
The experimental `act_precision="int8"` path is now implemented for quantized
matrix products. It quantizes every activation row in 32-element blocks using
`32767 / absmax`, uses int16 `pmaddwd` MACs, and retains FP32 block scales for
the epilogue. It is numerically sound but is not enabled by default: on this
Ryzen 5 5600 (Zen 3, AVX2/FMA, no VNNI), the quantization work dominates small
layer GEMMs. Supra-Mini-8M, one thread, measured about 1,753 tok/s with INT8
weights and FP32 activations versus 603 tok/s with the experimental path;
hybrid-INT4 measured 1,297 versus 607 tok/s, and hybrid-FP4 1,238 versus 574.
Custom PPL comparisons put the activation-mode delta near +0.01 on Supra-Mini
and at or below +0.1 on SmolLM2-135M. Keep it as an explicit experimental API
until a fused/coarser quantizer or a CPU with suitable integer-dot instructions
shows a net benefit.
### Unpack-tax reduction, phase 1 (2026-09-12)
Shipped: skip the nibble interleave (even/odd streams dot a permuted
activation scratch built once per call from plain local buffers, no TLS),
halving shuffle-port work per 32 weights; dead `matvec_*` family deleted
(~350 lines, verified uncalled); 5-arg scalar adapters keep non-AVX2 builds
working. Tried and reverted the same session: folding per-8 scale multiplies
into a per-block scalar epilogue — the serial per-block reduce loses to four
parallel mulps on Zen 3 out-of-order execution (measured slower, not faster).
Measured on Ryzen 5 5600: full suite green, Mini PPL identical to 3 decimals,
production AutoBench once showed fp4-2T beating int4-2T 1254 to 1015 (+24%)
then losing the rematch — direction flips are inside this box's ±15-20%
multithread noise, so neither is claimed. Drift-controlled A/B (alternating
int8/fp4 in one process): 0.87 vs 0.90 pre-change. Verdict: the remaining
~26 µops/32 (converts, scale muls, act loads) are structural on AVX2 without
VNNI; 4-bit cannot beat int8 kernel efficiency on this chip. Realistic next
levers: int8-container upcast at load (int4 quality, near-int8 speed, int8
size — wins wherever cache residency makes bytes free), the q8 integer-MAC
track, stronger ISA (VNNI/AMX/tensor cores), or DRAM-bound regimes where the
2x byte saving dominates.
Fused epilogues tried and reverted (2026-09-12): accumulating o/down dots
directly onto x_ and SiLU into the up-gemm epilogue measured ~7% SLOWER on
Supra2-Medium int8-1T in a back-to-back stash A/B (580 vs 623, tight
clusters, wrong order for a thermal explanation). Likely cause: branches plus
a temp in the FMA loop beat clean separate vector passes on wide rows, and an
unconditional permute scratch taxed int8/fp32 paths (since made lazy).
Lesson: on Zen 3 streaming hardware, an extra cache-resident vector pass can
be cheaper than a branch in the hot loop — fuse only with A/B proof.
Session scratch arena (2026-09-12): per-call permute/row mallocs replaced by
Session-persistent grow-only buffers (caller-owned, pool-safe; q8 TLS left
alone). PPL identical; 4-bit/int8 tok/s ratios up ~0.1 at 2T on both models
(Medium int4 0.83 -> 0.98, fp4 0.83 -> 0.96) — consistent with removing
malloc-lock contention across pool threads rather than raw bytes. Model now
reports info["compiled"] = {key, hit: False, mode: "generic", canonical} as
the Phase-0 key plumbing goes live (execution still generic).
Policy: AutoBench now defaults to 512 decode tokens and compiles nothing
unless asked (inductor warmup taxed every default run); thread guidance is 2T
on a loaded server, never 12T (oversubscription convoy collapses pooled
paths while serial paths sail through).
### Supra2-Medium-Base production run (2026-09-12, live server, 512 tokens)
First model that spills past L3 (fp32 96.78 MB, DRAM-streaming):
| Precision | tok/s (2T) | PPL | vs fp32 |
| --- | --- | --- | --- |
| int8 (7.65 MB L3) | 540 | 59.40 | +0.62% |
| hybrid-fp4 (20.32 MB L3) | 533 | 60.28 | +2.12% |
| hybrid-int4 (20.89 MB L3) | 475 | 62.61 | +6.07% |
| fp32 (96.78 MB DRAM) | 278 | 59.03 | — |
| HF eager fp32 | 93 | 59.03 | — |
| HF dynamic int8 | 72 | 95.45 | +61.7% |
hybrid-fp4 beats hybrid-int4 on speed (+12%) and quality (+2.1% vs +6.1%)
where matrices are large enough to amortize the unpack. int8 still leads
overall at near-zero quality cost. Speculation accepts at 96-99% and is the
largest single lever here (fp32 +65% at spec 8, int4 +17%); HF dynamic int8
is outclassed on both axes. Threads peak at 2-4T; 12T convulses as usual.
### Spec K curve (2026-09-12, Medium-Base int8, 2T, 64 tokens)
Cap raised 8 -> 16 to test whether near-100% acceptance favors longer
drafts. Measured: K=0: 844, K=2: 974 (84% acc), K=4: 1016 (76%),
K=8: 954 (71%), K=12: 915 (69%), K=16: 834 (71%). Peak at K=4:
acceptance decays with draft length while block-verify cost grows linearly,
so dilution loses past the peak. Optimal K is prompt-dependent (longer,
more repetitive generations accept more); the sweep now covers 0-16 with
per-thread accept% tables so each setup can read its own peak. Thread sweep
preset trimmed to 1,2,4 (above 4T never pays on these sizes).
### SmolLM2-135M precision x thread (2026-09-13, single-harness, act fp32, rev 93efa2f0)
Weight bytes (engine.info): int8 135.4 MB (4.2x L3), hybrid-int4 105.1 MB,
hybrid-fp4 100.2 MB; all spill, fits_l3=false. 512-token decode, greedy,
spec_k=0, one live engine at a time:
| Threads | int8 tok/s (GB/s) | hybrid-int4 tok/s (GB/s) | hybrid-fp4 tok/s (GB/s) |
| --- | --- | --- | --- |
| 1T | 88 (11.9) | 72 (7.5, 0.82x) | 73 (7.3, 0.83x) |
| 2T | 120 (16.2) | 107 (11.3, 0.89x) | 106 (10.6, 0.88x) |
| 4T | 124 (16.8) | 122 (12.8, 0.98x) | 122 (12.2, 0.99x) |
PPL (514 scored tokens, window 128): int8 34.74, hybrid-int4 42.07 (+21.1%
vs int8), hybrid-fp4 41.32 (+18.9% vs int8; fp4 beats int4 on quality at
fewer bytes, speed tied within noise). Central 64-token 1T check: int8 120.9,
int4 90.4, fp4 89.0 tok/s with identical weight bytes — same ordering.
128-token bracket int8-1T 114.5 tok/s / 15.5 GB/s matches the old 120-127
table; 512-window context growth costs ~21%, so ratios are intra-run only.
Verdict: int8 is bandwidth-bound (saturates at 2T, +3.6% to 4T), 4-bit paths
are kernel-bound at every thread count (1T->4T ~1.7x scaling, DRAM idle,
never cash the 1.29x byte edge). Kernel bet unchanged on Zen 3 / no-VNNI:
int8 default everywhere, fp4 the 4-bit pick. 360M not fetched (724 MB BF16
shard); expect the amplified same pattern. Next levers: spec on DRAM-bound
int8-135M, int8-container upcast at load for cache-resident sizes (test-gated),
VNNI-class ISA. Compile Phase-0 stays hit:False/mode:generic (plumbing only,
execution generic) — correct until a PPL-gated execution change ships.
## Threading Results
Persistent native pool, row-sliced matvecs, pause-spin barriers (yield/sleep
only after long idle). Supra-Mini-8M INT8, 256-token decode, medians:
| Threads | Tokens/s | vs 1T |
| --- | ---: | ---: |
| 1 | 1,548 | 1.00x |
| 2 | 2,148 | 1.39x |
| 4 | 2,036 | 1.31x (regresses) |
FP32 (31.5 MB, DRAM-bound) shows no thread scaling: 998 -> 1,004 tok/s at 4T.
The Ryzen 5600 sweet spot is 2 threads for cache-sized quantized models;
beyond that, shared L3/LSU bandwidth saturates. gVisor-class sandboxes showed
~19 us sync round trips, so threading there may need larger models to pay off.
`cism bench/serve --threads N` sets the pool size; default remains 1.
## Speculative Decoding (Prompt Lookup, Greedy)
Implemented: n-gram drafter over full token history (up to 4-gram tails),
block verification via per-matrix GEMM (all T candidates share one weight
stream), greedy argmax acceptance with KV rollback, budget/EOS boundaries.
Speculative greedy output is bitwise identical to plain greedy decode (tested
across architectures, precisions, rejection paths, and budgets).
SmolLM2-135M (135 MB INT8 / 540 MB FP32 - exceeds the 32 MB L3), one thread,
256 tokens, repetitive prompt for high draft acceptance:
| Configuration | Tokens/s | Gain |
| --- | ---: | ---: |
| FP32 plain | 51 | 1.00x |
| FP32 spec_k=4 | 83 | 1.63x |
| INT8 plain | 106 | 1.00x |
| INT8 spec_k=4 | 130 | 1.23x |
Reason: speculation amortizes weight-bandwidth (one stream per verify block),
so gains appear when decoding is DRAM-bound - the FP32 regime. The INT8 kernel
is compute-bound (conversion-heavy SIMD), so amortizing bandwidth buys little
until integer kernels land. On the AVX-512 VNNI cloud target, integer kernels
should make INT8 bandwidth-bound and transfer the full speculation multiplier.
Two structural costs: the bonus token costs one extra single-token pass (ideal
(K+1)/2, measured 1.6x at K=4), and low-acceptance text can make speculation
slightly negative - disable it when drafts do not match the domain.
## Memory Control Verification
Real-model page locking and touch on this machine (Supra-Mini-8M INT8):
- `lock_pages()`: 8,022,880 of 8,022,784 weight bytes locked (page rounding).
- `touch()`: 8,022,880 bytes swept.
- Generation while keep-warm active and pages locked: normal.
- `unlock_pages()`: 6,268,768 bytes reported. Smaller than the lock count
because adjacent weight regions share 4 KiB pages; a page released by one
region cannot be released again by its neighbor. Counts are approximate byte
sums of successful OS calls, not exact page accounting.
x86 cannot pin cache lines; locking prevents OS paging, and sweeping only
refreshes residency while running. Suite at time of writing: 571 tests passed.
## Cache-Residency Probe Results
`cism cache-report` measures a warm full-weight scan, a cold scan after
evicting caches with a 96 MiB scrub buffer, and decode effective bandwidth
(weight bytes x decode tok/s). AMD Ryzen 5 5600, INT8, one thread:
| Model | Weights | Warm scan | Cold scan | Decode | Effective | Verdict |
| --- | --- | --- | --- | --- | --- | --- |
| Supra-Mini-8M | 8.0 MB | 28.5 GB/s | 25.1 GB/s | 1,646 tok/s | 13.2 GB/s | Compute-bound |
| SmolLM2-135M | 135.4 MB | 21.4 GB/s | 23.1 GB/s | 120 tok/s | 16.3 GB/s | Compute-bound |
Interpretation:
- Decode bandwidth is well below even the cold (post-eviction) scan rate in
both cases. The decoder does not saturate DRAM, so **cache residency is not
the current bottleneck; kernel efficiency is** (scalar conversion-heavy
quantized kernels, sequential prefill, single thread).
- The warm/cold scan ratio is small for these linear probes; hardware
prefetchers hide much of DRAM latency on sequential streams, so this probe
bounds bandwidth rather than proving residency. Definitive per-event counts
(LLC miss) require hardware PMU tooling such as WindowsPerf or AMD uProf.
- The earlier session's implied "SRAM 3x-6.5x multiplier" was not observed by
this probe on this hardware and remains unproven.
Speed work priorities implied by this data: quantized kernel quality, prefill
matrix-matrix, and multithreading before further cache-tuning.
## Measured Comparison Against PyTorch 2.14 And torch.compile
Supra-Mini-v5-8M, revision bb98e356, FP32 checkpoint, greedy, one torch/CISM
thread, 6-token prompt, 256-token fixed-length decode (EOS ignored both sides),
warmup absorbed, Dynamo counters verified zero recompilations during
measurement. Torch baselines use a KV cache (never use_cache=False).
Script: `scripts/bench_vs_torch.py`; CISM: `cism bench --max-tokens 256`.
| Decoder | Tokens/s | CISM FP32 (998.5) | CISM INT8 (1,716.4) |
| --- | ---: | ---: | ---: |
| HF generate(use_cache=True) eager | 158.9 | 6.28x | 10.80x |
| Manual eager loop + StaticCache | 195.8 | 5.10x | 8.77x |
| torch.compile default, dynamic=False | 260.1 | 3.84x | 6.60x |
| torch.compile default, dynamic=True | 224.9 | 4.44x | 7.63x |
The fair comparison is against torch.compile dynamic=False, its best
configuration here: **3.84x at matched FP32 precision, 6.60x for CISM INT8.**
Reasons, with evidence:
1. **Per-token dispatch.** torch.compile lowers eager's per-op Python/ATen
dispatch (159 -> 260 tok/s, +64%) but still launches dozens of generated
kernels through a Python-level decode loop each token. CISM executes the
whole token in one C++ call. Eager's effective bandwidth is
31.5 MB x 159 = 5.0 GB/s and compile's is 8.2 GB/s - both overhead-bound,
while CISM FP32 runs memory-bound at ~30 GB/s kernel rate.
2. **Memory traffic.** CISM INT8 reads 8.0 MB/token versus 31.5 MB for the
FP32 checkpoint, cutting the memory-bound floor ~4x on top of the
dispatch elimination.
3. **Quality.** CISM FP32 matches eager within 1e-5 max logit error
(established above). INT8 quality reporting remains a beta gate.
Caveats: compile mode is "default" (max-autotune not used); one compile
configuration initially crashed with an Inductor codegen bug on this model
(index-out-of-bounds in a generated kernel) and only ran after cache-reset
handling; multithreaded torch was not measured; prefill/TTFT is excluded from
decode rates. This is one cache-sized model on one machine - not yet the
agreed multi-model llama.cpp gate.
## 4-bit quantization scheme study (2026-09-12)
NumPy simulation mirroring the native hybrid-int4 quantizer (round-half-away,
blocks of 32 along the row, per-block scale; validated against the native
engine to 8 decimal places on all checkpoints). PPL over the fixed AutoBench
corpus (514 tokens, window 128), MLP matrices 4-bit, all other 2D matrices
per-row INT8, vectors FP32.
Machine: Ryzen 5 5600, local checkpoints, scripts/quant4_study.py.
| scheme (MLP element + scale) | bits/w | 8M ΔPPL | 100M ΔPPL | 135M ΔPPL |
|-------------------------------------|-------:|--------:|----------:|----------:|
| int4 + FP32 b32 (hybrid-int4) | 5.00 | +15.4% | +6.1% | +23.2% |
| int4 + E4M3 b32 | 4.25 | +15.5% | +6.0% | +26.7% |
| int4 + E8M0 b32 | 4.25 | +308% | +80.8% | +537% |
| E2M1 + E8M0 b32 (OCP MXFP4) | 4.25 | +318% | +92.4% | +323% |
| E2M1 + E4M3 b32 | 4.25 | +18.7% | +3.5% | +23.3% |
| E2M1 + E4M3 b16 (NVFP4 granularity) | 4.50 | +15.0% | +2.1% | +21.0% |
| int4 + E4M3 b16 | 4.50 | n/a | +8.1% | n/a |
| E2M1 + FP32 b32 | 5.00 | +13.7% | +4.4% | +19.9% |
| NF4 + FP32 b32 (QLoRA) | 5.00 | +44.1% | +8.8% | +32.0% |
| int4 + error feedback b32 | 5.00 | +36.5% | +16.9% | +80.9% |
Findings:
- Scale precision dominates element format; power-of-two (E8M0) scales are
catastrophic for both uniform and FP elements.
- E2M1 elements beat uniform INT4 at equal scale precision and block size
(100M model, b16+e4m3: +2.1% vs +8.1%).
- E2M1 + E4M3 b16 is the only scheme that improves on hybrid-int4 at every
model size while using fewer bits (4.50 vs 5.00): tied on 8M, 3x better on
100M, better on 135M.
- NF4 (QLoRA) loses badly on these trained small models; naive per-block
error feedback with a fixed scale inflates residuals and hurts everywhere.
Why hybrid-int4 decodes slower than int8 (measured, 1 thread, same
checkpoints): the nibble unpack is a ~10-instruction dependent chain per
block, so hybrid-int4 delivers only 7.6-8.1 GB/s of weight-equivalents vs
12.5-14.0 for int8 and 26.6 for fp32, despite moving 2.5x fewer bytes than
fp32. The E2M1 LUT kernel replaces that chain with a 16-entry table lookup
and is the reason to build the mode.
### Compile one-shot canonical parity (2026-09-14)
Phase-0 key plumbing is now cross-language exact. Python
`compile_cache.canonical_model_config()` builds the native 12-field pipe
`model_type|hidden|intermediate|layers|heads|kv_heads|head_dim|vocab|
context|hexfloat(eps)|hexfloat(rope_theta)|tied` with FP32-round-trip
`struct.pack('f')` hexfloat, matching `Model::refresh_compile_key`
bit-for-bit. Verified: tiny Llama 16/32/1/4/2/4/32/32 probe Python
canonical equals `Model.info["compiled"]["canonical"]`, and
`native.compile_key(python_canonical, precision, act, threads)` equals
`Model.info["compiled"]["key"]` (`d4547092c73a7aac` on the probe).
Legacy `canonical_config()` sorted-JSON remains for early tests only.
Kept as a stable model fingerprint (zero runtime cost, tested); the
packed-artifact load idea is dropped with the project (see below).
### SmolLM2-135M full AutoBench matrix (2026-09-13, 512 tokens, act fp32)
Local file `autobench_135M_512.json` (untracked measurement, torch
baselines disabled by policy `include_compile=False`): 4 precisions x
threads (1,2,4) x spec_k (0,2,4,8,12,16) = 72 decode cells, PPL once per
precision (514 scored tokens, window 128). Revision `93efa2f0` (report
does not yet stamp `revision`; next run should copy
`engine.info["revision"]`). All `fits_l3=false` (L3 cache size 32.0 MB).
Weight storage in megabytes: int8 129.16 MB (135,439,104 bytes),
hybrid-int4 100.27 MB (105,141,504 bytes), hybrid-fp4 95.52 MB
(100,164,864 bytes), fp32 513.13 MB.
Decode rate in tokens/s at spec_k=0 (effective bandwidth in GB/s =
weight bytes x tok/s):
| Precision | 1 thread tok/s (GB/s) | 2 threads tok/s (GB/s) | 4 threads tok/s (GB/s) | Best tok/s (threads, spec_k) | PPL (+vs int8) |
| --- | --- | --- | --- | --- | --- |
| fp32 | 43.78 (23.6) | 57.44 (30.9) | 55.34 (29.8) | 106.41 (4T, K=8) | 34.15 |
| int8 | 91.53 (12.4) | 124.82 (16.9) | 130.03 (17.6) | 147.19 (4T, K=2) | 34.74 (--) |
| hybrid-int4 | 71.93 (7.6) | 106.54 (11.2) | 115.69 (12.2) | 140.59 (4T, K=16) | 42.07 (+21.1%) |
| hybrid-fp4 | 71.85 (7.2) | 102.10 (10.2) | 117.51 (11.8) | 132.90 (4T, K=2) | 41.32 (+18.9%) |
Ratios int8 vs 4-bit at spec 0: 1T 1.27x, 2T 1.17-1.22x, 4T
1.11-1.12x. Scaling 1T to 4T: int8 1.42x but 2T to 4T only +4.2%
(saturates, bandwidth-bound); int4 1.61x, fp4 1.64x, still climbing
with DRAM idle (kernel-bound nibble unpack, no VNNI). Verdict unchanged:
int8 default everywhere, fp4 the 4-bit pick (better quality at fewer
bytes, speed tied within ±15-20% multithread noise; ratios
single-harness only).
### Spec K-curve on 135M DRAM-bound int8 (512 tokens, greedy)
Accept rate in percent is deterministic across threads per K
(88.0% at K=2 decaying to 72.4% at K=16).
| Threads | K=0 tok/s | K=2 tok/s | K=4 tok/s | K=8 tok/s | K=12 tok/s | K=16 tok/s | Peak gain vs K=0 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 1T | 91.53 | 98.42 | 100.47 | 99.45 | 96.14 | 93.67 | K=4, +9.8% |
| 2T | 124.82 | 136.05 | 132.62 | 131.34 | 129.72 | 126.94 | K=2, +9.0% |
| 4T | 130.03 | 147.19 | 146.55 | 135.13 | 136.81 | 143.53 | K=2, +13.2% |
Global best: int8 147.19 tok/s at 4 threads, spec_k=2. Contrast fp32
+73-92% peaking K=8 (bandwidth-bound, block verify amortizes DRAM) vs
4-bit +5-21% erratic (kernel-bound). Peak at small K because acceptance
decays with draft length while GEMM verify cost grows linearly, plus the
bonus-token extra pass. Prompt-lookup drafter (up to 4-gram, greedy-only,
bitwise identical) helps repetitive prompts only; it does not fix the
135M slow complaint (still ~147 tok/s vs ~1,600 tok/s cache-resident 8M).
Default 135M int8: 4 threads, spec_k=2 (2-thread servers also K=2).
### trust_remote_code custom-arch assessment (2026-09-14)
`trust_remote_code=True` authorizes tokenizer Python only, never custom
modeling operators; trusted code itself is not sandboxed. Enforcement:
loader `_ARCHITECTURES={"llama","qwen3"}`, `validate_config` rejects
non-llama/qwen3 `model_type` and any `auto_map` beyond `AutoTokenizer`
even with trust enabled, cross-repo refs rejected, `_expected_shapes`
plus `_load_weights` reject extra/missing tensors, native
`Config::validate`/`parse_config` reject bias/MoE/sliding-window/
non-silu/non-default-RoPE. Custom Negative/CMA/Ember2/Spark2A/Blaze
targets (ROADMAP gate 4: recurrence, routing, lane mixing, Engram state)
need new native Config fields, operators, KV layouts, quant definitions,
and PPL/parity gates per arch — large work, not a flag flip. Keep
explicit `UnsupportedModelError`; server stays opt-in
`--allow-remote-code-load` (403 otherwise).
### Compile project dropped (2026-09-14, verdict: correctly built, not worth it)
Target was one-time load cost for faster batch-1 CPU decode (generic
AVX2, PPL-gated). Verdict: the plumbing is correct and tested, but the
goal is unreachable on this CPU, so the project is dropped and its
speedup fiction removed:
- What was real: cross-language canonical parity (bit-exact, tested),
hoisted dispatch loops (same kernels/order, bitwise identical, kept as
plain code), manifest helpers (kept, unused by load path).
- What was fiction: manifest-hit as "compile hit" (bookkeeping, not a
faster layout), `compiled.hit/mode` in Engine info and tester cards
(removed), any decode tok/s attributed to compile (zero: mode was
generic everywhere; the +13.2% is spec K=2).
- Why not worth it: remaining ~26 µops/32 are structural on AVX2/no-VNNI;
int8 is bandwidth-bound (2T saturates), 4-bit kernel-bound; fresh 512
runs show hoisting moved 4-bit ~+23% but int8 K0 flat (-0.9%),
confounded by idle-machine effect — no clean decode win to claim;
inductor costs 107 s warmup for 32 tok/s against CISM's 190 (6-8x
lead from the pre-existing single-call C++ path + quant, not compile).
VNNI/VBMI or integer-MAC kernels would reopen the question; until
then the line is closed. Fingerprint key + hoisted loops stay; hit/UI/
future-speedup prose goes.
### Fresh 512-token 4T confirmation (SmolLM2-135M, all precisions, spec sweep)
Re-ran 24 cells (4 precisions x 4T x spec 0,2,4,8,12,16), 512 tokens,
127.3 s wall, torch off. Accept% deterministic and identical to the old
file; PPL/weights identical (fp32 34.15, int8 34.74, int4 42.07, fp4
41.32; 513.13/129.16/100.27/95.52 MB, all fits_l3=false). Absolute tok/s
ran +11-30% hotter than the old file (idle machine); cells >25% off are
marked SUSPECT — use this table, do not mix files.
| Precision | K=0 | K=2 | K=4 | K=8 | K=12 | K=16 | Best |
| --- | --- | --- | --- | --- | --- | --- | --- |
| fp32 | 58.77 | 72.89 | 101.92 | 125.09 | 122.38 | 121.18 | K=8 125.09 |
| int8 | 128.90 | 159.99 | 162.64 | 171.76 | 171.37 | 177.03 (SUSPECT vs old) | K=16 177.03 |
| hybrid-int4 | 142.63 | 155.69 | 151.97 | 159.86 | 160.29 | 156.40 | K=12 160.29 |
| hybrid-fp4 | 144.58 | 150.73 | 153.96 | 160.41 | 155.88 | 153.82 | K=8 160.41 |
Accept%: fp32 96.5/94.0/89.9/88.1/85.5; int8 88.0/83.4/77.1/73.0/72.4;
int4 90.1/84.4/79.1/75.6/75.3; fp4 84.1/76.6/70.3/66.7/67.6 (K=2..16).
Note 4-bit now leads int8 at K=0 (142.63/144.58 vs 128.90) in this
run — confounded by machine state (old file had int8 leading), so no
compile attribution claimed without a same-state A/B. fp4 vs int4: tied
within noise at every K (max gap 3.7%), fp4 keeps the quality edge
(+18.9% vs +21.1% PPL), so fp4 stays the 4-bit pick.
### 4T K=2 vs regular baselines (SmolLM2-135M int8, 512 tokens)
Global best cell is int8 4 threads spec_k=2 at 147.19 tok/s (accept
87.96%). Regular baselines from the same `autobench_135M_512.json` run
(`include_torch=False`, so no HF eager/compile in this file):
| Baseline (4T) | Tok/s | vs 4T K=2 (147.19) |
| --- | --- | --- |
| int8 K=0 (same weights, no spec) | 130.03 | +13.2% with spec |
| int8 2T K=0 | 124.82 | +17.9% (threads+spec) |
| int8 1T K=0 | 91.53 | +60.8% |
| fp32 K=0 | 55.34 | 2.66x |
| fp32 best (K=8) | 106.41 | 1.38x |
| hybrid-int4 K=0 | 115.69 | 1.27x |
| hybrid-int4 best (K=16) | 140.59 | 1.05x |
| hybrid-fp4 K=0 | 117.51 | 1.25x |
| hybrid-fp4 best (K=2) | 132.90 | 1.11x |
PPL in same run: int8 34.74, fp32 34.15, int4 42.07, fp4 41.32.
Spec pays (+13.2%) because 135M int8 spills (129.16 MB vs 32 MB L3):
block verify streams weights once per block via GEMM. It does not make
135M fast in absolute terms (still ~147 vs ~1,600 cache-resident 8M).
### HF eager/compile baselines (SmolLM2-135M, 512 tokens, 4T, same harness)
One `run_benchmark` (`src/cism/autobench.py:223`) with
`include_torch=True, include_compile=True`, `max_new_tokens=512`,
threads 4, spec sweep 0-16, rev `2c4d40c`. Whole sweep ran ~4-15% under
the fresh reference (mild thermal throttle); HF rows are baselines, not
throttle-sensitive conclusions.
| Decoder (512 tok, 4T) | Tok/s | PPL (514 scored tokens) |
| --- | --- | --- |
| CISM int8 K=0 | 131.88 | 34.74 (nll 3.5480, +1.73% vs eager) |
| CISM int8 best K=16 | 167.23 | same |
| CISM fp32 K=0 / best K=8 | 53.32 / 114.67 | 34.15 (+0.002%) |
| CISM hybrid-int4 K=0 / best K=8 | 137.20 / 153.07 | 42.07 (+23.18%) |
| CISM hybrid-fp4 K=0 / best K=16 | 137.38 / 150.25 | 41.32 (+21.00%) |
| HF eager fp32 | 23.79 (21.52 s) | 34.15 (nll 3.5308) |
| HF dynamic int8 | 19.74 (25.94 s) | 81.12 (broken, +137.5%) |
| HF torch.compile fp32 | 25.18 (20.33 s, warmup 147.3 s) | not scored |
| HF compile int8 | 16.86 (30.38 s) | not scored |
| torchao int4 | skipped (not installed) | -- |
Speedups (same harness): CISM int8 K=0 **5.54x eager / 5.24x compile**;
best K=16 **7.03x / 6.64x**; fp32 K=0 2.24x/2.12x, best K=8 4.82x/4.55x.
Compile barely beats eager (+5.9%) after a 147 s warmup; dynamic int8 is
slower than fp32 AND destroyed (PPL 81). Weights: 513.13/129.16/100.27/
95.52 MB, all fits_l3=false (L3 32 MB).
### Surjo hybrid native Phase1 (loader+fp32 decode)
Scope: SurjoLabs Surjo hybrid import plus native fp32 decode runs. No
Surjo tok/s, no Surjo PPL, no speedup claim in this section. Reference
dense numbers below are context only, single-harness, ±15-20%
multithread noise (see 135M matrix notes); do not mix files or
attribute cross-run deltas to Surjo.
Config fingerprint (Surjo-50m, `tests/test_surjo.py:20`): `model_type`
surjo / `SurjoForCausalLM`, vocab 32768, hidden 512, intermediate
1536, layers 10, context 2048, prelude 1 / recurrent 8 / coda 1,
groups 2 / passes 2 / gdn_per_xsa 3, heads 8 / kv 4 / head_dim 64,
gdn_v_heads 8 / gdn_k_dim 64 / gdn_v_dim 64 / kernel 4,
`gdn_allow_neg_eigval` false, `xsa_projection` true, bias false,
dropout 0.0, `hidden_act` silu, eps 1e-5, theta 10000.0, default RoPE
(`partial_rotary_factor` 1.0), tied true, `layer_types`
`[F,L,L,L,F,L,L,L,F,F]`, `auto_map` pinned to
`configuration_surjo`/`modeling_surjo` (never executed). Base
checkpoints ship `use_cache=false`; `import_model`
(`src/cism/loader.py:666`) preserves the flag verbatim and it does
not disable native caching (XSA KV slots + GDN recurrent/conv state
always allocate).
Checkpoint header: 178xF32 tensors in `model.safetensors`, tied
`lm_head` omitted, no rotary `inv_freq` buffers (RoPE is computed in
`forward()`). `_expected_shapes_surjo` length is 179 including the
tied-head alias entry; spot-checked in
`test_surjo_50m_expected_shapes_match_header` (embed `(32768,512)`,
XSA `q (512,512)` / `k (256,512)` / `q_norm (64,)`, GDN `q/k (512,512)`
/ `v (512,512)` / conv `(512,1,4)` / `f0 (64,512)` / `f1 (512,64)` /
`A_log (8,)` / `dt_bias (512,)` / `g1.bias (512,)` / `o_norm (64,)`).
Loader (`src/cism/loader.py`): `validate_surjo_config`,
`surjo_exec_plan` (18 steps: 1 prelude + 2 passes x 2 groups x
(3 GDN + 1 XSA) + 1 coda; XSA slots `[0..5]`, 12 GDN `(layer,pass)`
steps), `surjo_num_xsa_slots` (=6), `_expected_shapes_surjo:393`,
`_load_weights` Surjo branch (no buffers, strict shape/dtype/finite
checks, tied-head alias by identity), `import_model` dispatch
(`model_type=="surjo"` -> Surjo path, else dense). `_SURJO_SUPPORT`
string still reads "next phase" in this snapshot; Phase1 below is
what landed natively.
Native (`native/runtime.{hpp,cpp}`, `native/bindings.cpp:468`,
`src/cism/engine.py:80` dispatch to `_native.SurjoModel`):
`SurjoModel`/`SurjoSession` fp32 decode. XSA: 6 slots, KV
`[slots,capacity,kv_width]` with `kv_width=256`; at max context 2048
that is 4 MiB/slot (K+V), 24 MiB total (session-sized by
`prompt+max_new`, so 24 MiB is the upper bound). GDN: 12 states
(6 GDN layers x 2 passes), each `S` is `H*K*V=8*64*64=32768` floats =
128 KiB (12x128 KiB = 1.5 MiB), plus depthwise-conv raw FIFO
`(kernel-1)=3`: q/k/v FIFOs ~72 KiB each, ~0.21 MiB total. MLP is
ClampedMLP (gate clamped to `[-15,15]` then `silu*up`,
`runtime.cpp:2010`). `spec_k` is rejected for Surjo
(`next_tokens` throws unless `spec_k==0`; `Engine.stream(spec_k=2)`
raises `ValueError`, covered by `test_surjo_engine_decode_runs`).
Prefill is sequential per-token (`forward_tokens` loop); `logits` /
`nll` teacher-forcing paths exist but are unscored against HF so far.
Quant mapping reuse, no new format: `protected_storage` (embed, head,
all XSA/GDN projections incl. `f0/f1/g0/g1/o_proj`) is fp32 or int8;
MLP `gate/up/down` follow the dense hybrid rule (hybrid-int4 -> int4,
hybrid-fp4 -> fp4). Same `Matrix` kernels/scratch/act_q8 path as
dense; only norms/vectors/conv/A_log/dt_bias stay fp32.
Parity gate: dense bar is max abs logit error <1e-4 vs HF eager.
Surjo fp32 decode/smoke runs (tiny shapes fixture, tied-head
identity, `Engine.logits/generate` finite, spec rejection). HF
`modeling_surjo.py` parity script and Surjo PPL are pending — no
parity number claimed here.
Batch/paging: Phase0 files unchanged. Single-request serialized
`Engine` generation lock, `create_session`/`next_tokens`,
cancel/finish, `lock_pages`/`touch`/`scan` weight-memory controls
are reused as-is. No continuous batching, no paged-KV layout, no new
server path for Surjo.
FLA CPU verdict: decode needs no Flash-Linear-Attention (single-token
GDN recurrence is O(1) state + per-slot causal XSA). Chunkwise/FLA
prefill is an optional long-prompt throughput optimization only, not
correctness; current sequential prefill stands.
Dense context (same `autobench_135M_512.json`, `include_torch=False`,
NOT Surjo results): 4T K=0 fp32 55.34 / int8 130.03 / hybrid-int4
115.69 / hybrid-fp4 117.51 tok/s; best cells fp32 106.41 (K=8), int8
147.19 (K=2, accept 87.96%), hybrid-int4 140.59 (K=16), hybrid-fp4
132.90 (K=2); PPL fp32 34.15 / int8 34.74 / int4 42.07 / fp4 41.32.
Ratios single-harness only; 4-bit vs int8 gaps sit inside ±15-20%
noise. Surjo inherits none of these numbers.
### Surjo-50m full AutoBench (Phase2, 2026-09-15, 4T K0)
File `autobench_surjo50m_512.json` (local snapshot
`35c9fa6e`, 512-token HF window, CISM greedy K0 only, act fp32,
threads 1/2/4). Weights (engine.info): fp32 205.21 MB, int8
51.84 MB, hybrid-int4 43.27 MB, hybrid-fp4 41.86 MB; L3 32.0 MB,
all `fits_l3=false`. PPL over the fixed AutoBench corpus (527 scored
tokens): fp32 59.29 / int8 59.46 (+0.30% vs HF eager) / hybrid-int4
62.90 (+6.10%) / hybrid-fp4 61.80 (+4.25%). CISM int8/fp32 PPL matches
HF eager fp32 59.28 to 0.3% — quant quality gate holds at int8, 4-bit
pays +4-6%.
CISM decode at 4T K0 (greedy; CISM stopped on EOS at 43-48 tokens,
HF forced 512 via min_new_tokens — lengths differ, ratios indicative):
| Precision | 1T tok/s | 2T tok/s | 4T tok/s | PPL |
| --- | --- | --- | --- | --- |
| fp32 | 62.18 | 81.03 | 81.06 | 59.29 |
| int8 | 129.86 | 158.39 | 207.80 | 59.46 |
| hybrid-int4 | 109.05 | 159.04 | 175.49 | 62.90 |
| hybrid-fp4 | 105.29 | 153.13 | 171.31 | 61.80 |
HF baselines (same harness, FP32 CPU, same thread count):
| Decoder (4T) | Tok/s | PPL |
| --- | --- | --- |
| HF eager fp32 | 26.27 (1T 28.08 / 2T 27.74 — inverse scaling, MT noise) | 59.28 |
| HF dynamic int8 | 18.18 (main cell 19.93; sweep 17.42/18.95/18.18) | 66.35 destroyed (+11.9%) |
| HF torch.compile fp32 | ~~31.72~~ INVALID — StaticCache path, see below | not scored |
| HF compile int8 | 16.73 (same StaticCache invalidity) | not scored |
Speedups at 4T K0 vs HF eager 26.27: int8 7.91x, hybrid-int4 6.68x,
hybrid-fp4 6.52x, fp32 3.09x. `speedup_vs_compile` in the JSON divides
by the invalid ~~31.72~~ and is void — do not cite.
Compile correction: `modeling_surjo.py:645` discards non-SurjoCache
past (`if not isinstance(past, SurjoCache): past = None`), so the
autobench StaticCache loop allocated a fresh cache per forward call and
state never carried — ~~31.72 tok/s~~ is not a Surjo compile number.
`src/cism/autobench.py` now detects `surjo_mode` and threads returned
`past_key_values` (SurjoCache) under `torch.compile(dynamic=True)`,
recording `{"skipped": "StaticCache invalid for Surjo"}` if the
threaded path cannot run — never a StaticCache tok/s. Correct path
proven by `scripts/bench_compiler.py`: threaded SurjoCache +
`dynamic=True` ~41 tok/s (compile warmup 43-119 s across runs;
189.3 s in the old JSON was the invalid-path warmup);
`dynamic=False` unusable (~0.06 tok/s, recompiles per token as XSA KV
length grows). torchao int4 blocked on Windows (no mslk wheel —
`Requires mslk >= 1.0.0`); torchao int8wo is the CPU-working control
only. FLA CPU dead here (Triton needs GPU; `fla.ops.gdn2` chunk/fused
probe does not run on CPU) — decode needs no FLA anyway (single-token
GDN recurrence is O(1) state + per-slot causal XSA).
Spec: unsupported for Surjo hybrid (GDN state destructive) — K0 only
(`spec_note` in JSON); K2≈K0 where probed, no speculation gain claimed.
Native `SurjoSession` rejects `spec_k!=0`; `Engine.stream(spec_k=2)`
raises.
Rigor: this JSON predates the harness (`matrix` rows lack
`rep_rates`/`thermal_throttled`/`cpu_snapshot`). Current
`src/cism/autobench.py:_timed` reports 1 warmup + median-of-`reps`
(`--reps`/`--runs` wired through `cism bench`; default 3) with min/max/
std, `thermal_throttled` when max-min spread exceeds 15% (wider than the
±15-20% MT noise band), and best-effort `cpu_snapshot` (clock/temp;
`{}` when unavailable, never fails the run). Next Surjo rerun picks up
medians + thermal flags automatically; do not mix old single-sample
cells with new median cells.
### Surjo-50m rerun on current code (2026-09-15 night, 72 cells, reps=3)
`autobench_surjo50m_512.json` overwritten with current native
(GDN-OPT + XSA-blocked + `forward_block` + spec verify). Threads
1/2/4 x spec 0/2/4/8/12/16, 512tok, act fp32, median-of-3, all
`thermal_throttled=false` except two fp32-4T spec cells. CISM EOS-stops
43-48tok vs HF forced-512 — same caveat as before.
K0 medians: fp32 60.32/84.73/88.86, int8 141.76/197.34/234.87,
int4 114.59/169.65/197.28, fp4 113.54/169.98/214.61. Best cells:
fp32 K2 92.66, int8 K4 236.91, int4 K4 201.33, fp4 K0 214.61.
Spec accept on bench prompt is 0.0 for fp32/int8 (no gain, no loss
beyond noise); int4 6.25%->1.7%, fp4 8.3%->2.7% decaying with K, and
spec K>0 is slower than K0 there (verify cost, no amortization win
on this prompt). PPL unchanged: 59.29/59.46/62.90/61.80.
HF equal-opportunity, same harness, threads swept, median-of-3:
eager 31.22/30.43/29.50 (4T 29.50 PPL 59.28),
dynamic-int8 20.27/22.59/22.37 (4T 22.69 PPL 66.35),
compile threaded-SurjoCache dynamic=True 38.77 warmup 100.6s,
compile-int8 21.10, torchao-int4 failed `Requires mslk>=1.0.0`.
Speedups vs eager 29.50: int8 K0 7.96x, best K4 8.03x;
vs compile 38.77: K0 6.06x. Old invalid StaticCache 31.72 is
superseded — do not cite it.
### Supra-50M-Base full AutoBench (dense Llama, 2026-09-15 night)
File `autobench_supra50M_512.json`, Hub rev
`521bfd3d3901fbeefa943a2f3d461a9d193c52ec`, dense Llama
(hidden 512, layers 12, heads 8/kv 4/head_dim 64, intermediate
1408, vocab 32000, ctx 1024, tied, `use_cache=false` preserved).
Threads 1/2/4 x spec 0/2/4/8/12/16 = 72 cells, 512tok, act
fp32, median-of-3. Weights: fp32 ~209MB / int8 ~53MB class
(all `fits_l3=false`). PPL (527 scored tokens): fp32 63.32 /
int8 63.79 (+0.7%) / int4 65.36 (+3.2%) / fp4 68.90 (+8.8%).
CISM K0 medians: fp32 108.29/131.54/123.55, int8
214.81/284.36/292.58, int4 178.04/247.49/250.87, fp4
171.67/234.54/251.38. Best cells: fp32 K8 268.19 (2T),
int8 K12 349.10 (4T), int4 K2 257.43, fp4 K8 274.25.
Accept on bench prompt is 100% for fp32/int8/fp4 at every K
(hence spec pays: int8 4T K0 292.58 -> K12 349.10, +19.3%;
fp32 K0 123.55 -> K8 268.19, +117%); int4 86%->70%.
HF same harness, threads swept, median-of-3: eager
48.59/47.90/43.61 (4T 43.61 PPL 63.32), dynamic-int8
37.70/40.49/37.78 (4T 34.91 PPL 101.68 destroyed),
compile StaticCache dynamic=False 55.23 warmup 143.0s,
compile-int8 33.14, torchao-int4 `Requires mslk>=1.0.0`.
Speedups vs eager 43.61: int8 K0 6.71x, best K12 8.00x;
vs compile 55.23: K0 5.30x, best 6.32x.
### FWKV Myosotis-1-base full support + AutoBench (2026-09-15 night)
Files inspected in snapshot `031fe84a34f667bc8bfff7dc3f8a77ef3c17efb2`:
`config.json` (model_type fwkv, d_model 768, d_emb 192,
n_layers 13, ffn_mult 4, vocab 50257, wkv_floor 0.1, tied),
`configuration_fwkv.py` (FWKVConfig + HF aliases),
`modeling_fwkv.py` (FactorizedTiedHead, FWKVBlock with
proj_k/v/r/out + W recurrence + GELU FFN + 2x LayerNorm,
`_supports_cache_class=False`, tuple-state threading),
`model.safetensors` (147 F32 tensors, no rotary buffers).
Native support landed, mirroring the Surjo vertical slice:
`loader.validate_fwkv_config` + `_expected_shapes_fwkv` (147
shapes) + `_load_weights` fwkv branch (also fixed an eager
`config["head_dim"]` default crash affecting any arch without
head_dim) + `import_model` dispatch; `native` FwkvConfig /
FwkvModel / FwkvSession (WKV recurrence, exact GELU,
LayerNorm eps 1e-5, factorized head, fp32/int8/int4/fp4 reuse
via Matrix, `forward_block` batched projections + sequential
recurrence, spec verify with ~40KB state checkpoint/restore);
`bindings` FwkvModel/FwkvSession; `engine` dispatch;
`tests/test_fwkv.py` 11 passed; `scripts/parity_fwkv.py`:
tiny 1.19e-07, live vs HF fp32 5.72e-06 (<1e-4 OK).
Autobench: `trust_remote_code=True` for fwkv HF loads,
threaded-FWKV-state compile path (StaticCache invalid, same
class of bug as Surjo), `_hf_context` fallback to `seq_len`
(FWKVConfig has no `max_position_embeddings`).
File `autobench_fwkv_512.json`, 72 cells, reps=3 (v4 drafter rerun,
2026-09-15 late night; supersedes the v1-drafter run below). Weights:
fp32 425.38MB / int8 107.64MB. PPL: fp32 156.68 (= HF eager
156.68, gate holds) / int8 156.80 (+0.07%) / int4 156.63
(-0.04%) / fp4 158.53 (+1.18%). High absolute PPL is base
model quality, faithfully reproduced.
CISM K0: fp32 64.65/93.73/87.24, int8 169.01/268.22/342.27,
int4 121.54/215.32/284.66, fp4 119.21/205.16/303.52. Best:
fp32 K16 177.34, int8 K4 382.74, int4 K2 290.19, fp4 K2
305.98. Accept (bench prompt): fp32 85%->61%, int8
85%->62%, int4 77%->45%, fp4 74%->44% (K2->K16). Spec pays
on fp32/int8: int8 4T K0 342.27 -> K4 382.74 (+11.8%);
fp32 K0 87.24 -> K16 177.34 (+103%); 4-bit int4 still
peaks near K0-K2, fp4 peaks K2.
HF same harness, threads swept, median-of-3: eager
42.72/45.59/43.42 (PPL 156.68), blocks-only dynamic-int8
34.74/43.88/51.26 (4T 51.22 PPL 158.42 +1.1% ok — fixed by
quantizing `blocks` Linears only; full-model quantize fails
on `FactorizedTiedHead.proj.weight.t()`), compile
threaded-FWKV-state dynamic=True 49.73 warmup 45.9s,
compile-int8 same path 46.93 warmup 36.2s, torchao-int4
`Requires mslk>=1.0.0` (no Windows wheel). Speedups vs
eager 43.42: int8 K0 7.88x, best K4 8.81x; vs compile
49.73: K0 6.88x, best 7.70x. Previous run (v1 drafter):
K0 59.82/83.70/87.54 etc., best int8 K8 359.97 — v4 gains
come from restored accept (19%->85% @K2).
### Spec drafter v4 + doom-loop check (2026-09-15 late night)
v1 (most-recent single match, verbatim tail truncated at history
end): Surjo bench 0% (correct text, nothing to match), Supra
bench 100% (degenerate loop flatters it). v2 (pure frequency
vote) collapsed on drifting text (FWKV bench int8 K2
85%->19%, K8 69%->9%). v3 (recency-decayed vote) same failure.
v4 (landed): most-recent strictly-inside occurrence per
position + iterative extension (fixes v1 truncation) +
rejected-block latch REMOVED (the 32-propose/<1/3-accept
latch fired during the short-history warmup and blinded the
session to later 85%-accept regions; rejected blocks stream
weights once per block via forward_block, so worst case on
0%-accept text is ~K0 speed). Shared by dense/Surjo/FWKV;
`tests/test_native.py::test_speculative_draft_hits_and_misses`
reimplemented against v4; suite 759 passed; K0==Kx exact
everywhere probed.
Doom-loop probe (greedy, 128 tok, fp32+int8, 3 prompts):
Surjo stops coherently on bench/code prompts (48/15 tok,
uniq 0.73/0.93, loop 0) but loops story prompts (128 tok,
uniq 0.09-0.16, 5/11-period loops); Supra loops ALL prompts
(128 tok, uniq 0.07-0.11, 7/9/12-period loops — bench,
story, AND code); FWKV never hard-loops bench/story
(uniq 0.28-0.54, loop 0; int8 bench drifts into a 10-period
body/brain cycle) but degenerates to 128 spaces on code
(uniq 0.02 both precisions). So Supra's 100% spec accept is
a looping artifact, Surjo's 0% is healthy text, FWKV sits
between. No repetition penalty exists anywhere (greedy
only); spec accept must always be read next to uniqueness.
### Portable kernels: Intel/AMD tiers + ARM64 phones (2026-09-16)
Tier ladder (runtime-dispatched, scalar fallback always works):
`scalar` (any CPU) < `avx2`+FMA (x86-64 baseline, all current
numbers) < `avx-vnni` (Zen4/Zen5, Alder Lake+: int8 VNNI fast
path) < `avx512-vnni` (server AVX512) / `neon` (ARM64 phones,
Apple Silicon, Raspberry Pi). New files: `native/kernels_vnni.cpp`
(dpbusd int8xint8, int8-activation ABI via `quantize_row_i8`),
`native/kernels_neon.cpp` (fp32/int8/int4/fp4 + uzp permute),
CPUID detection in `native/kernels.cpp` (AVX-VNNI = leaf7:1
EAX[4], AVX512-VNNI = leaf7:0 EBX[11] + foundation + ZMM state;
MSVC VNNI TU is EVEX so it dispatches only on AVX512-VNNI,
GCC VEX on AVX-VNNI), per-file codegen flags in CMakeLists
(dispatcher/runtime stay generic). `kernel_name()` keeps the
fp32 family; new `kernel_variant()` + `has_avx_vnni` /
`has_avx512_vnni` / `has_neon` bindings; `cpu_features` gains
`avx_vnni`; `Model.info["kernels"]` gains `variant`.
`tests/test_kernels_portable.py` (5 passed): tier reporting,
flag consistency, pre-VNNI fallback untouched, i8 quantizer
round-trip bounds, dpbusd bias-identity in numpy.
Proven on Ryzen 5 5600: builds clean (MSVC /arch:AVX512 VNNI TU
accepted), `kernel avx2 / variant avx2 / vnni False`,
suite 764 passed (759 + 5 new) — zero behavior change, as
designed (VNNI selector is null here so Matrix keeps the int16
Q8 / fp32-activation paths byte-for-byte). Pending hardware CI:
VNNI tok/s + int8-activation PPL re-gate on Zen4/ADL+ (expect
prefill/spec-verify wins ~2-3x integer MACs; decode stays
bandwidth-bound so batch-1 gains concentrate in K>0), NEON
tok/s + parity on ARM64 hardware (expects the usual
Silicon/Android numbers; tails fall back to scalar by
construction). No speedup is claimed for either until measured.