Spaces:
Sleeping
Sleeping
|
Download docs/VALIDATION.md from spitfire4794/test1111111: direct link, hf CLI and curl.
- Browser
- Download file 47.2 kB
-
https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/VALIDATION.md
- Command line
-
hf download hf://spaces/spitfire4794/test1111111/docs/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/spaces/spitfire4794/test1111111/resolve/main/docs/VALIDATION.md
47.2 kB
| # Development Snapshot Validation | |
| Date: 2026-09-08. This is not beta certification or a llama.cpp comparison. | |
| ## Environment | |
| - Windows x64, AMD Ryzen 5 5600, 6 physical cores / 12 logical processors. | |
| - CPython 3.13, MSVC Visual Studio Build Tools 2022 17.14.7. | |
| - One native inference thread, runtime-dispatched AVX2. | |
| - Installed wheel: cism 0.1.0.dev0. | |
| - NumPy 2.5.3, Transformers 5.16.1, PyTorch 2.14.0+cpu for reference tests only. | |
| - FastAPI 0.141.1, Uvicorn 0.52.4, OpenAI Python SDK 2.54.0. | |
| ## Tests | |
| `python -m pytest -q`: **555 passed**, two upstream TestClient deprecation | |
| warnings. Includes independent NumPy numerical references, Transformers tiny | |
| model parity, packed quantization, invalid checkpoint/config handling, generation | |
| metadata, cancellation, queue limits, SSE, official SDK, and a live Uvicorn TCP | |
| roundtrip using real native inference. | |
| `python -m pip check`: no broken requirements. | |
| Both scalar-only and AVX2 native builds were also tested separately during kernel | |
| development. Linux, sanitizers, long-running load tests, and full-context real | |
| checkpoint certification remain outstanding. | |
| ## Real Checkpoints | |
| Compared default FP32 eager HF computation with native last-token logits and | |
| four-token greedy continuations at prompt lengths 4, 8, and 16. All tested | |
| continuations agreed; maximum absolute logit error was below 0.0001. | |
| - SupraLabs/Supra-Mini-v5-8M: FP32, INT8, hybrid INT4. | |
| - veyra-ai/Veyra2-Blueberry-10M-Base: FP32. | |
| INT8/INT4 references used reconstructed quantized weights, not original FP32 | |
| weights. These results establish implementation parity only, not a PPL quality | |
| budget or task-quality equivalence. | |
| ## Local Throughput | |
| ```text | |
| cism bench SupraLabs/Supra-Mini-v5-8M --precision int8 --runs 5 --max-tokens 128 | |
| ``` | |
| Model revision: `bb98e3566a5ab3f24be16ee8db7f06f0cb884eb5`. | |
| Packed weight storage including scales/norms: 8,022,784 bytes. | |
| - Prompt: 6 tokens; generation: 128 tokens; EOS ignored for fixed-length timing. | |
| - One warmup, then five runs; native generation through coarse Python bindings. | |
| - Median decode throughput excluding prefill/first token: **1,622.86 tokens/s**. | |
| - Runs: 1,595.72 / 1,632.20 / 1,624.86 / 1,603.32 / 1,622.86 tokens/s. | |
| - Median prefill plus first token: **3.744 ms**. | |
| No tokenization, detokenization, HTTP, or queue latency is included in decode | |
| timing. These are growing-context token-generation results, not weight sweeps. | |
| CPU affinity and clocks were not locked. No comparison with a tuned llama.cpp | |
| baseline has been run; the required speed advantage is not established. | |
| ## Why These Numbers Differ From The Original Experiments | |
| The earlier session results (hundreds of thousands of tok/s) are not comparable | |
| to these, for verifiable reasons on both sides: | |
| - The supplied C harness (`decode_int8`/`decode_hybrid`) looped over flat weight | |
| buffers with no layers, no logits, no sampling, and no token feedback. Its | |
| "tok/s" measured buffer-processing rate, not decoding. | |
| - Its Hugging Face baseline ran `use_cache=False`, which undercounts real HF | |
| decode speed; its DDR4 column in the showcase table was computed from | |
| `34500 / model_MB`, not measured. | |
| - Kernel improvements have raised this snapshot's decode rate several-fold from | |
| the first working build, and it remains compute-bound (see the probe results | |
| above), so neither the old synthetic numbers nor the current 1,974 tok/s | |
| represent a final ceiling. | |
| ## Kernel Optimization History | |
| Baseline (per-row indirect dot calls, no FMA, serial INT4 block reduces), | |
| Supra-Mini-8M, one thread, 128-token decode: FP32 757, INT8 1,623, | |
| hybrid INT4 1,152 tokens/s. | |
| After the first kernel pass (explicit FMA with 4 independent accumulator | |
| chains, AVX2 matvec kernels that absorb the row loop so one native call serves | |
| a whole matrix, FMA required in the CPU dispatch check): | |
| | Precision | Kernel-only rate | Full-model decode | Before | Gain | | |
| | --- | --- | --- | --- | --- | | |
| | FP32 | 30.8 GB/s | 936 tok/s (31.5 MB weights) | 757 | +24% | | |
| | INT8 | 24.2 GB/s | 1,974 tok/s (8.0 MB weights) | 1,623 | +22% | | |
| | Hybrid INT4 | 4.8 GB/s per byte | 1,382 tok/s (6.6 MB weights) | 1,152 | +20% | | |
| | SmolLM2-135M INT8 | - | 127 tok/s (135.4 MB weights) | 120 | +6% | | |
| Cache-residency probe (unchanged conclusion): decode effective bandwidth | |
| (15.8 GB/s at INT8) remains below even the cold DRAM scan rate (25.1 GB/s), | |
| so decode is still compute-bound. The INT4 kernel is the current kernel | |
| bottleneck (nibble unpacking without VNNI); full-model non-dot work (attention, | |
| norms, 16k-logit sampling, epilogue) accounts for roughly a third of decode | |
| time at INT8. | |
| Next levers, by expected impact: multithreaded matvec (row-parallel), | |
| prompt-prefill matrix-matrix batching, and further INT4 unpack improvements. | |
| ### Activation-quantized kernel follow-up (2026-09-12) | |
| The experimental `act_precision="int8"` path is now implemented for quantized | |
| matrix products. It quantizes every activation row in 32-element blocks using | |
| `32767 / absmax`, uses int16 `pmaddwd` MACs, and retains FP32 block scales for | |
| the epilogue. It is numerically sound but is not enabled by default: on this | |
| Ryzen 5 5600 (Zen 3, AVX2/FMA, no VNNI), the quantization work dominates small | |
| layer GEMMs. Supra-Mini-8M, one thread, measured about 1,753 tok/s with INT8 | |
| weights and FP32 activations versus 603 tok/s with the experimental path; | |
| hybrid-INT4 measured 1,297 versus 607 tok/s, and hybrid-FP4 1,238 versus 574. | |
| Custom PPL comparisons put the activation-mode delta near +0.01 on Supra-Mini | |
| and at or below +0.1 on SmolLM2-135M. Keep it as an explicit experimental API | |
| until a fused/coarser quantizer or a CPU with suitable integer-dot instructions | |
| shows a net benefit. | |
| ### Unpack-tax reduction, phase 1 (2026-09-12) | |
| Shipped: skip the nibble interleave (even/odd streams dot a permuted | |
| activation scratch built once per call from plain local buffers, no TLS), | |
| halving shuffle-port work per 32 weights; dead `matvec_*` family deleted | |
| (~350 lines, verified uncalled); 5-arg scalar adapters keep non-AVX2 builds | |
| working. Tried and reverted the same session: folding per-8 scale multiplies | |
| into a per-block scalar epilogue — the serial per-block reduce loses to four | |
| parallel mulps on Zen 3 out-of-order execution (measured slower, not faster). | |
| Measured on Ryzen 5 5600: full suite green, Mini PPL identical to 3 decimals, | |
| production AutoBench once showed fp4-2T beating int4-2T 1254 to 1015 (+24%) | |
| then losing the rematch — direction flips are inside this box's ±15-20% | |
| multithread noise, so neither is claimed. Drift-controlled A/B (alternating | |
| int8/fp4 in one process): 0.87 vs 0.90 pre-change. Verdict: the remaining | |
| ~26 µops/32 (converts, scale muls, act loads) are structural on AVX2 without | |
| VNNI; 4-bit cannot beat int8 kernel efficiency on this chip. Realistic next | |
| levers: int8-container upcast at load (int4 quality, near-int8 speed, int8 | |
| size — wins wherever cache residency makes bytes free), the q8 integer-MAC | |
| track, stronger ISA (VNNI/AMX/tensor cores), or DRAM-bound regimes where the | |
| 2x byte saving dominates. | |
| Fused epilogues tried and reverted (2026-09-12): accumulating o/down dots | |
| directly onto x_ and SiLU into the up-gemm epilogue measured ~7% SLOWER on | |
| Supra2-Medium int8-1T in a back-to-back stash A/B (580 vs 623, tight | |
| clusters, wrong order for a thermal explanation). Likely cause: branches plus | |
| a temp in the FMA loop beat clean separate vector passes on wide rows, and an | |
| unconditional permute scratch taxed int8/fp32 paths (since made lazy). | |
| Lesson: on Zen 3 streaming hardware, an extra cache-resident vector pass can | |
| be cheaper than a branch in the hot loop — fuse only with A/B proof. | |
| Session scratch arena (2026-09-12): per-call permute/row mallocs replaced by | |
| Session-persistent grow-only buffers (caller-owned, pool-safe; q8 TLS left | |
| alone). PPL identical; 4-bit/int8 tok/s ratios up ~0.1 at 2T on both models | |
| (Medium int4 0.83 -> 0.98, fp4 0.83 -> 0.96) — consistent with removing | |
| malloc-lock contention across pool threads rather than raw bytes. Model now | |
| reports info["compiled"] = {key, hit: False, mode: "generic", canonical} as | |
| the Phase-0 key plumbing goes live (execution still generic). | |
| Policy: AutoBench now defaults to 512 decode tokens and compiles nothing | |
| unless asked (inductor warmup taxed every default run); thread guidance is 2T | |
| on a loaded server, never 12T (oversubscription convoy collapses pooled | |
| paths while serial paths sail through). | |
| ### Supra2-Medium-Base production run (2026-09-12, live server, 512 tokens) | |
| First model that spills past L3 (fp32 96.78 MB, DRAM-streaming): | |
| | Precision | tok/s (2T) | PPL | vs fp32 | | |
| | --- | --- | --- | --- | | |
| | int8 (7.65 MB L3) | 540 | 59.40 | +0.62% | | |
| | hybrid-fp4 (20.32 MB L3) | 533 | 60.28 | +2.12% | | |
| | hybrid-int4 (20.89 MB L3) | 475 | 62.61 | +6.07% | | |
| | fp32 (96.78 MB DRAM) | 278 | 59.03 | — | | |
| | HF eager fp32 | 93 | 59.03 | — | | |
| | HF dynamic int8 | 72 | 95.45 | +61.7% | | |
| hybrid-fp4 beats hybrid-int4 on speed (+12%) and quality (+2.1% vs +6.1%) | |
| where matrices are large enough to amortize the unpack. int8 still leads | |
| overall at near-zero quality cost. Speculation accepts at 96-99% and is the | |
| largest single lever here (fp32 +65% at spec 8, int4 +17%); HF dynamic int8 | |
| is outclassed on both axes. Threads peak at 2-4T; 12T convulses as usual. | |
| ### Spec K curve (2026-09-12, Medium-Base int8, 2T, 64 tokens) | |
| Cap raised 8 -> 16 to test whether near-100% acceptance favors longer | |
| drafts. Measured: K=0: 844, K=2: 974 (84% acc), K=4: 1016 (76%), | |
| K=8: 954 (71%), K=12: 915 (69%), K=16: 834 (71%). Peak at K=4: | |
| acceptance decays with draft length while block-verify cost grows linearly, | |
| so dilution loses past the peak. Optimal K is prompt-dependent (longer, | |
| more repetitive generations accept more); the sweep now covers 0-16 with | |
| per-thread accept% tables so each setup can read its own peak. Thread sweep | |
| preset trimmed to 1,2,4 (above 4T never pays on these sizes). | |
| ### SmolLM2-135M precision x thread (2026-09-13, single-harness, act fp32, rev 93efa2f0) | |
| Weight bytes (engine.info): int8 135.4 MB (4.2x L3), hybrid-int4 105.1 MB, | |
| hybrid-fp4 100.2 MB; all spill, fits_l3=false. 512-token decode, greedy, | |
| spec_k=0, one live engine at a time: | |
| | Threads | int8 tok/s (GB/s) | hybrid-int4 tok/s (GB/s) | hybrid-fp4 tok/s (GB/s) | | |
| | --- | --- | --- | --- | | |
| | 1T | 88 (11.9) | 72 (7.5, 0.82x) | 73 (7.3, 0.83x) | | |
| | 2T | 120 (16.2) | 107 (11.3, 0.89x) | 106 (10.6, 0.88x) | | |
| | 4T | 124 (16.8) | 122 (12.8, 0.98x) | 122 (12.2, 0.99x) | | |
| PPL (514 scored tokens, window 128): int8 34.74, hybrid-int4 42.07 (+21.1% | |
| vs int8), hybrid-fp4 41.32 (+18.9% vs int8; fp4 beats int4 on quality at | |
| fewer bytes, speed tied within noise). Central 64-token 1T check: int8 120.9, | |
| int4 90.4, fp4 89.0 tok/s with identical weight bytes — same ordering. | |
| 128-token bracket int8-1T 114.5 tok/s / 15.5 GB/s matches the old 120-127 | |
| table; 512-window context growth costs ~21%, so ratios are intra-run only. | |
| Verdict: int8 is bandwidth-bound (saturates at 2T, +3.6% to 4T), 4-bit paths | |
| are kernel-bound at every thread count (1T->4T ~1.7x scaling, DRAM idle, | |
| never cash the 1.29x byte edge). Kernel bet unchanged on Zen 3 / no-VNNI: | |
| int8 default everywhere, fp4 the 4-bit pick. 360M not fetched (724 MB BF16 | |
| shard); expect the amplified same pattern. Next levers: spec on DRAM-bound | |
| int8-135M, int8-container upcast at load for cache-resident sizes (test-gated), | |
| VNNI-class ISA. Compile Phase-0 stays hit:False/mode:generic (plumbing only, | |
| execution generic) — correct until a PPL-gated execution change ships. | |
| ## Threading Results | |
| Persistent native pool, row-sliced matvecs, pause-spin barriers (yield/sleep | |
| only after long idle). Supra-Mini-8M INT8, 256-token decode, medians: | |
| | Threads | Tokens/s | vs 1T | | |
| | --- | ---: | ---: | | |
| | 1 | 1,548 | 1.00x | | |
| | 2 | 2,148 | 1.39x | | |
| | 4 | 2,036 | 1.31x (regresses) | | |
| FP32 (31.5 MB, DRAM-bound) shows no thread scaling: 998 -> 1,004 tok/s at 4T. | |
| The Ryzen 5600 sweet spot is 2 threads for cache-sized quantized models; | |
| beyond that, shared L3/LSU bandwidth saturates. gVisor-class sandboxes showed | |
| ~19 us sync round trips, so threading there may need larger models to pay off. | |
| `cism bench/serve --threads N` sets the pool size; default remains 1. | |
| ## Speculative Decoding (Prompt Lookup, Greedy) | |
| Implemented: n-gram drafter over full token history (up to 4-gram tails), | |
| block verification via per-matrix GEMM (all T candidates share one weight | |
| stream), greedy argmax acceptance with KV rollback, budget/EOS boundaries. | |
| Speculative greedy output is bitwise identical to plain greedy decode (tested | |
| across architectures, precisions, rejection paths, and budgets). | |
| SmolLM2-135M (135 MB INT8 / 540 MB FP32 - exceeds the 32 MB L3), one thread, | |
| 256 tokens, repetitive prompt for high draft acceptance: | |
| | Configuration | Tokens/s | Gain | | |
| | --- | ---: | ---: | | |
| | FP32 plain | 51 | 1.00x | | |
| | FP32 spec_k=4 | 83 | 1.63x | | |
| | INT8 plain | 106 | 1.00x | | |
| | INT8 spec_k=4 | 130 | 1.23x | | |
| Reason: speculation amortizes weight-bandwidth (one stream per verify block), | |
| so gains appear when decoding is DRAM-bound - the FP32 regime. The INT8 kernel | |
| is compute-bound (conversion-heavy SIMD), so amortizing bandwidth buys little | |
| until integer kernels land. On the AVX-512 VNNI cloud target, integer kernels | |
| should make INT8 bandwidth-bound and transfer the full speculation multiplier. | |
| Two structural costs: the bonus token costs one extra single-token pass (ideal | |
| (K+1)/2, measured 1.6x at K=4), and low-acceptance text can make speculation | |
| slightly negative - disable it when drafts do not match the domain. | |
| ## Memory Control Verification | |
| Real-model page locking and touch on this machine (Supra-Mini-8M INT8): | |
| - `lock_pages()`: 8,022,880 of 8,022,784 weight bytes locked (page rounding). | |
| - `touch()`: 8,022,880 bytes swept. | |
| - Generation while keep-warm active and pages locked: normal. | |
| - `unlock_pages()`: 6,268,768 bytes reported. Smaller than the lock count | |
| because adjacent weight regions share 4 KiB pages; a page released by one | |
| region cannot be released again by its neighbor. Counts are approximate byte | |
| sums of successful OS calls, not exact page accounting. | |
| x86 cannot pin cache lines; locking prevents OS paging, and sweeping only | |
| refreshes residency while running. Suite at time of writing: 571 tests passed. | |
| ## Cache-Residency Probe Results | |
| `cism cache-report` measures a warm full-weight scan, a cold scan after | |
| evicting caches with a 96 MiB scrub buffer, and decode effective bandwidth | |
| (weight bytes x decode tok/s). AMD Ryzen 5 5600, INT8, one thread: | |
| | Model | Weights | Warm scan | Cold scan | Decode | Effective | Verdict | | |
| | --- | --- | --- | --- | --- | --- | --- | | |
| | Supra-Mini-8M | 8.0 MB | 28.5 GB/s | 25.1 GB/s | 1,646 tok/s | 13.2 GB/s | Compute-bound | | |
| | SmolLM2-135M | 135.4 MB | 21.4 GB/s | 23.1 GB/s | 120 tok/s | 16.3 GB/s | Compute-bound | | |
| Interpretation: | |
| - Decode bandwidth is well below even the cold (post-eviction) scan rate in | |
| both cases. The decoder does not saturate DRAM, so **cache residency is not | |
| the current bottleneck; kernel efficiency is** (scalar conversion-heavy | |
| quantized kernels, sequential prefill, single thread). | |
| - The warm/cold scan ratio is small for these linear probes; hardware | |
| prefetchers hide much of DRAM latency on sequential streams, so this probe | |
| bounds bandwidth rather than proving residency. Definitive per-event counts | |
| (LLC miss) require hardware PMU tooling such as WindowsPerf or AMD uProf. | |
| - The earlier session's implied "SRAM 3x-6.5x multiplier" was not observed by | |
| this probe on this hardware and remains unproven. | |
| Speed work priorities implied by this data: quantized kernel quality, prefill | |
| matrix-matrix, and multithreading before further cache-tuning. | |
| ## Measured Comparison Against PyTorch 2.14 And torch.compile | |
| Supra-Mini-v5-8M, revision bb98e356, FP32 checkpoint, greedy, one torch/CISM | |
| thread, 6-token prompt, 256-token fixed-length decode (EOS ignored both sides), | |
| warmup absorbed, Dynamo counters verified zero recompilations during | |
| measurement. Torch baselines use a KV cache (never use_cache=False). | |
| Script: `scripts/bench_vs_torch.py`; CISM: `cism bench --max-tokens 256`. | |
| | Decoder | Tokens/s | CISM FP32 (998.5) | CISM INT8 (1,716.4) | | |
| | --- | ---: | ---: | ---: | | |
| | HF generate(use_cache=True) eager | 158.9 | 6.28x | 10.80x | | |
| | Manual eager loop + StaticCache | 195.8 | 5.10x | 8.77x | | |
| | torch.compile default, dynamic=False | 260.1 | 3.84x | 6.60x | | |
| | torch.compile default, dynamic=True | 224.9 | 4.44x | 7.63x | | |
| The fair comparison is against torch.compile dynamic=False, its best | |
| configuration here: **3.84x at matched FP32 precision, 6.60x for CISM INT8.** | |
| Reasons, with evidence: | |
| 1. **Per-token dispatch.** torch.compile lowers eager's per-op Python/ATen | |
| dispatch (159 -> 260 tok/s, +64%) but still launches dozens of generated | |
| kernels through a Python-level decode loop each token. CISM executes the | |
| whole token in one C++ call. Eager's effective bandwidth is | |
| 31.5 MB x 159 = 5.0 GB/s and compile's is 8.2 GB/s - both overhead-bound, | |
| while CISM FP32 runs memory-bound at ~30 GB/s kernel rate. | |
| 2. **Memory traffic.** CISM INT8 reads 8.0 MB/token versus 31.5 MB for the | |
| FP32 checkpoint, cutting the memory-bound floor ~4x on top of the | |
| dispatch elimination. | |
| 3. **Quality.** CISM FP32 matches eager within 1e-5 max logit error | |
| (established above). INT8 quality reporting remains a beta gate. | |
| Caveats: compile mode is "default" (max-autotune not used); one compile | |
| configuration initially crashed with an Inductor codegen bug on this model | |
| (index-out-of-bounds in a generated kernel) and only ran after cache-reset | |
| handling; multithreaded torch was not measured; prefill/TTFT is excluded from | |
| decode rates. This is one cache-sized model on one machine - not yet the | |
| agreed multi-model llama.cpp gate. | |
| ## 4-bit quantization scheme study (2026-09-12) | |
| NumPy simulation mirroring the native hybrid-int4 quantizer (round-half-away, | |
| blocks of 32 along the row, per-block scale; validated against the native | |
| engine to 8 decimal places on all checkpoints). PPL over the fixed AutoBench | |
| corpus (514 tokens, window 128), MLP matrices 4-bit, all other 2D matrices | |
| per-row INT8, vectors FP32. | |
| Machine: Ryzen 5 5600, local checkpoints, scripts/quant4_study.py. | |
| | scheme (MLP element + scale) | bits/w | 8M ΔPPL | 100M ΔPPL | 135M ΔPPL | | |
| |-------------------------------------|-------:|--------:|----------:|----------:| | |
| | int4 + FP32 b32 (hybrid-int4) | 5.00 | +15.4% | +6.1% | +23.2% | | |
| | int4 + E4M3 b32 | 4.25 | +15.5% | +6.0% | +26.7% | | |
| | int4 + E8M0 b32 | 4.25 | +308% | +80.8% | +537% | | |
| | E2M1 + E8M0 b32 (OCP MXFP4) | 4.25 | +318% | +92.4% | +323% | | |
| | E2M1 + E4M3 b32 | 4.25 | +18.7% | +3.5% | +23.3% | | |
| | E2M1 + E4M3 b16 (NVFP4 granularity) | 4.50 | +15.0% | +2.1% | +21.0% | | |
| | int4 + E4M3 b16 | 4.50 | n/a | +8.1% | n/a | | |
| | E2M1 + FP32 b32 | 5.00 | +13.7% | +4.4% | +19.9% | | |
| | NF4 + FP32 b32 (QLoRA) | 5.00 | +44.1% | +8.8% | +32.0% | | |
| | int4 + error feedback b32 | 5.00 | +36.5% | +16.9% | +80.9% | | |
| Findings: | |
| - Scale precision dominates element format; power-of-two (E8M0) scales are | |
| catastrophic for both uniform and FP elements. | |
| - E2M1 elements beat uniform INT4 at equal scale precision and block size | |
| (100M model, b16+e4m3: +2.1% vs +8.1%). | |
| - E2M1 + E4M3 b16 is the only scheme that improves on hybrid-int4 at every | |
| model size while using fewer bits (4.50 vs 5.00): tied on 8M, 3x better on | |
| 100M, better on 135M. | |
| - NF4 (QLoRA) loses badly on these trained small models; naive per-block | |
| error feedback with a fixed scale inflates residuals and hurts everywhere. | |
| Why hybrid-int4 decodes slower than int8 (measured, 1 thread, same | |
| checkpoints): the nibble unpack is a ~10-instruction dependent chain per | |
| block, so hybrid-int4 delivers only 7.6-8.1 GB/s of weight-equivalents vs | |
| 12.5-14.0 for int8 and 26.6 for fp32, despite moving 2.5x fewer bytes than | |
| fp32. The E2M1 LUT kernel replaces that chain with a 16-entry table lookup | |
| and is the reason to build the mode. | |
| ### Compile one-shot canonical parity (2026-09-14) | |
| Phase-0 key plumbing is now cross-language exact. Python | |
| `compile_cache.canonical_model_config()` builds the native 12-field pipe | |
| `model_type|hidden|intermediate|layers|heads|kv_heads|head_dim|vocab| | |
| context|hexfloat(eps)|hexfloat(rope_theta)|tied` with FP32-round-trip | |
| `struct.pack('f')` hexfloat, matching `Model::refresh_compile_key` | |
| bit-for-bit. Verified: tiny Llama 16/32/1/4/2/4/32/32 probe Python | |
| canonical equals `Model.info["compiled"]["canonical"]`, and | |
| `native.compile_key(python_canonical, precision, act, threads)` equals | |
| `Model.info["compiled"]["key"]` (`d4547092c73a7aac` on the probe). | |
| Legacy `canonical_config()` sorted-JSON remains for early tests only. | |
| Kept as a stable model fingerprint (zero runtime cost, tested); the | |
| packed-artifact load idea is dropped with the project (see below). | |
| ### SmolLM2-135M full AutoBench matrix (2026-09-13, 512 tokens, act fp32) | |
| Local file `autobench_135M_512.json` (untracked measurement, torch | |
| baselines disabled by policy `include_compile=False`): 4 precisions x | |
| threads (1,2,4) x spec_k (0,2,4,8,12,16) = 72 decode cells, PPL once per | |
| precision (514 scored tokens, window 128). Revision `93efa2f0` (report | |
| does not yet stamp `revision`; next run should copy | |
| `engine.info["revision"]`). All `fits_l3=false` (L3 cache size 32.0 MB). | |
| Weight storage in megabytes: int8 129.16 MB (135,439,104 bytes), | |
| hybrid-int4 100.27 MB (105,141,504 bytes), hybrid-fp4 95.52 MB | |
| (100,164,864 bytes), fp32 513.13 MB. | |
| Decode rate in tokens/s at spec_k=0 (effective bandwidth in GB/s = | |
| weight bytes x tok/s): | |
| | Precision | 1 thread tok/s (GB/s) | 2 threads tok/s (GB/s) | 4 threads tok/s (GB/s) | Best tok/s (threads, spec_k) | PPL (+vs int8) | | |
| | --- | --- | --- | --- | --- | --- | | |
| | fp32 | 43.78 (23.6) | 57.44 (30.9) | 55.34 (29.8) | 106.41 (4T, K=8) | 34.15 | | |
| | int8 | 91.53 (12.4) | 124.82 (16.9) | 130.03 (17.6) | 147.19 (4T, K=2) | 34.74 (--) | | |
| | hybrid-int4 | 71.93 (7.6) | 106.54 (11.2) | 115.69 (12.2) | 140.59 (4T, K=16) | 42.07 (+21.1%) | | |
| | hybrid-fp4 | 71.85 (7.2) | 102.10 (10.2) | 117.51 (11.8) | 132.90 (4T, K=2) | 41.32 (+18.9%) | | |
| Ratios int8 vs 4-bit at spec 0: 1T 1.27x, 2T 1.17-1.22x, 4T | |
| 1.11-1.12x. Scaling 1T to 4T: int8 1.42x but 2T to 4T only +4.2% | |
| (saturates, bandwidth-bound); int4 1.61x, fp4 1.64x, still climbing | |
| with DRAM idle (kernel-bound nibble unpack, no VNNI). Verdict unchanged: | |
| int8 default everywhere, fp4 the 4-bit pick (better quality at fewer | |
| bytes, speed tied within ±15-20% multithread noise; ratios | |
| single-harness only). | |
| ### Spec K-curve on 135M DRAM-bound int8 (512 tokens, greedy) | |
| Accept rate in percent is deterministic across threads per K | |
| (88.0% at K=2 decaying to 72.4% at K=16). | |
| | Threads | K=0 tok/s | K=2 tok/s | K=4 tok/s | K=8 tok/s | K=12 tok/s | K=16 tok/s | Peak gain vs K=0 | | |
| | --- | --- | --- | --- | --- | --- | --- | --- | | |
| | 1T | 91.53 | 98.42 | 100.47 | 99.45 | 96.14 | 93.67 | K=4, +9.8% | | |
| | 2T | 124.82 | 136.05 | 132.62 | 131.34 | 129.72 | 126.94 | K=2, +9.0% | | |
| | 4T | 130.03 | 147.19 | 146.55 | 135.13 | 136.81 | 143.53 | K=2, +13.2% | | |
| Global best: int8 147.19 tok/s at 4 threads, spec_k=2. Contrast fp32 | |
| +73-92% peaking K=8 (bandwidth-bound, block verify amortizes DRAM) vs | |
| 4-bit +5-21% erratic (kernel-bound). Peak at small K because acceptance | |
| decays with draft length while GEMM verify cost grows linearly, plus the | |
| bonus-token extra pass. Prompt-lookup drafter (up to 4-gram, greedy-only, | |
| bitwise identical) helps repetitive prompts only; it does not fix the | |
| 135M slow complaint (still ~147 tok/s vs ~1,600 tok/s cache-resident 8M). | |
| Default 135M int8: 4 threads, spec_k=2 (2-thread servers also K=2). | |
| ### trust_remote_code custom-arch assessment (2026-09-14) | |
| `trust_remote_code=True` authorizes tokenizer Python only, never custom | |
| modeling operators; trusted code itself is not sandboxed. Enforcement: | |
| loader `_ARCHITECTURES={"llama","qwen3"}`, `validate_config` rejects | |
| non-llama/qwen3 `model_type` and any `auto_map` beyond `AutoTokenizer` | |
| even with trust enabled, cross-repo refs rejected, `_expected_shapes` | |
| plus `_load_weights` reject extra/missing tensors, native | |
| `Config::validate`/`parse_config` reject bias/MoE/sliding-window/ | |
| non-silu/non-default-RoPE. Custom Negative/CMA/Ember2/Spark2A/Blaze | |
| targets (ROADMAP gate 4: recurrence, routing, lane mixing, Engram state) | |
| need new native Config fields, operators, KV layouts, quant definitions, | |
| and PPL/parity gates per arch — large work, not a flag flip. Keep | |
| explicit `UnsupportedModelError`; server stays opt-in | |
| `--allow-remote-code-load` (403 otherwise). | |
| ### Compile project dropped (2026-09-14, verdict: correctly built, not worth it) | |
| Target was one-time load cost for faster batch-1 CPU decode (generic | |
| AVX2, PPL-gated). Verdict: the plumbing is correct and tested, but the | |
| goal is unreachable on this CPU, so the project is dropped and its | |
| speedup fiction removed: | |
| - What was real: cross-language canonical parity (bit-exact, tested), | |
| hoisted dispatch loops (same kernels/order, bitwise identical, kept as | |
| plain code), manifest helpers (kept, unused by load path). | |
| - What was fiction: manifest-hit as "compile hit" (bookkeeping, not a | |
| faster layout), `compiled.hit/mode` in Engine info and tester cards | |
| (removed), any decode tok/s attributed to compile (zero: mode was | |
| generic everywhere; the +13.2% is spec K=2). | |
| - Why not worth it: remaining ~26 µops/32 are structural on AVX2/no-VNNI; | |
| int8 is bandwidth-bound (2T saturates), 4-bit kernel-bound; fresh 512 | |
| runs show hoisting moved 4-bit ~+23% but int8 K0 flat (-0.9%), | |
| confounded by idle-machine effect — no clean decode win to claim; | |
| inductor costs 107 s warmup for 32 tok/s against CISM's 190 (6-8x | |
| lead from the pre-existing single-call C++ path + quant, not compile). | |
| VNNI/VBMI or integer-MAC kernels would reopen the question; until | |
| then the line is closed. Fingerprint key + hoisted loops stay; hit/UI/ | |
| future-speedup prose goes. | |
| ### Fresh 512-token 4T confirmation (SmolLM2-135M, all precisions, spec sweep) | |
| Re-ran 24 cells (4 precisions x 4T x spec 0,2,4,8,12,16), 512 tokens, | |
| 127.3 s wall, torch off. Accept% deterministic and identical to the old | |
| file; PPL/weights identical (fp32 34.15, int8 34.74, int4 42.07, fp4 | |
| 41.32; 513.13/129.16/100.27/95.52 MB, all fits_l3=false). Absolute tok/s | |
| ran +11-30% hotter than the old file (idle machine); cells >25% off are | |
| marked SUSPECT — use this table, do not mix files. | |
| | Precision | K=0 | K=2 | K=4 | K=8 | K=12 | K=16 | Best | | |
| | --- | --- | --- | --- | --- | --- | --- | --- | | |
| | fp32 | 58.77 | 72.89 | 101.92 | 125.09 | 122.38 | 121.18 | K=8 125.09 | | |
| | int8 | 128.90 | 159.99 | 162.64 | 171.76 | 171.37 | 177.03 (SUSPECT vs old) | K=16 177.03 | | |
| | hybrid-int4 | 142.63 | 155.69 | 151.97 | 159.86 | 160.29 | 156.40 | K=12 160.29 | | |
| | hybrid-fp4 | 144.58 | 150.73 | 153.96 | 160.41 | 155.88 | 153.82 | K=8 160.41 | | |
| Accept%: fp32 96.5/94.0/89.9/88.1/85.5; int8 88.0/83.4/77.1/73.0/72.4; | |
| int4 90.1/84.4/79.1/75.6/75.3; fp4 84.1/76.6/70.3/66.7/67.6 (K=2..16). | |
| Note 4-bit now leads int8 at K=0 (142.63/144.58 vs 128.90) in this | |
| run — confounded by machine state (old file had int8 leading), so no | |
| compile attribution claimed without a same-state A/B. fp4 vs int4: tied | |
| within noise at every K (max gap 3.7%), fp4 keeps the quality edge | |
| (+18.9% vs +21.1% PPL), so fp4 stays the 4-bit pick. | |
| ### 4T K=2 vs regular baselines (SmolLM2-135M int8, 512 tokens) | |
| Global best cell is int8 4 threads spec_k=2 at 147.19 tok/s (accept | |
| 87.96%). Regular baselines from the same `autobench_135M_512.json` run | |
| (`include_torch=False`, so no HF eager/compile in this file): | |
| | Baseline (4T) | Tok/s | vs 4T K=2 (147.19) | | |
| | --- | --- | --- | | |
| | int8 K=0 (same weights, no spec) | 130.03 | +13.2% with spec | | |
| | int8 2T K=0 | 124.82 | +17.9% (threads+spec) | | |
| | int8 1T K=0 | 91.53 | +60.8% | | |
| | fp32 K=0 | 55.34 | 2.66x | | |
| | fp32 best (K=8) | 106.41 | 1.38x | | |
| | hybrid-int4 K=0 | 115.69 | 1.27x | | |
| | hybrid-int4 best (K=16) | 140.59 | 1.05x | | |
| | hybrid-fp4 K=0 | 117.51 | 1.25x | | |
| | hybrid-fp4 best (K=2) | 132.90 | 1.11x | | |
| PPL in same run: int8 34.74, fp32 34.15, int4 42.07, fp4 41.32. | |
| Spec pays (+13.2%) because 135M int8 spills (129.16 MB vs 32 MB L3): | |
| block verify streams weights once per block via GEMM. It does not make | |
| 135M fast in absolute terms (still ~147 vs ~1,600 cache-resident 8M). | |
| ### HF eager/compile baselines (SmolLM2-135M, 512 tokens, 4T, same harness) | |
| One `run_benchmark` (`src/cism/autobench.py:223`) with | |
| `include_torch=True, include_compile=True`, `max_new_tokens=512`, | |
| threads 4, spec sweep 0-16, rev `2c4d40c`. Whole sweep ran ~4-15% under | |
| the fresh reference (mild thermal throttle); HF rows are baselines, not | |
| throttle-sensitive conclusions. | |
| | Decoder (512 tok, 4T) | Tok/s | PPL (514 scored tokens) | | |
| | --- | --- | --- | | |
| | CISM int8 K=0 | 131.88 | 34.74 (nll 3.5480, +1.73% vs eager) | | |
| | CISM int8 best K=16 | 167.23 | same | | |
| | CISM fp32 K=0 / best K=8 | 53.32 / 114.67 | 34.15 (+0.002%) | | |
| | CISM hybrid-int4 K=0 / best K=8 | 137.20 / 153.07 | 42.07 (+23.18%) | | |
| | CISM hybrid-fp4 K=0 / best K=16 | 137.38 / 150.25 | 41.32 (+21.00%) | | |
| | HF eager fp32 | 23.79 (21.52 s) | 34.15 (nll 3.5308) | | |
| | HF dynamic int8 | 19.74 (25.94 s) | 81.12 (broken, +137.5%) | | |
| | HF torch.compile fp32 | 25.18 (20.33 s, warmup 147.3 s) | not scored | | |
| | HF compile int8 | 16.86 (30.38 s) | not scored | | |
| | torchao int4 | skipped (not installed) | -- | | |
| Speedups (same harness): CISM int8 K=0 **5.54x eager / 5.24x compile**; | |
| best K=16 **7.03x / 6.64x**; fp32 K=0 2.24x/2.12x, best K=8 4.82x/4.55x. | |
| Compile barely beats eager (+5.9%) after a 147 s warmup; dynamic int8 is | |
| slower than fp32 AND destroyed (PPL 81). Weights: 513.13/129.16/100.27/ | |
| 95.52 MB, all fits_l3=false (L3 32 MB). | |
| ### Surjo hybrid native Phase1 (loader+fp32 decode) | |
| Scope: SurjoLabs Surjo hybrid import plus native fp32 decode runs. No | |
| Surjo tok/s, no Surjo PPL, no speedup claim in this section. Reference | |
| dense numbers below are context only, single-harness, ±15-20% | |
| multithread noise (see 135M matrix notes); do not mix files or | |
| attribute cross-run deltas to Surjo. | |
| Config fingerprint (Surjo-50m, `tests/test_surjo.py:20`): `model_type` | |
| surjo / `SurjoForCausalLM`, vocab 32768, hidden 512, intermediate | |
| 1536, layers 10, context 2048, prelude 1 / recurrent 8 / coda 1, | |
| groups 2 / passes 2 / gdn_per_xsa 3, heads 8 / kv 4 / head_dim 64, | |
| gdn_v_heads 8 / gdn_k_dim 64 / gdn_v_dim 64 / kernel 4, | |
| `gdn_allow_neg_eigval` false, `xsa_projection` true, bias false, | |
| dropout 0.0, `hidden_act` silu, eps 1e-5, theta 10000.0, default RoPE | |
| (`partial_rotary_factor` 1.0), tied true, `layer_types` | |
| `[F,L,L,L,F,L,L,L,F,F]`, `auto_map` pinned to | |
| `configuration_surjo`/`modeling_surjo` (never executed). Base | |
| checkpoints ship `use_cache=false`; `import_model` | |
| (`src/cism/loader.py:666`) preserves the flag verbatim and it does | |
| not disable native caching (XSA KV slots + GDN recurrent/conv state | |
| always allocate). | |
| Checkpoint header: 178xF32 tensors in `model.safetensors`, tied | |
| `lm_head` omitted, no rotary `inv_freq` buffers (RoPE is computed in | |
| `forward()`). `_expected_shapes_surjo` length is 179 including the | |
| tied-head alias entry; spot-checked in | |
| `test_surjo_50m_expected_shapes_match_header` (embed `(32768,512)`, | |
| XSA `q (512,512)` / `k (256,512)` / `q_norm (64,)`, GDN `q/k (512,512)` | |
| / `v (512,512)` / conv `(512,1,4)` / `f0 (64,512)` / `f1 (512,64)` / | |
| `A_log (8,)` / `dt_bias (512,)` / `g1.bias (512,)` / `o_norm (64,)`). | |
| Loader (`src/cism/loader.py`): `validate_surjo_config`, | |
| `surjo_exec_plan` (18 steps: 1 prelude + 2 passes x 2 groups x | |
| (3 GDN + 1 XSA) + 1 coda; XSA slots `[0..5]`, 12 GDN `(layer,pass)` | |
| steps), `surjo_num_xsa_slots` (=6), `_expected_shapes_surjo:393`, | |
| `_load_weights` Surjo branch (no buffers, strict shape/dtype/finite | |
| checks, tied-head alias by identity), `import_model` dispatch | |
| (`model_type=="surjo"` -> Surjo path, else dense). `_SURJO_SUPPORT` | |
| string still reads "next phase" in this snapshot; Phase1 below is | |
| what landed natively. | |
| Native (`native/runtime.{hpp,cpp}`, `native/bindings.cpp:468`, | |
| `src/cism/engine.py:80` dispatch to `_native.SurjoModel`): | |
| `SurjoModel`/`SurjoSession` fp32 decode. XSA: 6 slots, KV | |
| `[slots,capacity,kv_width]` with `kv_width=256`; at max context 2048 | |
| that is 4 MiB/slot (K+V), 24 MiB total (session-sized by | |
| `prompt+max_new`, so 24 MiB is the upper bound). GDN: 12 states | |
| (6 GDN layers x 2 passes), each `S` is `H*K*V=8*64*64=32768` floats = | |
| 128 KiB (12x128 KiB = 1.5 MiB), plus depthwise-conv raw FIFO | |
| `(kernel-1)=3`: q/k/v FIFOs ~72 KiB each, ~0.21 MiB total. MLP is | |
| ClampedMLP (gate clamped to `[-15,15]` then `silu*up`, | |
| `runtime.cpp:2010`). `spec_k` is rejected for Surjo | |
| (`next_tokens` throws unless `spec_k==0`; `Engine.stream(spec_k=2)` | |
| raises `ValueError`, covered by `test_surjo_engine_decode_runs`). | |
| Prefill is sequential per-token (`forward_tokens` loop); `logits` / | |
| `nll` teacher-forcing paths exist but are unscored against HF so far. | |
| Quant mapping reuse, no new format: `protected_storage` (embed, head, | |
| all XSA/GDN projections incl. `f0/f1/g0/g1/o_proj`) is fp32 or int8; | |
| MLP `gate/up/down` follow the dense hybrid rule (hybrid-int4 -> int4, | |
| hybrid-fp4 -> fp4). Same `Matrix` kernels/scratch/act_q8 path as | |
| dense; only norms/vectors/conv/A_log/dt_bias stay fp32. | |
| Parity gate: dense bar is max abs logit error <1e-4 vs HF eager. | |
| Surjo fp32 decode/smoke runs (tiny shapes fixture, tied-head | |
| identity, `Engine.logits/generate` finite, spec rejection). HF | |
| `modeling_surjo.py` parity script and Surjo PPL are pending — no | |
| parity number claimed here. | |
| Batch/paging: Phase0 files unchanged. Single-request serialized | |
| `Engine` generation lock, `create_session`/`next_tokens`, | |
| cancel/finish, `lock_pages`/`touch`/`scan` weight-memory controls | |
| are reused as-is. No continuous batching, no paged-KV layout, no new | |
| server path for Surjo. | |
| FLA CPU verdict: decode needs no Flash-Linear-Attention (single-token | |
| GDN recurrence is O(1) state + per-slot causal XSA). Chunkwise/FLA | |
| prefill is an optional long-prompt throughput optimization only, not | |
| correctness; current sequential prefill stands. | |
| Dense context (same `autobench_135M_512.json`, `include_torch=False`, | |
| NOT Surjo results): 4T K=0 fp32 55.34 / int8 130.03 / hybrid-int4 | |
| 115.69 / hybrid-fp4 117.51 tok/s; best cells fp32 106.41 (K=8), int8 | |
| 147.19 (K=2, accept 87.96%), hybrid-int4 140.59 (K=16), hybrid-fp4 | |
| 132.90 (K=2); PPL fp32 34.15 / int8 34.74 / int4 42.07 / fp4 41.32. | |
| Ratios single-harness only; 4-bit vs int8 gaps sit inside ±15-20% | |
| noise. Surjo inherits none of these numbers. | |
| ### Surjo-50m full AutoBench (Phase2, 2026-09-15, 4T K0) | |
| File `autobench_surjo50m_512.json` (local snapshot | |
| `35c9fa6e`, 512-token HF window, CISM greedy K0 only, act fp32, | |
| threads 1/2/4). Weights (engine.info): fp32 205.21 MB, int8 | |
| 51.84 MB, hybrid-int4 43.27 MB, hybrid-fp4 41.86 MB; L3 32.0 MB, | |
| all `fits_l3=false`. PPL over the fixed AutoBench corpus (527 scored | |
| tokens): fp32 59.29 / int8 59.46 (+0.30% vs HF eager) / hybrid-int4 | |
| 62.90 (+6.10%) / hybrid-fp4 61.80 (+4.25%). CISM int8/fp32 PPL matches | |
| HF eager fp32 59.28 to 0.3% — quant quality gate holds at int8, 4-bit | |
| pays +4-6%. | |
| CISM decode at 4T K0 (greedy; CISM stopped on EOS at 43-48 tokens, | |
| HF forced 512 via min_new_tokens — lengths differ, ratios indicative): | |
| | Precision | 1T tok/s | 2T tok/s | 4T tok/s | PPL | | |
| | --- | --- | --- | --- | --- | | |
| | fp32 | 62.18 | 81.03 | 81.06 | 59.29 | | |
| | int8 | 129.86 | 158.39 | 207.80 | 59.46 | | |
| | hybrid-int4 | 109.05 | 159.04 | 175.49 | 62.90 | | |
| | hybrid-fp4 | 105.29 | 153.13 | 171.31 | 61.80 | | |
| HF baselines (same harness, FP32 CPU, same thread count): | |
| | Decoder (4T) | Tok/s | PPL | | |
| | --- | --- | --- | | |
| | HF eager fp32 | 26.27 (1T 28.08 / 2T 27.74 — inverse scaling, MT noise) | 59.28 | | |
| | HF dynamic int8 | 18.18 (main cell 19.93; sweep 17.42/18.95/18.18) | 66.35 destroyed (+11.9%) | | |
| | HF torch.compile fp32 | ~~31.72~~ INVALID — StaticCache path, see below | not scored | | |
| | HF compile int8 | 16.73 (same StaticCache invalidity) | not scored | | |
| Speedups at 4T K0 vs HF eager 26.27: int8 7.91x, hybrid-int4 6.68x, | |
| hybrid-fp4 6.52x, fp32 3.09x. `speedup_vs_compile` in the JSON divides | |
| by the invalid ~~31.72~~ and is void — do not cite. | |
| Compile correction: `modeling_surjo.py:645` discards non-SurjoCache | |
| past (`if not isinstance(past, SurjoCache): past = None`), so the | |
| autobench StaticCache loop allocated a fresh cache per forward call and | |
| state never carried — ~~31.72 tok/s~~ is not a Surjo compile number. | |
| `src/cism/autobench.py` now detects `surjo_mode` and threads returned | |
| `past_key_values` (SurjoCache) under `torch.compile(dynamic=True)`, | |
| recording `{"skipped": "StaticCache invalid for Surjo"}` if the | |
| threaded path cannot run — never a StaticCache tok/s. Correct path | |
| proven by `scripts/bench_compiler.py`: threaded SurjoCache + | |
| `dynamic=True` ~41 tok/s (compile warmup 43-119 s across runs; | |
| 189.3 s in the old JSON was the invalid-path warmup); | |
| `dynamic=False` unusable (~0.06 tok/s, recompiles per token as XSA KV | |
| length grows). torchao int4 blocked on Windows (no mslk wheel — | |
| `Requires mslk >= 1.0.0`); torchao int8wo is the CPU-working control | |
| only. FLA CPU dead here (Triton needs GPU; `fla.ops.gdn2` chunk/fused | |
| probe does not run on CPU) — decode needs no FLA anyway (single-token | |
| GDN recurrence is O(1) state + per-slot causal XSA). | |
| Spec: unsupported for Surjo hybrid (GDN state destructive) — K0 only | |
| (`spec_note` in JSON); K2≈K0 where probed, no speculation gain claimed. | |
| Native `SurjoSession` rejects `spec_k!=0`; `Engine.stream(spec_k=2)` | |
| raises. | |
| Rigor: this JSON predates the harness (`matrix` rows lack | |
| `rep_rates`/`thermal_throttled`/`cpu_snapshot`). Current | |
| `src/cism/autobench.py:_timed` reports 1 warmup + median-of-`reps` | |
| (`--reps`/`--runs` wired through `cism bench`; default 3) with min/max/ | |
| std, `thermal_throttled` when max-min spread exceeds 15% (wider than the | |
| ±15-20% MT noise band), and best-effort `cpu_snapshot` (clock/temp; | |
| `{}` when unavailable, never fails the run). Next Surjo rerun picks up | |
| medians + thermal flags automatically; do not mix old single-sample | |
| cells with new median cells. | |
| ### Surjo-50m rerun on current code (2026-09-15 night, 72 cells, reps=3) | |
| `autobench_surjo50m_512.json` overwritten with current native | |
| (GDN-OPT + XSA-blocked + `forward_block` + spec verify). Threads | |
| 1/2/4 x spec 0/2/4/8/12/16, 512tok, act fp32, median-of-3, all | |
| `thermal_throttled=false` except two fp32-4T spec cells. CISM EOS-stops | |
| 43-48tok vs HF forced-512 — same caveat as before. | |
| K0 medians: fp32 60.32/84.73/88.86, int8 141.76/197.34/234.87, | |
| int4 114.59/169.65/197.28, fp4 113.54/169.98/214.61. Best cells: | |
| fp32 K2 92.66, int8 K4 236.91, int4 K4 201.33, fp4 K0 214.61. | |
| Spec accept on bench prompt is 0.0 for fp32/int8 (no gain, no loss | |
| beyond noise); int4 6.25%->1.7%, fp4 8.3%->2.7% decaying with K, and | |
| spec K>0 is slower than K0 there (verify cost, no amortization win | |
| on this prompt). PPL unchanged: 59.29/59.46/62.90/61.80. | |
| HF equal-opportunity, same harness, threads swept, median-of-3: | |
| eager 31.22/30.43/29.50 (4T 29.50 PPL 59.28), | |
| dynamic-int8 20.27/22.59/22.37 (4T 22.69 PPL 66.35), | |
| compile threaded-SurjoCache dynamic=True 38.77 warmup 100.6s, | |
| compile-int8 21.10, torchao-int4 failed `Requires mslk>=1.0.0`. | |
| Speedups vs eager 29.50: int8 K0 7.96x, best K4 8.03x; | |
| vs compile 38.77: K0 6.06x. Old invalid StaticCache 31.72 is | |
| superseded — do not cite it. | |
| ### Supra-50M-Base full AutoBench (dense Llama, 2026-09-15 night) | |
| File `autobench_supra50M_512.json`, Hub rev | |
| `521bfd3d3901fbeefa943a2f3d461a9d193c52ec`, dense Llama | |
| (hidden 512, layers 12, heads 8/kv 4/head_dim 64, intermediate | |
| 1408, vocab 32000, ctx 1024, tied, `use_cache=false` preserved). | |
| Threads 1/2/4 x spec 0/2/4/8/12/16 = 72 cells, 512tok, act | |
| fp32, median-of-3. Weights: fp32 ~209MB / int8 ~53MB class | |
| (all `fits_l3=false`). PPL (527 scored tokens): fp32 63.32 / | |
| int8 63.79 (+0.7%) / int4 65.36 (+3.2%) / fp4 68.90 (+8.8%). | |
| CISM K0 medians: fp32 108.29/131.54/123.55, int8 | |
| 214.81/284.36/292.58, int4 178.04/247.49/250.87, fp4 | |
| 171.67/234.54/251.38. Best cells: fp32 K8 268.19 (2T), | |
| int8 K12 349.10 (4T), int4 K2 257.43, fp4 K8 274.25. | |
| Accept on bench prompt is 100% for fp32/int8/fp4 at every K | |
| (hence spec pays: int8 4T K0 292.58 -> K12 349.10, +19.3%; | |
| fp32 K0 123.55 -> K8 268.19, +117%); int4 86%->70%. | |
| HF same harness, threads swept, median-of-3: eager | |
| 48.59/47.90/43.61 (4T 43.61 PPL 63.32), dynamic-int8 | |
| 37.70/40.49/37.78 (4T 34.91 PPL 101.68 destroyed), | |
| compile StaticCache dynamic=False 55.23 warmup 143.0s, | |
| compile-int8 33.14, torchao-int4 `Requires mslk>=1.0.0`. | |
| Speedups vs eager 43.61: int8 K0 6.71x, best K12 8.00x; | |
| vs compile 55.23: K0 5.30x, best 6.32x. | |
| ### FWKV Myosotis-1-base full support + AutoBench (2026-09-15 night) | |
| Files inspected in snapshot `031fe84a34f667bc8bfff7dc3f8a77ef3c17efb2`: | |
| `config.json` (model_type fwkv, d_model 768, d_emb 192, | |
| n_layers 13, ffn_mult 4, vocab 50257, wkv_floor 0.1, tied), | |
| `configuration_fwkv.py` (FWKVConfig + HF aliases), | |
| `modeling_fwkv.py` (FactorizedTiedHead, FWKVBlock with | |
| proj_k/v/r/out + W recurrence + GELU FFN + 2x LayerNorm, | |
| `_supports_cache_class=False`, tuple-state threading), | |
| `model.safetensors` (147 F32 tensors, no rotary buffers). | |
| Native support landed, mirroring the Surjo vertical slice: | |
| `loader.validate_fwkv_config` + `_expected_shapes_fwkv` (147 | |
| shapes) + `_load_weights` fwkv branch (also fixed an eager | |
| `config["head_dim"]` default crash affecting any arch without | |
| head_dim) + `import_model` dispatch; `native` FwkvConfig / | |
| FwkvModel / FwkvSession (WKV recurrence, exact GELU, | |
| LayerNorm eps 1e-5, factorized head, fp32/int8/int4/fp4 reuse | |
| via Matrix, `forward_block` batched projections + sequential | |
| recurrence, spec verify with ~40KB state checkpoint/restore); | |
| `bindings` FwkvModel/FwkvSession; `engine` dispatch; | |
| `tests/test_fwkv.py` 11 passed; `scripts/parity_fwkv.py`: | |
| tiny 1.19e-07, live vs HF fp32 5.72e-06 (<1e-4 OK). | |
| Autobench: `trust_remote_code=True` for fwkv HF loads, | |
| threaded-FWKV-state compile path (StaticCache invalid, same | |
| class of bug as Surjo), `_hf_context` fallback to `seq_len` | |
| (FWKVConfig has no `max_position_embeddings`). | |
| File `autobench_fwkv_512.json`, 72 cells, reps=3 (v4 drafter rerun, | |
| 2026-09-15 late night; supersedes the v1-drafter run below). Weights: | |
| fp32 425.38MB / int8 107.64MB. PPL: fp32 156.68 (= HF eager | |
| 156.68, gate holds) / int8 156.80 (+0.07%) / int4 156.63 | |
| (-0.04%) / fp4 158.53 (+1.18%). High absolute PPL is base | |
| model quality, faithfully reproduced. | |
| CISM K0: fp32 64.65/93.73/87.24, int8 169.01/268.22/342.27, | |
| int4 121.54/215.32/284.66, fp4 119.21/205.16/303.52. Best: | |
| fp32 K16 177.34, int8 K4 382.74, int4 K2 290.19, fp4 K2 | |
| 305.98. Accept (bench prompt): fp32 85%->61%, int8 | |
| 85%->62%, int4 77%->45%, fp4 74%->44% (K2->K16). Spec pays | |
| on fp32/int8: int8 4T K0 342.27 -> K4 382.74 (+11.8%); | |
| fp32 K0 87.24 -> K16 177.34 (+103%); 4-bit int4 still | |
| peaks near K0-K2, fp4 peaks K2. | |
| HF same harness, threads swept, median-of-3: eager | |
| 42.72/45.59/43.42 (PPL 156.68), blocks-only dynamic-int8 | |
| 34.74/43.88/51.26 (4T 51.22 PPL 158.42 +1.1% ok — fixed by | |
| quantizing `blocks` Linears only; full-model quantize fails | |
| on `FactorizedTiedHead.proj.weight.t()`), compile | |
| threaded-FWKV-state dynamic=True 49.73 warmup 45.9s, | |
| compile-int8 same path 46.93 warmup 36.2s, torchao-int4 | |
| `Requires mslk>=1.0.0` (no Windows wheel). Speedups vs | |
| eager 43.42: int8 K0 7.88x, best K4 8.81x; vs compile | |
| 49.73: K0 6.88x, best 7.70x. Previous run (v1 drafter): | |
| K0 59.82/83.70/87.54 etc., best int8 K8 359.97 — v4 gains | |
| come from restored accept (19%->85% @K2). | |
| ### Spec drafter v4 + doom-loop check (2026-09-15 late night) | |
| v1 (most-recent single match, verbatim tail truncated at history | |
| end): Surjo bench 0% (correct text, nothing to match), Supra | |
| bench 100% (degenerate loop flatters it). v2 (pure frequency | |
| vote) collapsed on drifting text (FWKV bench int8 K2 | |
| 85%->19%, K8 69%->9%). v3 (recency-decayed vote) same failure. | |
| v4 (landed): most-recent strictly-inside occurrence per | |
| position + iterative extension (fixes v1 truncation) + | |
| rejected-block latch REMOVED (the 32-propose/<1/3-accept | |
| latch fired during the short-history warmup and blinded the | |
| session to later 85%-accept regions; rejected blocks stream | |
| weights once per block via forward_block, so worst case on | |
| 0%-accept text is ~K0 speed). Shared by dense/Surjo/FWKV; | |
| `tests/test_native.py::test_speculative_draft_hits_and_misses` | |
| reimplemented against v4; suite 759 passed; K0==Kx exact | |
| everywhere probed. | |
| Doom-loop probe (greedy, 128 tok, fp32+int8, 3 prompts): | |
| Surjo stops coherently on bench/code prompts (48/15 tok, | |
| uniq 0.73/0.93, loop 0) but loops story prompts (128 tok, | |
| uniq 0.09-0.16, 5/11-period loops); Supra loops ALL prompts | |
| (128 tok, uniq 0.07-0.11, 7/9/12-period loops — bench, | |
| story, AND code); FWKV never hard-loops bench/story | |
| (uniq 0.28-0.54, loop 0; int8 bench drifts into a 10-period | |
| body/brain cycle) but degenerates to 128 spaces on code | |
| (uniq 0.02 both precisions). So Supra's 100% spec accept is | |
| a looping artifact, Surjo's 0% is healthy text, FWKV sits | |
| between. No repetition penalty exists anywhere (greedy | |
| only); spec accept must always be read next to uniqueness. | |
| ### Portable kernels: Intel/AMD tiers + ARM64 phones (2026-09-16) | |
| Tier ladder (runtime-dispatched, scalar fallback always works): | |
| `scalar` (any CPU) < `avx2`+FMA (x86-64 baseline, all current | |
| numbers) < `avx-vnni` (Zen4/Zen5, Alder Lake+: int8 VNNI fast | |
| path) < `avx512-vnni` (server AVX512) / `neon` (ARM64 phones, | |
| Apple Silicon, Raspberry Pi). New files: `native/kernels_vnni.cpp` | |
| (dpbusd int8xint8, int8-activation ABI via `quantize_row_i8`), | |
| `native/kernels_neon.cpp` (fp32/int8/int4/fp4 + uzp permute), | |
| CPUID detection in `native/kernels.cpp` (AVX-VNNI = leaf7:1 | |
| EAX[4], AVX512-VNNI = leaf7:0 EBX[11] + foundation + ZMM state; | |
| MSVC VNNI TU is EVEX so it dispatches only on AVX512-VNNI, | |
| GCC VEX on AVX-VNNI), per-file codegen flags in CMakeLists | |
| (dispatcher/runtime stay generic). `kernel_name()` keeps the | |
| fp32 family; new `kernel_variant()` + `has_avx_vnni` / | |
| `has_avx512_vnni` / `has_neon` bindings; `cpu_features` gains | |
| `avx_vnni`; `Model.info["kernels"]` gains `variant`. | |
| `tests/test_kernels_portable.py` (5 passed): tier reporting, | |
| flag consistency, pre-VNNI fallback untouched, i8 quantizer | |
| round-trip bounds, dpbusd bias-identity in numpy. | |
| Proven on Ryzen 5 5600: builds clean (MSVC /arch:AVX512 VNNI TU | |
| accepted), `kernel avx2 / variant avx2 / vnni False`, | |
| suite 764 passed (759 + 5 new) — zero behavior change, as | |
| designed (VNNI selector is null here so Matrix keeps the int16 | |
| Q8 / fp32-activation paths byte-for-byte). Pending hardware CI: | |
| VNNI tok/s + int8-activation PPL re-gate on Zen4/ADL+ (expect | |
| prefill/spec-verify wins ~2-3x integer MACs; decode stays | |
| bandwidth-bound so batch-1 gains concentrate in K>0), NEON | |
| tok/s + parity on ARM64 hardware (expects the usual | |
| Silicon/Android numbers; tails fall back to scalar by | |
| construction). No speedup is claimed for either until measured. | |