Benchmark artifacts for the tiyuvta serving engine

Reproduction artifacts for the performance claims of the tiyuvta serving engine (OpenAI-compatible serving), tuned for RTX PRO 6000 Blackwell (sm_120a) and RTX 5090, with a compile-gated H100 (sm_90a) lane. If a claim depends on a specific artifact, the artifact is public here.

Refreshed 2026-08-21 against serving engine commit 73d79a3cc8 (release v0.99.0 at that refresh). The engine and its env prefix were both renamed since this card's first revision: every BW24_* incantation in older revisions of this card is dead.

Board update β€” v0.101.0 headline cell (added 2026-08-22)

Serving engine release v0.101.0 (tag 23d5ae3ffb) lands the DFlash2 drafter and the round-cost engine together; its headline cell, quoted from the release's own measured receipts (the "DFlash2 drafter and the round-cost engine" release entry):

Cell Result Protocol
Qwen3.8-27B NVFP4, DSpark route + DFlash2 drafter, c=1 accept 3.57 tok/round at 13.7 ms/round β‰ˆ 260 tok/s decode-class greedy; 217–240 at t0.6 with thinking on; spec-vs-plain 2.09Γ— at engine terms RTX PRO 6000 Blackwell (sm_120a), agentic pack, single stream. Bench context, opt-in route (the served default remains the MTP/FR-Spec route). Accept vectors pinned per-request byte-identical to the pre-merge banks across both drafters, both window arms, greedy and t0.6.

This is an engine bench claim on a named card under stated conditions β€” not a hosted-service throughput figure.

Drafts (drafts/<model>/) β€” the standard regime, one file per model

Every spec board row runs one trimmed draft file built by the documented draft regime: FR-Spec ranks derived from the model's own generations (never transferred between models β€” foreign ranks measured βˆ’12 acceptance pts on an identical tokenizer), MTP block extracted byte-verbatim from the published model GGUF, head requantized NVFP4 after trimming (measured zero acceptance cost), block Q4_K_M (measured faster AND higher acceptance than Q8_0). The serving engine attaches the draft file next to its trunk with no other flags (winners are defaults). Proof of attach is a log line, not the absence of an error: [mtp-draft] loading external MTP draft: <path> or [worker] <name>: regime draft attached (<path>).

Directory Source model (exact bytes) Tracked board row (e2e tok/s short / medium / long-agentic)
drafts/qwen35-9b-nvfp4/ Qwen3.5-9B NVFP4 MTP GGUF K=3: 281.0 / 211.7 / 187.1
drafts/qwen36-27b-nvfp4/ nvidia/Qwen3.6-27B-NVFP4 β†’ Q4_K_M GGUF K=3: 116.4 / 101.2 / 86.0
drafts/qwen36-35b-a3b-iq4xs/ unsloth Qwen3.6-35B-A3B UD-IQ4_XS K=2: 302.4 / 253.0 / 270.7
drafts/qwen36-27b-unsloth-nvfp4/ unsloth/Qwen3.6-27B-NVFP4 β†’ GGUF no board row β€” plain parity with the nvidia artifact; spec long-agentic win / medium loss (jsonl 2026-07-17)

Rows are the tracked RTX 5090 Laptop spec cells at serving engine commit 73d79a3cc8 (medians of N=5 same-session interleaved reps, measured 2026-08-02, no flags). The long-agentic column is SAMPLED (temp 0.7): rejection-sampling spec decode is distribution-exact; short/medium are greedy.

Each directory carries the draft (draft-owntrim-nvfp4head-q4blk.gguf) and the rank file it was built from (owngen-ranks-32768.txt, one token id per line, rank order β€” derived 2026-07-17/18, 218-prompt mixed corpus, ~110k own-generated tokens per model).

Use ours for these exact models. For any other model, requant, or finetune, build your own (a finetune's distribution moved, so its draft must too): frspec-owngen derives 32,768 ranks from the model's OWN generations, and make-trimmed-draft.sh extracts, trims and quantizes the head into the draft file. Validate before trusting: frspec-owngen --validate A/Bs baseline-vs-trimmed spec e2e and prints a GOOD/WASH/BAD verdict.

drafts/kf4/ β€” archived experiment, no released config consumes these

These directories were built for an experimental NVFP4 KV-cache format arm (kv-fp4 lane, 2026-07-20). That arm was never merged: at serving engine commit 73d79a3cc8 the KV-cache format accepts K = q8_0 | fp8 and V = q5_1 | q4_0 | fp8 only (defaults q8_0/q5_1). The files are kept as evidence of regime law 1 β€” a new numeric config re-derives each model's draft from its own generations under that config. That lane's 2026-07-20 verdicts, for the record: 9B wash, 27B kf4-draft win (+2.5% e2e), 35B old-draft stays.

Gemma ranks (drafts/gemma4-*/)

Gemma drafters are separate assistant GGUFs, not NextN/MTP heads β€” the serving engine attaches them through its assistant-drafter seam (the seam trap: the qwen-family MTP attach refuses this format; with a drafter attached, the gemma spec route arms at K=5 by default). The FR-Spec trim applies at load from the ranks file; the serve-time adaptive trim (on by default with ranks β€” 512 spare head slots; a static trim is the alternative) learns coverage escapes live from prompt + verify tokens and persists them to <ranks>.learned. Own-gen rank files per model (regime law 1):

Directory Model Note
drafts/gemma4-26b-a4b-qat/ Gemma-4 26B-A4B QAT Q4_0 trim adopted (own-gen e2e-neutral vs corpus ranks, correct provenance)
drafts/gemma4-31b-qat/ Gemma-4 31B QAT Q4_0 the adaptive trim flipped the 31B chat cell from βˆ’17% to +2.5% and made trim β‰₯ untrimmed on every measured cell (2026-07-19)
drafts/gemma4-e4b-qat/ Gemma-4 E4B QAT Q4_0 ranks published; serving stays untrimmed (measured e2e wash β€” small head)

Legacy trims (top level, Qwen3.6-27B)

The pre-regime artifacts the earlier boards used β€” superseded by drafts/ but kept because published claims referenced them:

File Ranking Status
mtp-Qwen3.6-27B-Q4_K_M-frspec32768.gguf generic corpus-frequency superseded by drafts/qwen36-27b-nvfp4/
mtp-Qwen3.6-27B-Q4_K_M-frspec-code75-32768.gguf code-skewed corpus superseded
mtp-Qwen3.6-27B-Q4_K_M-frspec-balanced32768.gguf balanced corpus superseded

Prompts + configs

prompts/ holds the exact board prompts β€” p1-code-short.txt, p2-code-medium.txt, p3-agentic-long.txt plus the well-formed p3-agentic-long-v2.txt the sampled protocol uses β€” and the long-context depth documents p4-16k.txt / p5-32k.txt / p6-64k.txt. CONFIGS.md carries the per-cell engine configs. Protocol: interleaved same-session reps (Nβ‰₯2–3 medians, both orders), power state pinned per window.

Related

Downloads last month
2,664
GGUF
Model size
32.8k params
Architecture
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support