Benchmark artifacts for the tiyuvta serving engine
Reproduction artifacts for the performance claims of the tiyuvta serving engine
(OpenAI-compatible serving), tuned for RTX PRO 6000 Blackwell (sm_120a) and RTX 5090,
with a compile-gated H100 (sm_90a) lane. If a claim depends on a specific artifact, the artifact is
public here.
Refreshed 2026-08-21 against serving engine commit 73d79a3cc8 (release v0.99.0 at that
refresh). The engine and its env prefix were both renamed since this card's first revision:
every BW24_* incantation in older revisions of this card is dead.
Board update β v0.101.0 headline cell (added 2026-08-22)
Serving engine release v0.101.0 (tag 23d5ae3ffb) lands the DFlash2 drafter and the
round-cost engine together; its headline cell, quoted from the release's own measured
receipts (the "DFlash2 drafter and the round-cost engine" release entry):
| Cell | Result | Protocol |
|---|---|---|
| Qwen3.8-27B NVFP4, DSpark route + DFlash2 drafter, c=1 | accept 3.57 tok/round at 13.7 ms/round β 260 tok/s decode-class greedy; 217β240 at t0.6 with thinking on; spec-vs-plain 2.09Γ at engine terms | RTX PRO 6000 Blackwell (sm_120a), agentic pack, single stream. Bench context, opt-in route (the served default remains the MTP/FR-Spec route). Accept vectors pinned per-request byte-identical to the pre-merge banks across both drafters, both window arms, greedy and t0.6. |
This is an engine bench claim on a named card under stated conditions β not a hosted-service throughput figure.
Drafts (drafts/<model>/) β the standard regime, one file per model
Every spec board row runs one trimmed draft file built by the documented draft regime:
FR-Spec ranks derived from the model's own generations (never transferred between
models β foreign ranks measured β12 acceptance pts on an identical tokenizer), MTP block
extracted byte-verbatim from the published model GGUF, head requantized NVFP4 after
trimming (measured zero acceptance cost), block Q4_K_M (measured faster AND higher
acceptance than Q8_0). The serving engine attaches the draft file next to its trunk with no
other flags (winners are defaults). Proof of attach is a log line, not the absence of an
error:
[mtp-draft] loading external MTP draft: <path> or
[worker] <name>: regime draft attached (<path>).
| Directory | Source model (exact bytes) | Tracked board row (e2e tok/s short / medium / long-agentic) |
|---|---|---|
drafts/qwen35-9b-nvfp4/ |
Qwen3.5-9B NVFP4 MTP GGUF | K=3: 281.0 / 211.7 / 187.1 |
drafts/qwen36-27b-nvfp4/ |
nvidia/Qwen3.6-27B-NVFP4 β Q4_K_M GGUF | K=3: 116.4 / 101.2 / 86.0 |
drafts/qwen36-35b-a3b-iq4xs/ |
unsloth Qwen3.6-35B-A3B UD-IQ4_XS | K=2: 302.4 / 253.0 / 270.7 |
drafts/qwen36-27b-unsloth-nvfp4/ |
unsloth/Qwen3.6-27B-NVFP4 β GGUF | no board row β plain parity with the nvidia artifact; spec long-agentic win / medium loss (jsonl 2026-07-17) |
Rows are the tracked RTX 5090 Laptop spec cells at serving engine commit 73d79a3cc8
(medians of N=5 same-session interleaved reps, measured 2026-08-02, no flags). The
long-agentic column is SAMPLED (temp 0.7): rejection-sampling spec decode is
distribution-exact; short/medium are greedy.
Each directory carries the draft (draft-owntrim-nvfp4head-q4blk.gguf) and the rank file
it was built from (owngen-ranks-32768.txt, one token id per line, rank order β derived
2026-07-17/18, 218-prompt mixed corpus, ~110k own-generated tokens per model).
Use ours for these exact models. For any other model, requant, or finetune, build your
own (a finetune's distribution moved, so its draft must too): frspec-owngen derives
32,768 ranks from the model's OWN generations, and make-trimmed-draft.sh extracts, trims
and quantizes the head into the draft file. Validate before trusting: frspec-owngen --validate A/Bs baseline-vs-trimmed spec e2e and prints a GOOD/WASH/BAD verdict.
drafts/kf4/ β archived experiment, no released config consumes these
These directories were built for an experimental NVFP4 KV-cache format arm (kv-fp4 lane,
2026-07-20). That arm was never merged: at serving engine commit 73d79a3cc8 the
KV-cache format accepts K = q8_0 | fp8 and V = q5_1 | q4_0 | fp8 only (defaults
q8_0/q5_1). The files are kept as evidence of regime law 1 β a new numeric config
re-derives each model's draft from its own generations under that config. That lane's
2026-07-20 verdicts, for the record: 9B wash, 27B kf4-draft win (+2.5% e2e), 35B old-draft
stays.
Gemma ranks (drafts/gemma4-*/)
Gemma drafters are separate assistant GGUFs, not NextN/MTP heads β the serving engine
attaches them through its assistant-drafter seam (the seam trap: the qwen-family MTP attach
refuses this format; with a drafter attached, the gemma spec route arms at K=5 by default).
The FR-Spec trim applies at load from the ranks file; the serve-time adaptive trim (on by
default with ranks β 512 spare head slots; a static trim is the alternative) learns coverage
escapes live from prompt + verify tokens and persists them to <ranks>.learned. Own-gen
rank files per model (regime law 1):
| Directory | Model | Note |
|---|---|---|
drafts/gemma4-26b-a4b-qat/ |
Gemma-4 26B-A4B QAT Q4_0 | trim adopted (own-gen e2e-neutral vs corpus ranks, correct provenance) |
drafts/gemma4-31b-qat/ |
Gemma-4 31B QAT Q4_0 | the adaptive trim flipped the 31B chat cell from β17% to +2.5% and made trim β₯ untrimmed on every measured cell (2026-07-19) |
drafts/gemma4-e4b-qat/ |
Gemma-4 E4B QAT Q4_0 | ranks published; serving stays untrimmed (measured e2e wash β small head) |
Legacy trims (top level, Qwen3.6-27B)
The pre-regime artifacts the earlier boards used β superseded by drafts/ but kept because
published claims referenced them:
| File | Ranking | Status |
|---|---|---|
mtp-Qwen3.6-27B-Q4_K_M-frspec32768.gguf |
generic corpus-frequency | superseded by drafts/qwen36-27b-nvfp4/ |
mtp-Qwen3.6-27B-Q4_K_M-frspec-code75-32768.gguf |
code-skewed corpus | superseded |
mtp-Qwen3.6-27B-Q4_K_M-frspec-balanced32768.gguf |
balanced corpus | superseded |
Prompts + configs
prompts/ holds the exact board prompts β p1-code-short.txt, p2-code-medium.txt,
p3-agentic-long.txt plus the well-formed p3-agentic-long-v2.txt the sampled protocol
uses β and the long-context depth documents p4-16k.txt / p5-32k.txt / p6-64k.txt.
CONFIGS.md carries the per-cell engine configs. Protocol: interleaved same-session reps
(Nβ₯2β3 medians, both orders), power state pinned per window.
Related
- Flagship (Qwen3.8-27B) ranks + drafts: tiyuvta/Qwen3.8-27B-NVFP4-MTP-GGUF β three ranks flavours; the safetensors path trims at load from the ranks file, no separate draft file.
- Avifenesh/Hy3-REAP-Layer103p5-bw24 β the ~100 GB Hy3 expert overlay the serving engine serves on a 24 GB card via VRAMβRAMβNVMe spill.
- No card? A hosted instance runs at inference.tiyuvta.ai.
- Downloads last month
- 2,664
4-bit