Download benchmark.md from Wiself/Voice: direct link, hf CLI and curl.
- Browser
- Download file 8.22 kB
-
https://huggingface.co/Wiself/Voice/resolve/main/benchmark.md
- Command line
-
hf download hf://Wiself/Voice/benchmark.md
-
curl -L -o benchmark.md https://huggingface.co/Wiself/Voice/resolve/main/benchmark.md
Benchmark — OG vs TEST across samplers (deep, anchor, 4 subagents)
Anchor: fanned out 4 isolated subagents (A OG1→OGg, B TEST1→TESTg, C OGg vs TESTg, D cross). All 10 prompts
max_tokens 2000seed 42same 10 prompts. Chat template Gemma4<|channel>thought → <channel|>parsed,thinkingscraped,textstripped. Files verified10/10each. Heuristic cliché list 50-phrase distinct substring (not Gryphe private regexcalculating/predatory xx).words = len(text.split()),trigram Jaccard = |a∩b|/|a∪b|,TTR=|unique|/|toks|,maxRep=Counter(trigrams).
Provenance note: §§1–4 draw on early probe runs whose draws are not archived (no temp/seed metadata on file). Treat directionally, not as reproducible. The §5/file-a/file-b tables in results.md reproduce from archives.
1. OG 1.0 → OG greedy (sampler effect on base, voice fixed)
Subagent A — Greedy shortens and denoises.
| metric | OG 1.0 | OG greedy | Δ |
|---|---|---|---|
| total words | 5463 | 5085 | -378 -6.9% |
| cliché hits distinct | 10 | 6 | -4 |
| per100 | 0.183 | 0.118 | -0.065 |
| thinking avg | 2654 | 2580 | -75 |
| text avg | 3099 | 2880 | -219 |
| th/txt | 0.857 | 0.896 | +0.039 |
| TTR avg | 0.500 | 0.494 | -0.006 |
| rep max | 3.4 | 3.0 | -0.4 |
| empty | 0/10 | 0/10 | 0 |
| overall Jaccard self | — | — | 3.67% 370/10076 per-prompt 0.8-4.8% |
whisper 7→2 (-71%), beautiful 4→3, echoed/suddenly gone, only shimmering added. Per-prompt Δw: lighthouse +28, tavern +13, forest +14, scholar -114, dragon -102, city -149, market -141. Sampler alone rephrases <5% trigrams, cuts cliché, no TTR collapse.
2. TEST 1.0 → TEST greedy (sampler effect on voiced)
Subagent B — Greedy lengthens and diverges more.
| metric | TEST 1.0 | TEST greedy | Δ |
|---|---|---|---|
| total words | 5563 | 5723 | +160 +2.9% |
| hits distinct | 8 | 6 (occ 9→10) | +1 (+0.014/100w) eval -2 (-0.04) |
| per100 (eval) | 0.14 | 0.10 | -0.04 |
| thinking avg | 2628 | 2663 | +35 +1.3% |
| text avg | 3336 | 3378 | +42 |
| th/txt | 0.788 | 0.788 | 0.00 |
| TTR | 0.535 | 0.506 | -0.029 |
| rep max | 2.6 | 3.4 | +0.8 |
| overall J self | — | — | 1.99% 221/11128 avg per-prompt 1.53% 0.8-3.1% |
whisper 5→2 (-60%) but beautiful 1→2, magnificent 0→1, suddenly 0→1 trade — net flat. Per-prompt id1 235→372 +137 flipped opening, id5 483→328 -155 dialogue swapped. Opposite length vs OG (-6.9% vs +2.9% ΔΔ +9.8pp), less self-consistent 1.99% < 3.67% — voice flattens temp sensitivity but amplifies phrasing shift.
3. OG greedy vs TEST greedy (fair voice benchmark, temp 0.0 matches Gryphe discussion/1 greedy)
Subagent C — Voice persists deterministically.
| id | OGg words | OGg cliché | OGg first80 | TESTg words | TESTg cliché | TESTg first80 | J | Δw |
|---|---|---|---|---|---|---|---|---|
| 1 | 360 | 0 | The old lighthouse had stood… She climbed… | 372 | 1 whisper |
The iron groaned… heartbeat… | 2.3% 17/731 | +12 |
| 2 | 901 | 1 whisper |
The rain hammered… The Crooked Tankard… | 880 | 0 | The rain outside the Rusty Nail… bony fingers… | 3.9% 69/1749 | -21 |
| 3 | 188 | 1 shimmering |
The sky bleeds… bruised violet… | 230 | 0 | Pale gold light fractures… honey… | 2.7% 11/412 | +42 |
| 4 | 615 | 0 | The answer depends… scholar's character… | 628 | 0 | Because this is a creative writing prompt… | 4.8% 55/1137 | +13 |
| 5 | 486 | 0 | "I already talked to the realtor… | " | 328 | 0 | "The realtor called again." | 2.3% 19/823 |
| 6 | 372 | 0 | They call him Vaelith… | 399 | 2 beautiful,suddenly |
You would expect a hoard of coin… | 1.0% 8/768 | +27 |
| 7 | 386 | 1 whisper |
The neon signs bled… | 581 | 1 beautiful |
Elias knew city's baseline… | 1.5% 15/971 | +195 |
| 8 | 415 | 1 beautiful |
Dearest Clara, … candle… | 479 | 0 | October 14th My Dearest Clara… garden… | 3.0% 26/881 | +64 |
| 9 | 665 | 1 beautiful |
The shutdown sequence began… | 914 | 1 shimmering |
Unit 7344 initiated… | 1.7% 27/1569 | +249 |
| 10 | 697 | 1 beautiful |
The market has no fixed location… | 912 | 1 whisper |
The market appears only when you stop looking… | 2.3% 36/1561 | +215 |
Aggregates greedy: OG 5085w 6 hits 0.12/100w thinking 2580 text 2880 0/10 TTR 0.494 vs TEST 5723w 6 hits 0.10/100w 2663 3378 0/10 TTR 0.506 **Δ +638w +12.5%, -0.013/100w, +83 thinking, +498 text, TTR +0.012, overall J 3.23% (336/10406) per-prompt 1.0-4.8% median 2.3%. All openings diverge within 40c even greedy. **Greedy shared 3.2% > 1.0 shared 1.62% — canonical convergence, still << author 19.9%shared (their 200 multi-turn certified, private regexcalculating/predatory xx→1.141→0.551). Our heuristic undercounts absolute (0.12→0.10vs1.141→0.551`) but direction matches.
4. Cross (voice vs sampler dominance)
Subagent D — Voice dominates sampler at 1.0.
| Pair | A w | B w | A/100 | B/100 | thA | thB | txtA | txtB | per-prompt J avg | overall J | TTR A→B | rep | |---|---|---|---|---|---|---|---|---| | c OG1 vs TEST1 |5463|5563|0.18|0.14|2654|2628|3099|3336|0.91%|1.62% 176/10877|0.500→0.535|3.4→2.6| | d OGg vs TESTg |5085|5723|0.12|0.10|2580|2663|2880|3378|2.57%|3.23% 336/10406|0.494→0.506|3.0→3.4| | a OG1 vs TESTg |5463|5723|0.18|0.10|2654|2663|3099|3378|1.94%|2.59% 282/10892|0.500→0.506|3.4→3.4| | b TEST1 vs OGg |5563|5085|0.14|0.12|2628|2580|3336|2880|1.28%|1.75% 183/10438|0.535→0.494|2.6→3.0| | self OG1 vs OGg |5463|5085|0.18|0.12|2654|2580|3099|2880|2.80%|3.67% 370/10076|0.500→0.494|3.4→3.0| | self TEST1 vs TESTg |5563|5723|0.14|0.10|2628|2663|3336|3378|1.53%|1.99% 221/11128|0.535→0.506|2.6→3.4|
Sorted overall J low→high (lower = more distance): 1.62% OG1vsTEST1 < 1.75% TEST1vsOGg < 1.99% TEST self < 2.59% OG1vsTESTg < 3.23% OGg vs TESTg < 3.67% OG self. Cross voice 1.75-2.59% < self sampler 1.99-3.67% and near same-temp voice 1.62% — voice cross-temp stays in voice regime, far below same-model sampler 3.67%. Voice @1.0 dominates sampler by 2.05pp (1.62 vs 3.67 0.44×); at greedy 3.23 vs 1.99 sampler TEST more divergent. No prompt collapses: cross per-prompt 0.7-3.2% same as same-temp 0.2-1.9%.
Cross first80 flips: OG1 "Clara's boots rang..." vs TESTg "The iron groaned..." 2.0% even cross-temp; TEST1 "The Rat and Whistle smelled of sour ale..." vs OGg "The rain hammered... The Crooked Tankard" 1.1% — voice changes tavern name entirely. Self OG anchor id4 4.8% both The answer depends entirely... shows deterministic outline.
5. QA Anchor Review
- Data gathering effective: 4 agents isolated, explicit 50-phrase heuristic,
words split()not regex,THINK_RESchannel,rawwith<|channel>thoughtreconstruction verified. All4×10/10present,empty 0/10after2000fix (wasTEST 2/10 emptyat 800),max 2000sufficient (thinking 1.9-3.5k+text 1-5k). - Consistency: Word totals
5463/5085/5563/5723reproduce within±14, J self3.67/1.99/3.23stable, per1000.18→0.12OG,0.14→0.10TEST direction in all agents (A/B/C/D agree). - Limitations:
50heuristic misses Gryphe privatecalculating/predatory xxregex, so absolute0.1-0.2vs1.14scale not comparable; smalln=10single-turn vs200multi-turn1-20intervals certified cliché-free;temp 1.0creative vs0.0greedy absolute not comparable — direction andsharedlow is signal. - Synthesis: Sampler effect: OG shortens/denses, TEST lengthens/diverges — voice-specific temp response. Voice effect persists deterministic (
OGg vs TESTg 3.2%all<5%), cross-temp~2%≈ same-temp voice, far below sampler3.7%— style transfer robust across samplers, not sampler artifact. Voice dominates at1.0, still strong at0.0. Canonical convergence1.6%→3.2%expected but far below19.9%author — our prompts more divergent or trigram Jaccard|a∩b|/|a∪b|stricter than their vocabulary %. - Next extra (per discussion/1): to replicate
1.141→0.551need200prompts + private regex at0.0greedy, certified set; also reportthinkingoverhead separate (Gryphe pre-thinking, not mentioned).
Files for reproduction
voice_test.pymax_tokens 2000- Style delta via
voice delta/voice_delta.py checkpoint.mddense,eval.md1.0 numbers, thisbenchmark.mddeep.