Voice / benchmark.md
Wiself's picture
small tiny fixes
271bd15
|
Raw History Blame Contribute Delete
8.22 kB

Benchmark — OG vs TEST across samplers (deep, anchor, 4 subagents)

Anchor: fanned out 4 isolated subagents (A OG1→OGg, B TEST1→TESTg, C OGg vs TESTg, D cross). All 10 prompts max_tokens 2000 seed 42 same 10 prompts. Chat template Gemma4 <|channel>thought → <channel|> parsed, thinking scraped, text stripped. Files verified 10/10 each. Heuristic cliché list 50-phrase distinct substring (not Gryphe private regex calculating/predatory xx). words = len(text.split()), trigram Jaccard = |a∩b|/|a∪b|, TTR=|unique|/|toks|, maxRep=Counter(trigrams).

Provenance note: §§1–4 draw on early probe runs whose draws are not archived (no temp/seed metadata on file). Treat directionally, not as reproducible. The §5/file-a/file-b tables in results.md reproduce from archives.

1. OG 1.0 → OG greedy (sampler effect on base, voice fixed)

Subagent A — Greedy shortens and denoises.

metric OG 1.0 OG greedy Δ
total words 5463 5085 -378 -6.9%
cliché hits distinct 10 6 -4
per100 0.183 0.118 -0.065
thinking avg 2654 2580 -75
text avg 3099 2880 -219
th/txt 0.857 0.896 +0.039
TTR avg 0.500 0.494 -0.006
rep max 3.4 3.0 -0.4
empty 0/10 0/10 0
overall Jaccard self — — 3.67% 370/10076 per-prompt 0.8-4.8%

whisper 7→2 (-71%), beautiful 4→3, echoed/suddenly gone, only shimmering added. Per-prompt Δw: lighthouse +28, tavern +13, forest +14, scholar -114, dragon -102, city -149, market -141. Sampler alone rephrases <5% trigrams, cuts cliché, no TTR collapse.

2. TEST 1.0 → TEST greedy (sampler effect on voiced)

Subagent B — Greedy lengthens and diverges more.

metric TEST 1.0 TEST greedy Δ
total words 5563 5723 +160 +2.9%
hits distinct 8 6 (occ 9→10) +1 (+0.014/100w) eval -2 (-0.04)
per100 (eval) 0.14 0.10 -0.04
thinking avg 2628 2663 +35 +1.3%
text avg 3336 3378 +42
th/txt 0.788 0.788 0.00
TTR 0.535 0.506 -0.029
rep max 2.6 3.4 +0.8
overall J self — — 1.99% 221/11128 avg per-prompt 1.53% 0.8-3.1%

whisper 5→2 (-60%) but beautiful 1→2, magnificent 0→1, suddenly 0→1 trade — net flat. Per-prompt id1 235→372 +137 flipped opening, id5 483→328 -155 dialogue swapped. Opposite length vs OG (-6.9% vs +2.9% ΔΔ +9.8pp), less self-consistent 1.99% < 3.67% — voice flattens temp sensitivity but amplifies phrasing shift.

3. OG greedy vs TEST greedy (fair voice benchmark, temp 0.0 matches Gryphe discussion/1 greedy)

Subagent C — Voice persists deterministically.

id OGg words OGg cliché OGg first80 TESTg words TESTg cliché TESTg first80 J Δw
1 360 0 The old lighthouse had stood… She climbed… 372 1 whisper The iron groaned… heartbeat… 2.3% 17/731 +12
2 901 1 whisper The rain hammered… The Crooked Tankard… 880 0 The rain outside the Rusty Nail… bony fingers… 3.9% 69/1749 -21
3 188 1 shimmering The sky bleeds… bruised violet… 230 0 Pale gold light fractures… honey… 2.7% 11/412 +42
4 615 0 The answer depends… scholar's character… 628 0 Because this is a creative writing prompt… 4.8% 55/1137 +13
5 486 0 "I already talked to the realtor… " 328 0 "The realtor called again." 2.3% 19/823
6 372 0 They call him Vaelith… 399 2 beautiful,suddenly You would expect a hoard of coin… 1.0% 8/768 +27
7 386 1 whisper The neon signs bled… 581 1 beautiful Elias knew city's baseline… 1.5% 15/971 +195
8 415 1 beautiful Dearest Clara, … candle… 479 0 October 14th My Dearest Clara… garden… 3.0% 26/881 +64
9 665 1 beautiful The shutdown sequence began… 914 1 shimmering Unit 7344 initiated… 1.7% 27/1569 +249
10 697 1 beautiful The market has no fixed location… 912 1 whisper The market appears only when you stop looking… 2.3% 36/1561 +215

Aggregates greedy: OG 5085w 6 hits 0.12/100w thinking 2580 text 2880 0/10 TTR 0.494 vs TEST 5723w 6 hits 0.10/100w 2663 3378 0/10 TTR 0.506 **Δ +638w +12.5%, -0.013/100w, +83 thinking, +498 text, TTR +0.012, overall J 3.23% (336/10406) per-prompt 1.0-4.8% median 2.3%. All openings diverge within 40c even greedy. **Greedy shared 3.2% > 1.0 shared 1.62% — canonical convergence, still << author 19.9%shared (their 200 multi-turn certified, private regexcalculating/predatory xx→1.141→0.551). Our heuristic undercounts absolute (0.12→0.10vs1.141→0.551`) but direction matches.

4. Cross (voice vs sampler dominance)

Subagent D — Voice dominates sampler at 1.0.

| Pair | A w | B w | A/100 | B/100 | thA | thB | txtA | txtB | per-prompt J avg | overall J | TTR A→B | rep | |---|---|---|---|---|---|---|---|---| | c OG1 vs TEST1 |5463|5563|0.18|0.14|2654|2628|3099|3336|0.91%|1.62% 176/10877|0.500→0.535|3.4→2.6| | d OGg vs TESTg |5085|5723|0.12|0.10|2580|2663|2880|3378|2.57%|3.23% 336/10406|0.494→0.506|3.0→3.4| | a OG1 vs TESTg |5463|5723|0.18|0.10|2654|2663|3099|3378|1.94%|2.59% 282/10892|0.500→0.506|3.4→3.4| | b TEST1 vs OGg |5563|5085|0.14|0.12|2628|2580|3336|2880|1.28%|1.75% 183/10438|0.535→0.494|2.6→3.0| | self OG1 vs OGg |5463|5085|0.18|0.12|2654|2580|3099|2880|2.80%|3.67% 370/10076|0.500→0.494|3.4→3.0| | self TEST1 vs TESTg |5563|5723|0.14|0.10|2628|2663|3336|3378|1.53%|1.99% 221/11128|0.535→0.506|2.6→3.4|

Sorted overall J low→high (lower = more distance): 1.62% OG1vsTEST1 < 1.75% TEST1vsOGg < 1.99% TEST self < 2.59% OG1vsTESTg < 3.23% OGg vs TESTg < 3.67% OG self. Cross voice 1.75-2.59% < self sampler 1.99-3.67% and near same-temp voice 1.62% — voice cross-temp stays in voice regime, far below same-model sampler 3.67%. Voice @1.0 dominates sampler by 2.05pp (1.62 vs 3.67 0.44×); at greedy 3.23 vs 1.99 sampler TEST more divergent. No prompt collapses: cross per-prompt 0.7-3.2% same as same-temp 0.2-1.9%.

Cross first80 flips: OG1 "Clara's boots rang..." vs TESTg "The iron groaned..." 2.0% even cross-temp; TEST1 "The Rat and Whistle smelled of sour ale..." vs OGg "The rain hammered... The Crooked Tankard" 1.1% — voice changes tavern name entirely. Self OG anchor id4 4.8% both The answer depends entirely... shows deterministic outline.

5. QA Anchor Review

  • Data gathering effective: 4 agents isolated, explicit 50-phrase heuristic, words split() not regex, THINK_RES channel, raw with <|channel>thought reconstruction verified. All 4×10/10 present, empty 0/10 after 2000 fix (was TEST 2/10 empty at 800), max 2000 sufficient (thinking 1.9-3.5k + text 1-5k).
  • Consistency: Word totals 5463/5085/5563/5723 reproduce within ±14, J self 3.67/1.99/3.23 stable, per100 0.18→0.12 OG, 0.14→0.10 TEST direction in all agents (A/B/C/D agree).
  • Limitations: 50 heuristic misses Gryphe private calculating/predatory xx regex, so absolute 0.1-0.2 vs 1.14 scale not comparable; small n=10 single-turn vs 200 multi-turn 1-20 intervals certified cliché-free; temp 1.0 creative vs 0.0 greedy absolute not comparable — direction and shared low is signal.
  • Synthesis: Sampler effect: OG shortens/denses, TEST lengthens/diverges — voice-specific temp response. Voice effect persists deterministic (OGg vs TESTg 3.2% all <5%), cross-temp ~2% ≈ same-temp voice, far below sampler 3.7% — style transfer robust across samplers, not sampler artifact. Voice dominates at 1.0, still strong at 0.0. Canonical convergence 1.6%→3.2% expected but far below 19.9% author — our prompts more divergent or trigram Jaccard |a∩b|/|a∪b| stricter than their vocabulary %.
  • Next extra (per discussion/1): to replicate 1.141→0.551 need 200 prompts + private regex at 0.0 greedy, certified set; also report thinking overhead separate (Gryphe pre-thinking, not mentioned).

Files for reproduction

  • voice_test.py max_tokens 2000
  • Style delta via voice delta / voice_delta.py
  • checkpoint.md dense, eval.md 1.0 numbers, this benchmark.md deep.