# Benchmark — OG vs TEST across samplers (deep, anchor, 4 subagents) > **Anchor:** fanned out 4 isolated subagents (A OG1→OGg, B TEST1→TESTg, C OGg vs TESTg, D cross). All 10 prompts `max_tokens 2000` `seed 42` same 10 prompts. Chat template Gemma4 `<|channel>thought → ` parsed, `thinking` scraped, `text` stripped. Files verified `10/10` each. Heuristic cliché list 50-phrase distinct substring (not Gryphe private regex `calculating/predatory xx`). `words = len(text.split())`, `trigram Jaccard = |a∩b|/|a∪b|`, `TTR=|unique|/|toks|`, `maxRep=Counter(trigrams)`. > Provenance note: §§1–4 draw on early probe runs whose draws are not archived (no temp/seed metadata on file). Treat directionally, not as reproducible. The §5/file-a/file-b tables in results.md reproduce from archives. ## 1. OG 1.0 → OG greedy (sampler effect on base, voice fixed) *Subagent A* — **Greedy shortens and denoises.** | metric | OG 1.0 | OG greedy | Δ | |---|---|---|---| | total words |5463|5085|-378 -6.9%| | cliché hits distinct |10|6|-4| | per100 |0.183|0.118|-0.065| | thinking avg|2654|2580|-75| | text avg|3099|2880|-219| | th/txt|0.857|0.896|+0.039| | TTR avg|0.500|0.494|-0.006| | rep max|3.4|3.0|-0.4| | empty|0/10|0/10|0| | overall Jaccard self| — | — | **3.67% 370/10076** per-prompt 0.8-4.8% | `whisper 7→2 (-71%)`, `beautiful 4→3`, `echoed/suddenly` gone, only `shimmering` added. Per-prompt Δw: `lighthouse +28`, `tavern +13`, `forest +14`, `scholar -114`, `dragon -102`, `city -149`, `market -141`. **Sampler alone rephrases <5% trigrams, cuts cliché, no TTR collapse.** ## 2. TEST 1.0 → TEST greedy (sampler effect on voiced) *Subagent B* — **Greedy lengthens and diverges more.** | metric | TEST 1.0 | TEST greedy | Δ | |---|---|---|---| | total words |5563|5723|+160 +2.9%| | hits distinct |8|6 (occ 9→10) |+1 (+0.014/100w) eval -2 (-0.04)| | per100 (eval)|0.14|0.10|-0.04| | thinking avg|2628|2663|+35 +1.3%| | text avg|3336|3378|+42| | th/txt|0.788|0.788|0.00| | TTR|0.535|0.506|-0.029| | rep max|2.6|3.4|+0.8| | overall J self| — | — | **1.99% 221/11128 avg per-prompt 1.53% 0.8-3.1%**| `whisper 5→2 (-60%)` but `beautiful 1→2, magnificent 0→1, suddenly 0→1` trade — net flat. Per-prompt `id1 235→372 +137` flipped opening, `id5 483→328 -155` dialogue swapped. **Opposite length vs OG (-6.9% vs +2.9% ΔΔ +9.8pp), less self-consistent 1.99% < 3.67% — voice flattens temp sensitivity but amplifies phrasing shift.** ## 3. OG greedy vs TEST greedy (fair voice benchmark, temp 0.0 matches Gryphe `discussion/1` greedy) *Subagent C* — **Voice persists deterministically.** | id | OGg words | OGg cliché | OGg first80 | TESTg words | TESTg cliché | TESTg first80 | J | Δw | |---|---|---|---|---|---|---|---:|---| |1|360|0|The old lighthouse had stood… She climbed…|372|1 `whisper`|The iron groaned… heartbeat…|2.3% 17/731|+12| |2|901|1 `whisper`|The rain hammered… The Crooked Tankard…|880|0|The rain outside the Rusty Nail… bony fingers…|3.9% 69/1749|-21| |3|188|1 `shimmering`|The sky bleeds… bruised violet…|230|0|Pale gold light fractures… honey…|2.7% 11/412|+42| |4|615|0|The answer depends… scholar's character…|628|0|Because this is a creative writing prompt…|4.8% 55/1137|+13| |5|486|0|"I already talked to the realtor…|" |328|0|"The realtor called again."|2.3% 19/823|-158| |6|372|0|They call him **Vaelith**…|399|2 `beautiful,suddenly`|You would expect a hoard of coin…|1.0% 8/768|+27| |7|386|1 `whisper`|The neon signs bled…|581|1 `beautiful`|Elias knew city's baseline…|1.5% 15/971|+195| |8|415|1 `beautiful`|Dearest Clara, … candle…|479|0|October 14th My Dearest Clara… garden…|3.0% 26/881|+64| |9|665|1 `beautiful`|The shutdown sequence began…|914|1 `shimmering`|Unit 7344 initiated…|1.7% 27/1569|+249| |10|697|1 `beautiful`|The market has no fixed location…|912|1 `whisper`|The market appears only when you stop looking…|2.3% 36/1561|+215| **Aggregates greedy:** `OG 5085w 6 hits 0.12/100w thinking 2580 text 2880 0/10 TTR 0.494` vs `TEST 5723w 6 hits 0.10/100w 2663 3378 0/10 TTR 0.506` **Δ +638w +12.5%, -0.013/100w, +83 thinking, +498 text, TTR +0.012, overall J 3.23% (336/10406) per-prompt 1.0-4.8% median 2.3%`. All openings diverge within 40c even greedy. **Greedy shared 3.2% > 1.0 shared 1.62% — canonical convergence, still << author `19.9%` shared (their 200 multi-turn certified, private regex `calculating/predatory xx` → `1.141→0.551`). Our heuristic undercounts absolute (`0.12→0.10` vs `1.141→0.551`) but direction matches. ## 4. Cross (voice vs sampler dominance) *Subagent D* — **Voice dominates sampler at 1.0.** | Pair | A w | B w | A/100 | B/100 | thA | thB | txtA | txtB | per-prompt J avg | overall J | TTR A→B | rep | |---|---|---|---|---|---|---|---|---| | **c OG1 vs TEST1** |5463|5563|0.18|0.14|2654|2628|3099|3336|0.91%|1.62% 176/10877|0.500→0.535|3.4→2.6| | **d OGg vs TESTg** |5085|5723|0.12|0.10|2580|2663|2880|3378|2.57%|3.23% 336/10406|0.494→0.506|3.0→3.4| | **a OG1 vs TESTg** |5463|5723|0.18|0.10|2654|2663|3099|3378|1.94%|2.59% 282/10892|0.500→0.506|3.4→3.4| | **b TEST1 vs OGg** |5563|5085|0.14|0.12|2628|2580|3336|2880|1.28%|1.75% 183/10438|0.535→0.494|2.6→3.0| | **self OG1 vs OGg** |5463|5085|0.18|0.12|2654|2580|3099|2880|2.80%|3.67% 370/10076|0.500→0.494|3.4→3.0| | **self TEST1 vs TESTg** |5563|5723|0.14|0.10|2628|2663|3336|3378|1.53%|1.99% 221/11128|0.535→0.506|2.6→3.4| Sorted overall J low→high (lower = more distance): `1.62% OG1vsTEST1 < 1.75% TEST1vsOGg < 1.99% TEST self < 2.59% OG1vsTESTg < 3.23% OGg vs TESTg < 3.67% OG self`. Cross voice `1.75-2.59%` < self sampler `1.99-3.67%` and near same-temp voice `1.62%` — **voice cross-temp stays in voice regime, far below same-model sampler 3.67%.** Voice @1.0 dominates sampler by `2.05pp` (`1.62 vs 3.67 0.44×`); at greedy `3.23 vs 1.99` sampler TEST more divergent. No prompt collapses: cross per-prompt `0.7-3.2%` same as same-temp `0.2-1.9%`. Cross first80 flips: `OG1 "Clara's boots rang..." vs TESTg "The iron groaned..."` 2.0% even cross-temp; `TEST1 "The Rat and Whistle smelled of sour ale..." vs OGg "The rain hammered... The Crooked Tankard"` 1.1% — voice changes tavern name entirely. Self OG anchor `id4 4.8%` both `The answer depends entirely...` shows deterministic outline. ## 5. QA Anchor Review - **Data gathering effective:** 4 agents isolated, explicit 50-phrase heuristic, `words split()` not regex, `THINK_RES` channel, `raw` with `<|channel>thought` reconstruction verified. All `4×10/10` present, `empty 0/10` after `2000` fix (was `TEST 2/10 empty` at 800), `max 2000` sufficient (`thinking 1.9-3.5k` + `text 1-5k`). - **Consistency:** Word totals `5463/5085/5563/5723` reproduce within `±14`, J self `3.67/1.99/3.23` stable, per100 `0.18→0.12` OG, `0.14→0.10` TEST direction in all agents (A/B/C/D agree). - **Limitations:** `50` heuristic misses Gryphe private `calculating/predatory xx` regex, so absolute `0.1-0.2` vs `1.14` scale not comparable; small `n=10` single-turn vs `200` multi-turn `1-20` intervals certified cliché-free; `temp 1.0` creative vs `0.0` greedy absolute not comparable — direction and `shared` low is signal. - **Synthesis:** Sampler effect: OG shortens/denses, TEST lengthens/diverges — voice-specific temp response. Voice effect persists deterministic (`OGg vs TESTg 3.2%` all `<5%`), cross-temp `~2%` ≈ same-temp voice, far below sampler `3.7%` — **style transfer robust across samplers, not sampler artifact**. Voice dominates at `1.0`, still strong at `0.0`. Canonical convergence `1.6%→3.2%` expected but far below `19.9%` author — our prompts more divergent or trigram Jaccard `|a∩b|/|a∪b|` stricter than their vocabulary %. - **Next extra (per discussion/1):** to replicate `1.141→0.551` need `200` prompts + private regex at `0.0` greedy, certified set; also report `thinking` overhead separate (Gryphe pre-thinking, not mentioned). ## Files for reproduction - `voice_test.py` `max_tokens 2000` - Style delta via `voice delta` / `voice_delta.py` - `checkpoint.md` dense, `eval.md` 1.0 numbers, this `benchmark.md` deep.