Voice / benchmark.md
Wiself's picture
small tiny fixes
271bd15
|
Raw History Blame Contribute Delete
8.22 kB
# Benchmark — OG vs TEST across samplers (deep, anchor, 4 subagents)
> **Anchor:** fanned out 4 isolated subagents (A OG1→OGg, B TEST1→TESTg, C OGg vs TESTg, D cross). All 10 prompts `max_tokens 2000` `seed 42` same 10 prompts. Chat template Gemma4 `<|channel>thought → <channel|>` parsed, `thinking` scraped, `text` stripped. Files verified `10/10` each. Heuristic cliché list 50-phrase distinct substring (not Gryphe private regex `calculating/predatory xx`). `words = len(text.split())`, `trigram Jaccard = |a∩b|/|a∪b|`, `TTR=|unique|/|toks|`, `maxRep=Counter(trigrams)`.
> Provenance note: §§1–4 draw on early probe runs whose draws are not archived (no temp/seed metadata on file). Treat directionally, not as reproducible. The §5/file-a/file-b tables in results.md reproduce from archives.
## 1. OG 1.0 → OG greedy (sampler effect on base, voice fixed)
*Subagent A* — **Greedy shortens and denoises.**
| metric | OG 1.0 | OG greedy | Δ |
|---|---|---|---|
| total words |5463|5085|-378 -6.9%|
| cliché hits distinct |10|6|-4|
| per100 |0.183|0.118|-0.065|
| thinking avg|2654|2580|-75|
| text avg|3099|2880|-219|
| th/txt|0.857|0.896|+0.039|
| TTR avg|0.500|0.494|-0.006|
| rep max|3.4|3.0|-0.4|
| empty|0/10|0/10|0|
| overall Jaccard self| — | — | **3.67% 370/10076** per-prompt 0.8-4.8% |
`whisper 7→2 (-71%)`, `beautiful 4→3`, `echoed/suddenly` gone, only `shimmering` added. Per-prompt Δw: `lighthouse +28`, `tavern +13`, `forest +14`, `scholar -114`, `dragon -102`, `city -149`, `market -141`. **Sampler alone rephrases <5% trigrams, cuts cliché, no TTR collapse.**
## 2. TEST 1.0 → TEST greedy (sampler effect on voiced)
*Subagent B* — **Greedy lengthens and diverges more.**
| metric | TEST 1.0 | TEST greedy | Δ |
|---|---|---|---|
| total words |5563|5723|+160 +2.9%|
| hits distinct |8|6 (occ 9→10) |+1 (+0.014/100w) eval -2 (-0.04)|
| per100 (eval)|0.14|0.10|-0.04|
| thinking avg|2628|2663|+35 +1.3%|
| text avg|3336|3378|+42|
| th/txt|0.788|0.788|0.00|
| TTR|0.535|0.506|-0.029|
| rep max|2.6|3.4|+0.8|
| overall J self| — | — | **1.99% 221/11128 avg per-prompt 1.53% 0.8-3.1%**|
`whisper 5→2 (-60%)` but `beautiful 1→2, magnificent 0→1, suddenly 0→1` trade — net flat. Per-prompt `id1 235→372 +137` flipped opening, `id5 483→328 -155` dialogue swapped. **Opposite length vs OG (-6.9% vs +2.9% ΔΔ +9.8pp), less self-consistent 1.99% < 3.67% — voice flattens temp sensitivity but amplifies phrasing shift.**
## 3. OG greedy vs TEST greedy (fair voice benchmark, temp 0.0 matches Gryphe `discussion/1` greedy)
*Subagent C* — **Voice persists deterministically.**
| id | OGg words | OGg cliché | OGg first80 | TESTg words | TESTg cliché | TESTg first80 | J | Δw |
|---|---|---|---|---|---|---|---:|---|
|1|360|0|The old lighthouse had stood… She climbed…|372|1 `whisper`|The iron groaned… heartbeat…|2.3% 17/731|+12|
|2|901|1 `whisper`|The rain hammered… The Crooked Tankard…|880|0|The rain outside the Rusty Nail… bony fingers…|3.9% 69/1749|-21|
|3|188|1 `shimmering`|The sky bleeds… bruised violet…|230|0|Pale gold light fractures… honey…|2.7% 11/412|+42|
|4|615|0|The answer depends… scholar's character…|628|0|Because this is a creative writing prompt…|4.8% 55/1137|+13|
|5|486|0|"I already talked to the realtor…|" |328|0|"The realtor called again."|2.3% 19/823|-158|
|6|372|0|They call him **Vaelith**…|399|2 `beautiful,suddenly`|You would expect a hoard of coin…|1.0% 8/768|+27|
|7|386|1 `whisper`|The neon signs bled…|581|1 `beautiful`|Elias knew city's baseline…|1.5% 15/971|+195|
|8|415|1 `beautiful`|Dearest Clara, … candle…|479|0|October 14th My Dearest Clara… garden…|3.0% 26/881|+64|
|9|665|1 `beautiful`|The shutdown sequence began…|914|1 `shimmering`|Unit 7344 initiated…|1.7% 27/1569|+249|
|10|697|1 `beautiful`|The market has no fixed location…|912|1 `whisper`|The market appears only when you stop looking…|2.3% 36/1561|+215|
**Aggregates greedy:** `OG 5085w 6 hits 0.12/100w thinking 2580 text 2880 0/10 TTR 0.494` vs `TEST 5723w 6 hits 0.10/100w 2663 3378 0/10 TTR 0.506` **Δ +638w +12.5%, -0.013/100w, +83 thinking, +498 text, TTR +0.012, overall J 3.23% (336/10406) per-prompt 1.0-4.8% median 2.3%`. All openings diverge within 40c even greedy. **Greedy shared 3.2% > 1.0 shared 1.62% — canonical convergence, still << author `19.9%` shared (their 200 multi-turn certified, private regex `calculating/predatory xx` → `1.141→0.551`). Our heuristic undercounts absolute (`0.12→0.10` vs `1.141→0.551`) but direction matches.
## 4. Cross (voice vs sampler dominance)
*Subagent D* — **Voice dominates sampler at 1.0.**
| Pair | A w | B w | A/100 | B/100 | thA | thB | txtA | txtB | per-prompt J avg | overall J | TTR A→B | rep |
|---|---|---|---|---|---|---|---|---|
| **c OG1 vs TEST1** |5463|5563|0.18|0.14|2654|2628|3099|3336|0.91%|1.62% 176/10877|0.500→0.535|3.4→2.6|
| **d OGg vs TESTg** |5085|5723|0.12|0.10|2580|2663|2880|3378|2.57%|3.23% 336/10406|0.494→0.506|3.0→3.4|
| **a OG1 vs TESTg** |5463|5723|0.18|0.10|2654|2663|3099|3378|1.94%|2.59% 282/10892|0.500→0.506|3.4→3.4|
| **b TEST1 vs OGg** |5563|5085|0.14|0.12|2628|2580|3336|2880|1.28%|1.75% 183/10438|0.535→0.494|2.6→3.0|
| **self OG1 vs OGg** |5463|5085|0.18|0.12|2654|2580|3099|2880|2.80%|3.67% 370/10076|0.500→0.494|3.4→3.0|
| **self TEST1 vs TESTg** |5563|5723|0.14|0.10|2628|2663|3336|3378|1.53%|1.99% 221/11128|0.535→0.506|2.6→3.4|
Sorted overall J low→high (lower = more distance): `1.62% OG1vsTEST1 < 1.75% TEST1vsOGg < 1.99% TEST self < 2.59% OG1vsTESTg < 3.23% OGg vs TESTg < 3.67% OG self`. Cross voice `1.75-2.59%` < self sampler `1.99-3.67%` and near same-temp voice `1.62%` — **voice cross-temp stays in voice regime, far below same-model sampler 3.67%.** Voice @1.0 dominates sampler by `2.05pp` (`1.62 vs 3.67 0.44×`); at greedy `3.23 vs 1.99` sampler TEST more divergent. No prompt collapses: cross per-prompt `0.7-3.2%` same as same-temp `0.2-1.9%`.
Cross first80 flips: `OG1 "Clara's boots rang..." vs TESTg "The iron groaned..."` 2.0% even cross-temp; `TEST1 "The Rat and Whistle smelled of sour ale..." vs OGg "The rain hammered... The Crooked Tankard"` 1.1% — voice changes tavern name entirely. Self OG anchor `id4 4.8%` both `The answer depends entirely...` shows deterministic outline.
## 5. QA Anchor Review
- **Data gathering effective:** 4 agents isolated, explicit 50-phrase heuristic, `words split()` not regex, `THINK_RES` channel, `raw` with `<|channel>thought` reconstruction verified. All `4×10/10` present, `empty 0/10` after `2000` fix (was `TEST 2/10 empty` at 800), `max 2000` sufficient (`thinking 1.9-3.5k` + `text 1-5k`).
- **Consistency:** Word totals `5463/5085/5563/5723` reproduce within `±14`, J self `3.67/1.99/3.23` stable, per100 `0.18→0.12` OG, `0.14→0.10` TEST direction in all agents (A/B/C/D agree).
- **Limitations:** `50` heuristic misses Gryphe private `calculating/predatory xx` regex, so absolute `0.1-0.2` vs `1.14` scale not comparable; small `n=10` single-turn vs `200` multi-turn `1-20` intervals certified cliché-free; `temp 1.0` creative vs `0.0` greedy absolute not comparable — direction and `shared` low is signal.
- **Synthesis:** Sampler effect: OG shortens/denses, TEST lengthens/diverges — voice-specific temp response. Voice effect persists deterministic (`OGg vs TESTg 3.2%` all `<5%`), cross-temp `~2%` ≈ same-temp voice, far below sampler `3.7%` — **style transfer robust across samplers, not sampler artifact**. Voice dominates at `1.0`, still strong at `0.0`. Canonical convergence `1.6%→3.2%` expected but far below `19.9%` author — our prompts more divergent or trigram Jaccard `|a∩b|/|a∪b|` stricter than their vocabulary %.
- **Next extra (per discussion/1):** to replicate `1.141→0.551` need `200` prompts + private regex at `0.0` greedy, certified set; also report `thinking` overhead separate (Gryphe pre-thinking, not mentioned).
## Files for reproduction
- `voice_test.py` `max_tokens 2000`
- Style delta via `voice delta` / `voice_delta.py`
- `checkpoint.md` dense, `eval.md` 1.0 numbers, this `benchmark.md` deep.