We Added Claude and ChatGPT to the "Show, Don't Tell" Detection Test. The Wall Held — But It Has Two Sides.

Community Article
Published July 25, 2026

Levent Bulut · Independent Researcher · ORCID 0009-0007-7500-2261

Conflict of interest, stated first: one of the two models evaluated below is Claude — and the assistant that helped assemble this analysis is also a Claude model. The original protocol for this study excluded Claude as a rater for exactly that circularity; this round deliberately includes it as a subject and reports the deviation instead of hiding it. Human labels were locked before any model labeling, prompts were fixed, and every statistic comes from a small deterministic script published with the data. Judge the numbers, not any model's authority — including this one's.


In the last post, I described a held-out test of whether machines can detect "show, don't tell" craft features: 100 fresh Turkish scenes from the objective-projection dataset, six feature definitions, one independent human rater as reference, and three machine raters — a rule-based detector, Gemini 2.5 Flash, and Grok. The result was uncomfortable in a useful way: surface features were trivial, and the two genuinely inferential features defeated every machine. I ended by asking whether any language model does better, or whether they all hit the same wall.

Two more model families have now taken the same test: Claude Fable 5 (web, "High" reasoning mode) and ChatGPT 5.5 (web, default), July 2026. Same 100 scenes, same masked IDs, same ten prompt blocks used for Gemini and Grok, scored against the same locked human labels.

The short version

Both models replicated the core failure. On materialized metaphor — an abstract inner state converted into a concrete, load-bearing physical detail, the diagnostic move of shown-mode writing — the human found 9 positives in 100 scenes. Claude scored κ = 0.03 against the human; ChatGPT κ = 0.02. That makes five machine raters in a row at κ ≈ 0 on this feature, across two studies.

But the shape of the failure is new, and it is the most informative thing this round produced. Here is every rater's positive count on the same definition:

Rater Flagged VAR (of 100) Of the human's 9, caught
Grok 0 0
Gemini 2.5 Flash ~2 1
ChatGPT 5.5 40 4
Claude Fable 5 78 8
Rule-based detector 72 7

One definition. Four frontier models. Positive rates from zero to nearly everything. Grok and Gemini under-flag toward silence; Claude flags almost everything — catching 8 of the human's 9 (the best recall of any rater) while burying them under 70 false positives, a profile strikingly close to the naive keyword detector it was supposed to outclass. ChatGPT sits in the middle.

In the last post I said the data supported two readings it could not separate: (a) the feature requires inference machines can't yet do at human level, or (b) the definition isn't operational enough for any rater to apply consistently. The 0 / 2 / 40 / 78 spread moves real weight onto (b) — or more precisely, it establishes (b) regardless of whether (a) is also true. A definition that four capable models apply at thresholds spanning two orders of magnitude is not yet a measurement instrument. The wall is real; part of the wall is ours.

The claim that didn't survive — and gets corrected

Last time I reported that the second inferential feature, atmosphere contradiction (an incongruous concrete detail cutting against the scene's emotional vector), was missed almost entirely by every machine: the human found 44, the detector caught 0, Gemini 2, Grok 3.

That sentence is now bounded to those raters, because it did not survive this round. Claude identified 31 of the human's 44 (κ = 0.27); ChatGPT identified 23 (κ = 0.18). Still weak in absolute terms — 24 and 19 false positives respectively, nowhere near a usable annotator — but categorically different from zero, on the one rule with a balanced class split. Either newer models are genuinely better at this class of inference, or the construct sits at a difficulty level where model generation matters. Either way: correction published, same prominence as the original claim.

What this says about Claude vs ChatGPT

Honestly: neither can annotate these features at human level, and "which is better" depends on which failure you prefer. ChatGPT tracked the human more closely overall (84.5% vs 81.0% raw agreement across 600 cells) and stayed near plausible base rates. Claude found more of what the human found on the hard features — best recall on both — at the cost of flagging far too much. As a high-recall screening pass, Claude's profile is arguably more useful; as a trustworthy label source, neither qualifies. On the four rules with extreme class skew, κ is uninformative by construction and I decline to rank models there.

A bonus finding: within-model stability is high. One block was re-run in each model; ChatGPT agreed with its own first run on 58/60 cells, Claude on 60/60. The inconsistency in this system is between raters, not within them.

What would actually settle it

The same thing as last time, now with more urgency: a second independent human rater. If two humans converge while five machine raters scatter from 0 to 78, reading (a) survives alongside (b). If the humans scatter too, the definitions go back to the workshop. That search is ongoing — if you work on annotation or literary evaluation and want to label 100 short scenes against six definitions, the dataset page has everything needed, and disagreement with the existing labels is the most valuable outcome.

Full write-up with methodology, tables, and limitations: Claude vs ChatGPT on Narrative Analysis. Both models' complete 100-scene label sets, the per-rule statistics, and the scoring script are published alongside, so every number here can be recomputed with no model in the loop.

Data: objective-projection — HF DOI 10.57967/hf/8960, Zenodo archive 10.5281/zenodo.19511369, CC BY-NC-ND 4.0.

Community

Sign up or log in to comment