Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
dipankarsarkar 
posted an update 10 days ago
Post
3172
I audited one of my own evaluations. The ranking did not hold up the way I expected.

Eight open models, one task: infer the structure of a prompt. Then ask again with the identical call. Caching off.

- Agreement between repeated identical calls (mean Jaccard) ranged from 0.39 to 0.96 across models.
- Only 35 of 127 prompt-model cells were perfectly reproducible on every run.
- I bootstrapped the reproducibility ranking over prompts. The two least reproducible models kept their rank in 99% and 86% of resamples. The middle four kept theirs in 27% to 48%.

So the table reliably finds the worst model. It does not reliably find the best.

Reproducible is also not the same as correct. F1 against gold annotations ran from 0.56 to 0.99.

By the audit date, 4 of the 8 model variants had been retired (HTTP 410). The study as specified can no longer be re-run. The saved outputs are what survives, so I published all of them: every inferred structure, every run, the prompts and the annotations.

Paper: How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure (2609.30074)
Dataset: dipankarsarkar/llm-evaluation-self-audit
Code: https://github.com/sarkar-dipankar/llm-evaluation-self-audit

How many of the leaderboard rankings you rely on would survive re-running the same calls?

This is very interesting! I saw a lot of variance in task outcomes for frontier models in our experiments on computer use and robotics tasks too. This really makes you wonder where eval should go beyond the benchmarks today.

btw, we just put out the full per-task breakdown, in case it's useful. Here's our report on it too. Would love your thoughts

I ran our bootstrap on your VisualWebArena table. It tells the same story.

173 tasks (the 4 with an invalid_reason dropped), corrected verdicts, resampled within site.

  • Kimi K3 keeps last place in 99% of resamples.
  • Astra keeps #1 in 82%. Fable 5 takes it in the other 18%.
  • The middle four (GPT-5.6, Opus 5.5, Opus 5, Opus 4.7) keep their rank in 41% to 60%.

Some of the site-level jaggedness is thinner than it looks. Fable 5 tops shopping and classifieds, but stays the winner there in only 65% and 60% of resamples. Astra on reddit holds at 97%.

Astra vs Fable 5 comes down to 18 of 173 tasks: 11 only Astra solved, 7 only Fable 5 solved.

Your limitations section names the missing piece: one pass per item. In our audit, on a different task, the same model agreed with itself on an identical call at a mean Jaccard of 0.39 to 0.96.

If you reran Astra on those 18 tasks, how many would flip against itself?