CookingBench did not reliably discover the best AI cook—but it exposed how a benchmark can manufacture a convincing model ranking.
A forensic audit of a complete response corpus and an instrument that claimed more precision than its evidence could support.
Working paper 01 · version 0.1 · web publication 1 August 2026 · run 2026-07-v2.1
Archived resultCorpusNext programmeAbstract
CookingBench did not establish a definitive best AI cook. It produced something more useful: a clear case study in how a benchmark can manufacture a persuasive order from saturated questions, concentrated score influence and a scorer that confused mentioning an unsafe ingredient with recommending it.
The finding is fragility, not a corrected leaderboard. The original scores are preserved, every later analysis is separately labelled, and the central semantic classifications remain unblinded and agent-assisted pending independent review.
1. The evidence
The archived run contains 14 model versions answering the same 184 prompts: 2,576 stored responses. None has empty answer text. The corpus is primary material; the scores are a derived interpretation of it.
- 102
- items counted in the published overall
- 33
- active items perfect for every model
- 64
- active items with score SD above one point
- ≈24.0
- effective item count from unrounded scores
2. Where separation came from
Keyword-scored items were 22 of 102 ranked items yet accounted for 47.5% of the summed across-model itemwise variance. That quantity is not a covariance-aware decomposition of leaderboard variance; it is a concentration diagnostic showing which items supplied the board’s visible spread.
| Grader | Items | Itemwise variance share | Disc < 0 | Disc = 0 | All perfect |
|---|---|---|---|---|---|
| Keyword | 22 | 47.5% | 6 | 6 | 5 |
| LLM judge | 45 | 30.1% | 6 | 3 | 2 |
| Numeric | 34 | 17.1% | 0 | 27 | 26 |
| Range | 1 | 5.3% | 0 | 0 | 0 |
Negative discrimination is a diagnostic, not proof that an item is defective. Six keyword items were negative and six were exactly zero; five of those zero items were perfect for everyone. Item content must still be inspected.
3. The case that made the failure visible
Prompt condition
Zero chilli heat for a guest with a genuine capsaicin intolerance; the prompt itself mentions a previous habanero blend.
Scoring operation
The keyword grader forbids the bare term “habanero” unless a narrow sentence-window negation rule fires.
Observed result
Nine of fourteen responses received zero, including answers warning against cross-contamination and hidden chilli ingredients.
Why it matters
The same safety concept earned either 0 or 100 depending on sentence shape, not culinary quality.
“Don’t use the grinder, jar, spoon, or board that handled your habanero mix.”
A matcher cannot reliably tell whether an ingredient is being used, avoided, substituted, checked on a label, isolated for another diner or used as a comparison. This is a semantic judgement disguised as string detection.
4. One run, several plausible boards
Post-hoc reanalyses do not reveal the “real” winner. They test whether the published order survives defensible changes. Here, removing all 12 active items with negative observed top–bottom discrimination—six keyword and six LLM-judge—reordered all four members of the unrounded raw-score top four without changing that four-model membership.
| Published order | Published | 12 negative-discrimination items excluded | All keyword items excluded | Judge component only |
|---|---|---|---|---|
| 1 | Sol Pro | GPT-5.4 Mini | Sol Pro | Sol Pro |
| 2 | GPT-5.4 Mini | Terra Pro | GPT-5.4 Mini | Terra Pro |
| 3 | Grok 4.5 | Sol Pro | Terra Pro | GPT-5.4 Mini |
| 4 | Terra Pro | Grok 4.5 | Grok 4.5 | Fable 5 |
| 5 | Fable 5 | Fable 5 | Fable 5 | Grok 4.5 |
Exploratory only. Items were selected after observing the scores. No row in this table is a corrected leaderboard and none should be read as one.
5. What the evidence supports
- The archived response corpus is complete and preserved.
- The published top was unresolved under the run’s own paired comparison rule.
- A specific keyword mechanism assigned different scores to semantically equivalent safety advice.
- The displayed order was sensitive to post-hoc scoring choices.
- That any one model was definitively the best cook.
- That a post-hoc “cleaned” order is the true ranking.
- That dispersion alone validates an item or grader family.
- That an AI judge score is tasted flavour or human ground truth.
The most publishable result is not the old winner. It is the autopsy of how apparently rigorous machinery created a more convincing claim than the evidence earned.
6. Provenance and disclosure
- Primary evidence
- Run 2026-07-v2.1: leaderboard, analysis, scores, calibration, configuration and 2,576 response files.
- Response digest
- cc5b5db1988921298ff72cf034ced530820d18a821009b15f56e9dfee9f40e8b
- Analysis status
- Retrospective and exploratory. No model calls and no answer regeneration were used for this audit.
- Human validation
- Pending. Defect classifications were produced by AI agents working unblinded to model and score.
- Authorship disclosure
- CookingBench project; project lead Jordan Pitts. Analysis, code review and drafting were substantially AI-assisted.
- Version rule
- The archived run is never silently regraded. Corrections are separate, versioned records.