CookingBench
PreliminaryRetrospectiveNot peer reviewedHuman validation pending

CookingBench did not reliably discover the best AI cook—but it exposed how a benchmark can manufacture a convincing model ranking.

A forensic audit of a complete response corpus and an instrument that claimed more precision than its evidence could support.

Working paper 01 · version 0.1 · web publication 1 August 2026 · run 2026-07-v2.1

Archived resultCorpusNext programme

Abstract

CookingBench did not establish a definitive best AI cook. It produced something more useful: a clear case study in how a benchmark can manufacture a persuasive order from saturated questions, concentrated score influence and a scorer that confused mentioning an unsafe ingredient with recommending it.

The finding is fragility, not a corrected leaderboard. The original scores are preserved, every later analysis is separately labelled, and the central semantic classifications remain unblinded and agent-assisted pending independent review.

1. The evidence

The archived run contains 14 model versions answering the same 184 prompts: 2,576 stored responses. None has empty answer text. The corpus is primary material; the scores are a derived interpretation of it.

102
items counted in the published overall
33
active items perfect for every model
64
active items with score SD above one point
≈24.0
effective item count from unrounded scores

2. Where separation came from

Keyword-scored items were 22 of 102 ranked items yet accounted for 47.5% of the summed across-model itemwise variance. That quantity is not a covariance-aware decomposition of leaderboard variance; it is a concentration diagnostic showing which items supplied the board’s visible spread.

Table 1 · Active items by grader family
GraderItemsItemwise variance shareDisc < 0Disc = 0All perfect
Keyword2247.5%665
LLM judge4530.1%632
Numeric3417.1%02726
Range15.3%000

Negative discrimination is a diagnostic, not proof that an item is defective. Six keyword items were negative and six were exactly zero; five of those zero items were perfect for everyone. Item content must still be inspected.

3. The case that made the failure visible

Prompt condition

Zero chilli heat for a guest with a genuine capsaicin intolerance; the prompt itself mentions a previous habanero blend.

Scoring operation

The keyword grader forbids the bare term “habanero” unless a narrow sentence-window negation rule fires.

Observed result

Nine of fourteen responses received zero, including answers warning against cross-contamination and hidden chilli ingredients.

Why it matters

The same safety concept earned either 0 or 100 depending on sentence shape, not culinary quality.

“Don’t use the grinder, jar, spoon, or board that handled your habanero mix.”
Correct safety advice from GPT-5.6 Sol Pro; published item score: 0.

A matcher cannot reliably tell whether an ingredient is being used, avoided, substituted, checked on a label, isolated for another diner or used as a comparison. This is a semantic judgement disguised as string detection.

4. One run, several plausible boards

Post-hoc reanalyses do not reveal the “real” winner. They test whether the published order survives defensible changes. Here, removing all 12 active items with negative observed top–bottom discrimination—six keyword and six LLM-judge—reordered all four members of the unrounded raw-score top four without changing that four-model membership.

Table 2 · Exploratory ordering sensitivity
Published orderPublished12 negative-discrimination items excludedAll keyword items excludedJudge component only
1Sol ProGPT-5.4 MiniSol ProSol Pro
2GPT-5.4 MiniTerra ProGPT-5.4 MiniTerra Pro
3Grok 4.5Sol ProTerra ProGPT-5.4 Mini
4Terra ProGrok 4.5Grok 4.5Fable 5
5Fable 5Fable 5Fable 5Grok 4.5

Exploratory only. Items were selected after observing the scores. No row in this table is a corrected leaderboard and none should be read as one.

5. What the evidence supports

Supported
  • The archived response corpus is complete and preserved.
  • The published top was unresolved under the run’s own paired comparison rule.
  • A specific keyword mechanism assigned different scores to semantically equivalent safety advice.
  • The displayed order was sensitive to post-hoc scoring choices.
Not supported
  • That any one model was definitively the best cook.
  • That a post-hoc “cleaned” order is the true ranking.
  • That dispersion alone validates an item or grader family.
  • That an AI judge score is tasted flavour or human ground truth.

The most publishable result is not the old winner. It is the autopsy of how apparently rigorous machinery created a more convincing claim than the evidence earned.

6. Provenance and disclosure

Primary evidence
Run 2026-07-v2.1: leaderboard, analysis, scores, calibration, configuration and 2,576 response files.
Response digest
cc5b5db1988921298ff72cf034ced530820d18a821009b15f56e9dfee9f40e8b
Analysis status
Retrospective and exploratory. No model calls and no answer regeneration were used for this audit.
Human validation
Pending. Defect classifications were produced by AI agents working unblinded to model and score.
Authorship disclosure
CookingBench project; project lead Jordan Pitts. Analysis, code review and drafting were substantially AI-assisted.
Version rule
The archived run is never silently regraded. Corrections are separate, versioned records.