Results / 2026-07-v2.1
CookingBench v2.1
Original published scores, preserved as historical evidence—not a current claim that one model was the best cook.
The raw responses, original scores and paired-comparison record can be reproduced. Broad separation between some models exists in this run, and the frontier subset contains a stronger signal than the headline total.
A unique fine-grained ordering of the leading models; a corrected winner after post-hoc item removal; or a claim that the overall score is latent cooking ability.
Original published scores
Row order reproduces the historical artifact. It is labelled “Published order” because the order is not a statistically established total ranking. Values have not been silently regraded.
| Published order | Model | Original overall | 95% interval | Frontier | Interpretation |
|---|---|---|---|---|---|
| 1 | GPT-5.4 MiniOpenAI | 96.0 | 93.7–98.1 | 94.5 | Unresolved leading group |
| 2 | GPT-5.6 Sol ProOpenAI | 96.0 | 93.2–98.4 | 93.8 | Unresolved leading group |
| 3 | Grok 4.5xAI | 96.0 | 93.3–98.2 | 98.6 | Unresolved leading group |
| 4 | GPT-5.6 Terra ProOpenAI | 95.4 | 91.7–98.3 | 91.3 | Unresolved leading group |
| 5 | Claude Fable 5Anthropic | 94.1 | 90.3–97.3 | 87.1 | Unresolved leading group |
| 6 | Kimi K3Moonshot AI | 93.9 | 90.4–96.9 | 86.9 | Historical score only |
| 7 | Gemini 3.6 FlashGoogle | 93.3 | 89.5–96.4 | 91.1 | Historical score only |
| 8 | Claude Opus 5Anthropic | 92.9 | 89.1–96.1 | 87.9 | Historical score only |
| 9 | Gemini 3.1 Pro PreviewGoogle | 92.8 | 88.9–96.5 | 91.6 | Historical score only |
| 10 | DeepSeek V4 ProDeepSeek | 92.7 | 88.8–96.0 | 88.2 | Historical score only |
| 11 | Claude Sonnet 5Anthropic | 90.9 | 86.6–94.8 | 86.7 | Historical score only |
| 12 | Qwen 3.7 MaxAlibaba | 90.7 | 86.2–94.6 | 85.8 | Historical score only |
| 13 | Mistral Large 3Mistral | 85.9 | 80.6–90.6 | 87.0 | Historical score only |
| 14 | Llama 4 MaverickMeta | 84.5 | 79.1–89.5 | 85.8 | Historical score only |
The top order is fragile.
Excluding all 12 active items with negative observed top–bottom discrimination—six keyword and six LLM-judge—reordered all four members of the unrounded raw-score top four without changing that group’s membership. Other defensible scoring choices produce other orders. This is evidence of sensitivity, not a corrected ranking.
Errata and known limitations
Scoring semantics
Keyword and constraint checks sometimes penalised warnings, substitutions and cross-contamination advice for naming a forbidden ingredient.
Item-bank saturation
Thirty-three of 102 active items were perfect for every model; unrounded effective item count is approximately 24.0.
Judge evidence
Panel disagreements were not independently adjudicated, and current audit classifications are agent-produced and unblinded.
Archived release · 29 July 2026 · original scores unchanged · response digest cc5b5db1988921298ff72cf034ced530820d18a821009b15f56e9dfee9f40e8b