CookingBench

Results / 2026-07-v2.1

CookingBench v2.1

Original published scores, preserved as historical evidence—not a current claim that one model was the best cook.

What this supports

The raw responses, original scores and paired-comparison record can be reproduced. Broad separation between some models exists in this run, and the frontier subset contains a stronger signal than the headline total.

What this does not support

A unique fine-grained ordering of the leading models; a corrected winner after post-hoc item removal; or a claim that the overall score is latent cooking ability.

Original published scores

Row order reproduces the historical artifact. It is labelled “Published order” because the order is not a statistically established total ranking. Values have not been silently regraded.

14 models · 184/184 prompts each
Original CookingBench v2.1 published scores
Published orderModelOriginal overall95% intervalFrontierInterpretation
1GPT-5.4 MiniOpenAI96.093.7–98.194.5Unresolved leading group
2GPT-5.6 Sol ProOpenAI96.093.2–98.493.8Unresolved leading group
3Grok 4.5xAI96.093.3–98.298.6Unresolved leading group
4GPT-5.6 Terra ProOpenAI95.491.7–98.391.3Unresolved leading group
5Claude Fable 5Anthropic94.190.3–97.387.1Unresolved leading group
6Kimi K3Moonshot AI93.990.4–96.986.9Historical score only
7Gemini 3.6 FlashGoogle93.389.5–96.491.1Historical score only
8Claude Opus 5Anthropic92.989.1–96.187.9Historical score only
9Gemini 3.1 Pro PreviewGoogle92.888.9–96.591.6Historical score only
10DeepSeek V4 ProDeepSeek92.788.8–96.088.2Historical score only
11Claude Sonnet 5Anthropic90.986.6–94.886.7Historical score only
12Qwen 3.7 MaxAlibaba90.786.2–94.685.8Historical score only
13Mistral Large 3Mistral85.980.6–90.687.0Historical score only
14Llama 4 MaverickMeta84.579.1–89.585.8Historical score only
Exploratory sensitivity

The top order is fragile.

Excluding all 12 active items with negative observed top–bottom discrimination—six keyword and six LLM-judge—reordered all four members of the unrounded raw-score top four without changing that group’s membership. Other defensible scoring choices produce other orders. This is evidence of sensitivity, not a corrected ranking.

See the analysis table

Errata and known limitations

Scoring semantics

Keyword and constraint checks sometimes penalised warnings, substitutions and cross-contamination advice for naming a forbidden ingredient.

Item-bank saturation

Thirty-three of 102 active items were perfect for every model; unrounded effective item count is approximately 24.0.

Judge evidence

Panel disagreements were not independently adjudicated, and current audit classifications are agent-produced and unblinded.

Archived release · 29 July 2026 · original scores unchanged · response digest cc5b5db1988921298ff72cf034ced530820d18a821009b15f56e9dfee9f40e8b