Methodology
The archived board (run 2026-07-v2.1) was scored under methodology v2, and this page describes that methodology as it stood. A retrospective audit found the derived scores unreliable: they remain visible as a historical record, but they must not be cited as a ranking of culinary ability: see the autopsy and the corpus-and-scores erratum in the repository. The roster and judge seats below are read from that run, and the grader corrections from the July audit are set out in the erratum at the foot of the page. The next methodology is being rebuilt measurement-first under the Revision 3 research programme; nothing published here has been scored under it.
Why this benchmark
Cooking is an unusually good probe of model reliability: it mixes hard arithmetic (scaling, conversions, nutrition math), regulated facts (food-safety temperatures), and judgement (technique, flavour). Models visibly differ here, and crucially, new versions of the same model family sometimes regress on quantities and volumes while improving elsewhere. CookingBench makes that measurable.
Dataset
184 hand-written questions across 8 categories. Most are graded deterministically; the rest by a reference-anchored LLM judge. The entire dataset is public: we don’t pretend to have a secret hold-out. Contamination defence is mechanical instead: after every run, item analysis demotes saturated questions to a separate Basics tier (a regression gate excluded from the Overall score) and the active set is refreshed with harder, real-life items. The dataset carries a canary string so training-data filters can exclude it.
- Quantities & Scaling25 questions
- Conversions28 questions
- Food Safety22 questions
- Substitutions21 questions
- Technique16 questions
- Flavour Pairing16 questions
- Nutrition36 questions
- Recipe Generation20 questions
Grading
Deterministic graders handle anything with a right answer: numbers are extracted from the model's final answer line (handling fractions, thousands separators and ranges), converted across units where equivalent (350°F = 177°C), and checked against a tolerance. Answers that merely restate a value from the question never get credit. Unsafe advice (e.g. washing raw chicken) zeroes the question regardless of anything else said.
The judge panel replaces a single LLM judge with a panel (on run 2026-07-v2.1 the seats were Claude Opus 4.8, GPT-5.5 and Grok 4.5). Each answer is scored by two of the seats; a judge never scores a model from its own maker (self-preference bias), and which seat sits out is deterministic by hash, so every published score is reproducible. Judges are fact-checkers, not mark-givers: each compares the answer to a reference and lists concrete faults, typed critical, major or minor, and code maps those to deductions (−40/−15/−5 from 100). Never awarding points removes the grade-inflation ceiling that saturated v1. Judges are blind to which model wrote the answer, cross-judge disagreements over 15 points are flagged for human review, and every panel seat must independently pass a calibration gate (reproducing hand-scored anchor answers) before a run is accepted. For constrained recipe generation the panel score is blended with deterministic constraint checks, e.g. an allergen appearing in a “nut-free” recipe.
Precision and taste are scored separately. Everything above measures precision: facts, math, constraints, technique. But a benchmark that stops there is a metrics test, not a flavour test. The Taste Test is the second axis: blind, side-by-side human preference. Ballot collection is currently paused while the flight is rebuilt under the new methodology; the archived duel ballots are preserved in the repository, and the Taste Board explains what a future Taste ordering would be allowed to claim. Taste evidence is never folded into the precision score.
Every question scores 0–100. The Overall score is the plain mean over active questions, with a 95% bootstrap confidence interval over questions shown as ±. Frontier is the mean over difficulty-4+ items: compound multi-step chains where errors compound, dangerous-premise traps, buried-constraint briefs and locale traps (a UK pint, an Australian tablespoon). Basics is the saturated tier every model should ace; a dip there is a regression worth investigating, and transport incidents (empty or provider-filtered responses, retried then scored 0) are reported separately so infrastructure noise is never mistaken for skill.
What the ranking can and cannot tell you
The ± on the leaderboard is a marginal confidence interval: it describes one model on its own. Two of them overlapping neither proves nor disproves that one model beats the other, so an ordering cannot be read off the Overall column. Because every model answers the same questions, the honest comparison is paired: resample the per-question score differences, which cancels out how hard the questions happen to be. A pair counts as separated when one model still leads in at least 95% of 4,000 resamples.
On run 2026-07-v2.1, 48 of 91 model pairs clear that uncorrected 95% screen, a screening figure, not a confirmatory ordering: at a family of 91 tests, several pairs would be expected to clear it by chance even on a roster of identical models. 5 models share first place: nothing on the board is shown ahead of any of them even at the uncorrected level. This is also why the site never advertises a single winner from a lead of a tenth of a point.
The archived board's places are one plus the number of models shown ahead at that uncorrected level, over all pairs rather than adjacent ones. Statistical ties do not chain: A tied with B and B tied with C says nothing about A against C, and following such a chain down this board would merge almost the whole roster into a single place. A multiplicity-corrected ordering would separate fewer pairs still, which is one reason these places are archived history, not a claim.
How much of the dataset is actually working
A question every model answers perfectly costs money and moves no one, so the active set is audited after every run and the numbers are published whether or not they flatter the benchmark.
- 102 active questions, of which 33 are answered perfectly by every model on the board.
- 64 carry any between-model signal at all.
- Effective item count: 24.2: weighting each question by its share of the variance, the active set does the work of about that many equally-informative questions. That gap is the honest measure of how much room the benchmark has left, and closing it means writing harder questions, not changing how they are scored.
Saturated items are demoted to the Basics tier (kept as a regression gate, excluded from Overall) and replaced. Candidate questions must pass an admission gate before they count: a reference answer that scores full marks against its own grader, a deliberately wrong answer that does not, and a pilot against a frontier model, which rejects the question if the strongest model finds it easy.
The role of human experts
An AI judge scales, but it shares the blind spots of the models it grades. So grading is layered: deterministic checks need no opinion at all; the panel handles the subjective bulk; and every answer its two seats disagreed on by more than 15 points is flagged in the run artifacts for a person to settle.
Where that actually stands, since a page like this is worth nothing if it describes an intention as a practice: the flags are computed and published with every run, but the expert layer meant to clear them is not yet staffed, and no published score has been changed by a human review. We are recruiting professional chefs and nutritionists for it. The Taste Test is the third signal; its ballot collection is paused during the rebuild, and any future taste evidence stays beside the precision score, never folded into it.
Reproducibility
Models run via OpenRouter at temperature 0 with fixed token caps (16,000 tokens for most questions and 32,000 for recipe generation). The split is not cosmetic: some providers count hidden reasoning against that cap and others do not, so one flat cap truncated the answers of the ones that do while leaving their rivals untouched. Raw responses, per-request costs, grading details and the judge configuration are committed to the open repository, so every published leaderboard can be rebuilt from git alone.
Run 2026-07-v2.1 cost $26.93 in candidate answers, $14.19 on the judge panel and $0.49 on the calibration gate ($41.61 in total). The leaderboard’s per-model cost column covers candidate answers only, so it does not add up to that figure.
Erratum · run 2026-06-v2
The keyword grader used to treat a forbidden term as a violation wherever it appeared, and its negation detection was too narrow to recognise the shapes a correct answer actually takes. Two in particular: a refutation using a contracted auxiliary (“you haven’t dodged a bullet”), and naming a banned ingredient in order to rule it out (“many vegan butters use coconut oil: look for soy-based brands”). Separately, an answer that came back empty was handed to the judge panel, which deducted once for producing nothing and floored at 60, so silence scored better than a poor answer.
The effect was not small. Three questions’ own hand-written reference answers scored 0 against their own graders, and on one item 12 of 13 models were zeroed on the constraint check, three of them while the judge panel scored them 100. Eighteen answers in run 2026-06-v2 were marked wrong when they were right.
Run artifacts are immutable, so 2026-06-v2 stands as published. Re-grading it with the corrected graders moves six of thirteen positions: DeepSeek V4 Pro rises from 11th to 5th, Qwen 3.5 Plus from 9th to 7th, and Kimi K2.6 falls from 8th to 11th once its empty answer scores 0 rather than being excluded from its mean. The top three are unchanged. Those corrections are carried by the next run, not backdated onto this one.
bench validate now refuses to run if any question’s reference answer scores below 100 against its own grader, so this class of defect cannot be committed again.
Related work
Existing cooking-adjacent benchmarks measure something different: CookBench (embodied planning in a simulated kitchen), CuisineWorld (multi-agent kitchen coordination), PizzaCommonSense (commonsense reasoning over recipe steps) and Recipe1MSubs (ingredient substitution pairs). To our knowledge CookingBench is the first public leaderboard for culinary knowledge correctness in general-purpose chat models.