CookingBench

Can AI cook?

Cooking joins physical truth with human judgement. A good answer must be safe, feasible and technically sound, but it must also understand flavour, culture, occasion and the person being fed.

The test, in one sequence

  1. 01

    Can it survive physics?

    Heat · time · ratios · safety

  2. 02

    Can it predict experience?

    Flavour · texture · aroma

  3. 03

    Can it respond to a person?

    Culture · occasion · care

What the benchmark exposed

A leaderboard can be reproducible and still measure the wrong thing.

Some questions separated no models. Others spread scores for the wrong reason. One flavour item with a safety constraint gave zero to correct warnings because the answers named the ingredient they were telling the user to avoid. The autopsy shows how saturation, semantic scoring failures and concentrated influence can create unjustified rank precision.

Read the forensic audit

Why cooking matters

Cooking is where physics becomes personal.

Heat, time, ratios and microbiology constrain what can work. Flavour, culture, memory and care shape whether the result is worth eating. Because both live inside the same task, cooking offers an unusually rich way to study what AI understands, and what it only sounds as if it understands.

Read “Can AI cook?”

The next instrument

What would culinary intelligence require?

There may be no single cooking ability. The next CookingBench will test a profile of connected capabilities and report uncertainty rather than forcing every difference into one rank.

01 · Material feasibility

Ingredients, tools, quantities, heat and time

Operation

Build an executable sequence

Evidence sought

The plan could work in a real kitchen

02 · Safety and responsibility

Hazards, allergies, storage and dangerous premises

Operation

Detect, refuse and adapt

Evidence sought

The eater is protected, not merely reassured

03 · Sensory causal reasoning

Flavour, aroma, texture and their physical causes

Operation

Predict the effect of a change

Evidence sought

Taste is reasoned about, not decorated with adjectives

04 · Culture and occasion

History, setting, convention and who defines authenticity

Operation

Interpret the meal in context

Evidence sought

Appropriateness is not mistaken for a universal rule

05 · Recipient-responsive care

Need, dignity, budget, energy, preference and purpose

Operation

Change the plan for this person

Evidence sought

Care is tested as observable responsiveness, not claimed feeling

Preserved primary material

The answers remain valuable even when the scores do not.

v2.1 contains every planned model–prompt response and no empty answer text. That does not validate the original ranking. It does preserve the primary material needed for blinded adjudication, alternative scoring and independent analysis.

14
model versions
184
prompts per model
2,576
response artifacts
1,998,922
answer-text Unicode code points
Scientific programme

The next run starts with better questions, better judging and a frozen claim.

CookingBench will treat question design as the instrument, validate AI judges against blinded expert judgements, measure dimensions separately and publish sensitivity, not declare a winner simply because a table can be sorted.

“Calories are survival. Flavour is art. Care is relationship. Cooking is where they meet.”

Historical result · July 2026

CookingBench v2.1

Original published scores, uncertainty, the unresolved leading group and every known limitation, preserved without silently rewriting the record.

View archived result

Archive record 2026-07-v2.1 · response digest cc5b5db198892129