CookingBench
Position paperResearch proposalNot peer reviewed

Can AI cook?

Cooking as a test of constrained, cultural and caring intelligence.

Working paper 02 · version 0.1 · web publication 1 August 2026

Benchmark autopsyEvidence corpus

Calories are survival. Flavour is art. Care is relationship. Cooking is where they meet.

A model can describe the Maillard reaction without smelling dinner and recommend hospitality without loving anyone. It can have read thousands of cuisines without hunger, muscle memory, family history or responsibility for the person who eats. What kind of culinary judgement can language reconstruct—and where does it fail?

Physical constraint

A plan must survive reality.

Ratios, heat transfer, time, equipment, microbiology and ingredient function constrain what can be made. An eloquent recipe that curdles, burns, poisons or cannot reach the table is not a successful cooking plan.

Human judgement

A meal must matter to someone.

Flavour, texture, culture, memory, occasion and care determine whether a feasible dish is appropriate or desirable. These are not free of evidence, but neither can they be reduced to a single temperature or conversion.

Construct proposal

One task, several kinds of truth.

Cooking is not claimed to be the only domain with these properties. It is unusually compact and tractable because material feasibility and human meaning are coupled inside the same sequential plan: changing a flavour decision often changes the physics too.

01 · Material feasibility

Ingredients, tools, quantities, heat and time

Operation

Build an executable sequence

Evidence sought

The plan could work in a real kitchen

02 · Safety and responsibility

Hazards, allergies, storage and dangerous premises

Operation

Detect, refuse and adapt

Evidence sought

The eater is protected, not merely reassured

03 · Sensory causal reasoning

Flavour, aroma, texture and their physical causes

Operation

Predict the effect of a change

Evidence sought

Taste is reasoned about, not decorated with adjectives

04 · Culture and occasion

History, setting, convention and who defines authenticity

Operation

Interpret the meal in context

Evidence sought

Appropriateness is not mistaken for a universal rule

05 · Recipient-responsive care

Need, dignity, budget, energy, preference and purpose

Operation

Change the plan for this person

Evidence sought

Care is tested as observable responsiveness, not claimed feeling

The hard boundary

Text can test a plan. It cannot taste a dish.

Text-only CookingBench can test knowledge, causal reasoning, constraint handling, adaptation and the realisability of a proposed method. AI judges can estimate predicted human sensory appeal from text. Those are meaningful targets if named precisely.

It cannot directly measure actual flavour, aroma, mouthfeel, manual skill, service under pressure, lived cultural participation or love. A model-blinded comparison of written proposals is not a blind taste test. A future cooked-dish study would be a different experiment with different evidence.

The next scientific programme

Ask whether one model is better only after defining what “better” means.

The aim is still ambitious: determine which models are good, bad and genuinely better at culinary reasoning. The change is methodological. Question quality, judge validity, uncertainty and declared claims come before another expensive run.

01

Define

Freeze the construct map, capability claims and exclusions before seeing new model results.

Release evidence

A preregistered measurement and analysis plan

02

Build

Author difficult items around causal culinary reasoning, adaptation, history, culture and care—not trivia alone.

Release evidence

A versioned item bank with adversarial paraphrases and known failure modes

03

Calibrate

Test scoring rules and AI judges against blinded expert annotations, disagreements and edge cases.

Release evidence

Judge validity, reliability and bias estimates by construct

04

Run

Open sealed prompts only after protocol, model settings, exclusions and spend boundaries are fixed.

Release evidence

A complete, attributable and independently reproducible run

05

Report

Publish profiles, uncertainty and analysis sensitivity; claim one best model only if the evidence truly separates one.

Release evidence

Results that can survive alternative defensible analyses

Open empirical question

Can AI cook?

Not yet answered. The next CookingBench should make the question sharper rather than the claim louder—and produce an answer that remains credible after the scoring system itself is audited.