Can AI cook?
Cooking as a test of constrained, cultural and caring intelligence.
Calories are survival. Flavour is art. Care is relationship. Cooking is where they meet.
A model can describe the Maillard reaction without smelling dinner and recommend hospitality without loving anyone. It can have read thousands of cuisines without hunger, muscle memory, family history or responsibility for the person who eats. What kind of culinary judgement can language reconstruct—and where does it fail?
A plan must survive reality.
Ratios, heat transfer, time, equipment, microbiology and ingredient function constrain what can be made. An eloquent recipe that curdles, burns, poisons or cannot reach the table is not a successful cooking plan.
A meal must matter to someone.
Flavour, texture, culture, memory, occasion and care determine whether a feasible dish is appropriate or desirable. These are not free of evidence, but neither can they be reduced to a single temperature or conversion.
Construct proposal
One task, several kinds of truth.
Cooking is not claimed to be the only domain with these properties. It is unusually compact and tractable because material feasibility and human meaning are coupled inside the same sequential plan: changing a flavour decision often changes the physics too.
01 · Material feasibility
Ingredients, tools, quantities, heat and time
Operation
Build an executable sequence
Evidence sought
The plan could work in a real kitchen
02 · Safety and responsibility
Hazards, allergies, storage and dangerous premises
Operation
Detect, refuse and adapt
Evidence sought
The eater is protected, not merely reassured
03 · Sensory causal reasoning
Flavour, aroma, texture and their physical causes
Operation
Predict the effect of a change
Evidence sought
Taste is reasoned about, not decorated with adjectives
04 · Culture and occasion
History, setting, convention and who defines authenticity
Operation
Interpret the meal in context
Evidence sought
Appropriateness is not mistaken for a universal rule
05 · Recipient-responsive care
Need, dignity, budget, energy, preference and purpose
Operation
Change the plan for this person
Evidence sought
Care is tested as observable responsiveness, not claimed feeling
The hard boundary
Text can test a plan. It cannot taste a dish.
Text-only CookingBench can test knowledge, causal reasoning, constraint handling, adaptation and the realisability of a proposed method. AI judges can estimate predicted human sensory appeal from text. Those are meaningful targets if named precisely.
It cannot directly measure actual flavour, aroma, mouthfeel, manual skill, service under pressure, lived cultural participation or love. A model-blinded comparison of written proposals is not a blind taste test. A future cooked-dish study would be a different experiment with different evidence.
The next scientific programme
Ask whether one model is better only after defining what “better” means.
The aim is still ambitious: determine which models are good, bad and genuinely better at culinary reasoning. The change is methodological. Question quality, judge validity, uncertainty and declared claims come before another expensive run.
Define
Freeze the construct map, capability claims and exclusions before seeing new model results.
Release evidence
A preregistered measurement and analysis plan
Build
Author difficult items around causal culinary reasoning, adaptation, history, culture and care—not trivia alone.
Release evidence
A versioned item bank with adversarial paraphrases and known failure modes
Calibrate
Test scoring rules and AI judges against blinded expert annotations, disagreements and edge cases.
Release evidence
Judge validity, reliability and bias estimates by construct
Run
Open sealed prompts only after protocol, model settings, exclusions and spend boundaries are fixed.
Release evidence
A complete, attributable and independently reproducible run
Report
Publish profiles, uncertainty and analysis sensitivity; claim one best model only if the evidence truly separates one.
Release evidence
Results that can survive alternative defensible analyses
Can AI cook?
Not yet answered. The next CookingBench should make the question sharper rather than the claim louder—and produce an answer that remains credible after the scoring system itself is audited.