The benchmark
CookingBench is rebuilding its instrument around causal culinary reasoning, explicit constructs, validated judging and uncertainty. The old question bank remains public as an archive, not a template to rerun unchanged.
v2.1 questions, answers, scores and analysis are fixed historical evidence.
Evidence boundaries preserve artifacts and prevent a draft from silently becoming a public result.
New constructs, item authoring, judge validation and confirmatory analysis must still be exercised.
The proposed construct
Future scores should describe a profile, not hide every capability inside one total.
01 · Material feasibility
Ingredients, tools, quantities, heat and time
Operation
Build an executable sequence
Evidence sought
The plan could work in a real kitchen
02 · Safety and responsibility
Hazards, allergies, storage and dangerous premises
Operation
Detect, refuse and adapt
Evidence sought
The eater is protected, not merely reassured
03 · Sensory causal reasoning
Flavour, aroma, texture and their physical causes
Operation
Predict the effect of a change
Evidence sought
Taste is reasoned about, not decorated with adjectives
04 · Culture and occasion
History, setting, convention and who defines authenticity
Operation
Interpret the meal in context
Evidence sought
Appropriateness is not mistaken for a universal rule
05 · Recipient-responsive care
Need, dignity, budget, energy, preference and purpose
Operation
Change the plan for this person
Evidence sought
Care is tested as observable responsiveness, not claimed feeling
Questions
Hard, varied tasks with explicit purpose, target evidence, failure modes, lifecycle state and adversarial paraphrases.
Judging
AI panels can scale predicted culinary judgement, but must be tested against blinded expert annotations and retained disagreement.
Method
Per-construct reporting, item health, paired uncertainty, frozen exclusions and analysis sensitivity before any winner claim.