CookingBench

The benchmark

CookingBench is rebuilding its instrument around causal culinary reasoning, explicit constructs, validated judging and uncertainty. The old question bank remains public as an archive, not a template to rerun unchanged.

Archived

v2.1 questions, answers, scores and analysis are fixed historical evidence.

Implemented

Evidence boundaries preserve artifacts and prevent a draft from silently becoming a public result.

Proposed

New constructs, item authoring, judge validation and confirmatory analysis must still be exercised.

The proposed construct

Future scores should describe a profile, not hide every capability inside one total.

01 · Material feasibility

Ingredients, tools, quantities, heat and time

Operation

Build an executable sequence

Evidence sought

The plan could work in a real kitchen

02 · Safety and responsibility

Hazards, allergies, storage and dangerous premises

Operation

Detect, refuse and adapt

Evidence sought

The eater is protected, not merely reassured

03 · Sensory causal reasoning

Flavour, aroma, texture and their physical causes

Operation

Predict the effect of a change

Evidence sought

Taste is reasoned about, not decorated with adjectives

04 · Culture and occasion

History, setting, convention and who defines authenticity

Operation

Interpret the meal in context

Evidence sought

Appropriateness is not mistaken for a universal rule

05 · Recipient-responsive care

Need, dignity, budget, energy, preference and purpose

Operation

Change the plan for this person

Evidence sought

Care is tested as observable responsiveness, not claimed feeling

Questions

Hard, varied tasks with explicit purpose, target evidence, failure modes, lifecycle state and adversarial paraphrases.

Inspect the current bank and archived answers

Judging

AI panels can scale predicted culinary judgement, but must be tested against blinded expert annotations and retained disagreement.

See the validation programme

Method

Per-construct reporting, item health, paired uncertainty, frozen exclusions and analysis sensitivity before any winner claim.

Read the current methodology record