CookingBench

Research

CookingBench studies culinary intelligence and the instruments used to measure it. Findings, failures, open questions, data and version history are published together.

Position paper & research proposal02

Can AI cook?

A construct framework for evaluating materially constrained, sensory, cultural and recipient-responsive culinary intelligence.

Working paper · proposed programme · not peer reviewed

Read publication

Evidence status

What exists today

Preserved

A complete, immutable primary-response corpus for v2.1, pinned by file and content digest.

Exploratory

The scoring autopsy and answer classifications were performed after seeing the data; the classifications are agent-assisted, unblinded and not independently adjudicated.

Proposed

The next benchmark construct, question bank, judging validation and analysis protocol. These are a research programme, not completed evidence.

The open evidence base

The scores are not a valid answer to “which model is the best cook?” The responses are still valuable evidence for independent reanalysis and better judging experiments.

14
model versions
184
culinary prompts each
2,576
preserved responses
0
empty answer texts