CookingBench

Results / v2.1 / OpenAI

Archived v2.1Unresolved leading group

GPT-5.4 Mini

This profile reports GPT-5.4 Mini’s original v2.1 result. It does not claim a timeless culinary rank. The run’s scoring limitations and later exploratory analyses are separate records.

Original overall

96.0

95% item-bootstrap interval

93.7–98.1

Frontier subset

94.5

Original score profile

Uncertainty and pairwise evidence

GPT-5.4 Mini belonged to a five-model group from which the run’s paired item-bootstrap rule did not identify one unique leader.

Scoring limitations

Saturated numeric items, high-influence keyword items and semantic constraint-check failures affected the instrument. The profile is retained as historical evidence, not validated culinary ability.

Run 2026-07-v2.1 · openai/gpt-5.4-mini · original artifact unchanged