CookingBench

Results / v2.1 / xAI

Archived v2.1Unresolved leading group

Grok 4.5

This profile reports Grok 4.5’s original v2.1 result. It does not claim a timeless culinary rank. The run’s scoring limitations and later exploratory analyses are separate records.

Original overall

96.0

95% item-bootstrap interval

93.3–98.2

Frontier subset

98.6

Original score profile

Uncertainty and pairwise evidence

Grok 4.5 belonged to a five-model group from which the run’s paired item-bootstrap rule did not identify one unique leader.

Scoring limitations

Saturated numeric items, high-influence keyword items and semantic constraint-check failures affected the instrument. The profile is retained as historical evidence, not validated culinary ability.

Run 2026-07-v2.1 · x-ai/grok-4.5 · original artifact unchanged