CookingBench

Results / v2.1 / DeepSeek

Archived v2.1

DeepSeek V4 Pro

This profile reports DeepSeek V4 Pro’s original v2.1 result. It does not claim a timeless culinary rank. The run’s scoring limitations and later exploratory analyses are separate records.

Original overall

92.7

95% item-bootstrap interval

88.8–96.0

Frontier subset

88.2

Original score profile

Uncertainty and pairwise evidence

This model’s position should be interpreted through direct paired comparisons, not its row number alone.

Scoring limitations

Saturated numeric items, high-influence keyword items and semantic constraint-check failures affected the instrument. The profile is retained as historical evidence, not validated culinary ability.

Run 2026-07-v2.1 · deepseek/deepseek-v4-pro · original artifact unchanged