CookingBench

Which AI model is the best chef?

CookingBench scores models on the things that actually go wrong in a kitchen: scaling quantities, converting units, food safety, substitutions, technique, flavour logic and nutrition math.

Leaderboard · run 2026-06-v2methodology v2

2026-06-12
#ModelOverallFrontierRun cost
1GPT-5.4 MiniOpenAI96.4±2.494.1$0.58
2Grok 4.3xAI95.4±2.894.4$0.42
3GPT-5.5OpenAI94.5±3.489.1$2.39
4Claude Fable 5Anthropic92.7±4.084.0$4.82
5Claude Opus 4.8Anthropic92.6±3.885.4$2.09
6Gemini 3.5 FlashGoogle92.4±3.987.6$1.49
7Gemini 3.1 Pro PreviewGoogle91.6±4.389.4$1.58
8Kimi K2.6Moonshot AI91.5±4.189.5$1.48
9Qwen 3.5 PlusAlibaba91.4±3.991.8$0.52
10Claude Sonnet 4.6Anthropic91.3±4.189.5$3.07
11DeepSeek V4 ProDeepSeek90.5±4.186.3$0.38
12Mistral Large 3Mistral88.9±4.693.3$0.05
13Llama 4 MaverickMeta82.0±5.582.1$0.03

Categories