Archived v2.1 result / categories
Recipe Generation
Generating complete recipes under constraints: servings, allergens, time, equipment.
Original v2.1 category means (archived, not a ranking)
Sorted by score for readability. Category means were never tested for statistical separation; over as few as a dozen questions, an ordering at this granularity would be noise presented as precision.
- GPT-5.4 Mini94.7
- GPT-5.6 Terra Pro94.0
- GPT-5.6 Sol Pro93.6
- DeepSeek V4 Pro92.9
- Gemini 3.1 Pro Preview91.6
- Grok 4.591.1
- Gemini 3.6 Flash90.3
- Claude Fable 588.5
- Kimi K385.2
- Claude Sonnet 583.3
- Qwen 3.7 Max82.9
- Llama 4 Maverick82.3
- Claude Opus 580.1
- Mistral Large 370.8
Question heatmap (public questions only)
| Model | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 | 012 | 013 | 014 | 015 | 016 | 017 | 018 | 019 | 020 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.4 Mini | ||||||||||||||||||||
| GPT-5.6 Terra Pro | ||||||||||||||||||||
| GPT-5.6 Sol Pro | ||||||||||||||||||||
| DeepSeek V4 Pro | ||||||||||||||||||||
| Gemini 3.1 Pro Preview | ||||||||||||||||||||
| Grok 4.5 | ||||||||||||||||||||
| Gemini 3.6 Flash | ||||||||||||||||||||
| Claude Fable 5 | ||||||||||||||||||||
| Kimi K3 | ||||||||||||||||||||
| Claude Sonnet 5 | ||||||||||||||||||||
| Qwen 3.7 Max | ||||||||||||||||||||
| Llama 4 Maverick | ||||||||||||||||||||
| Claude Opus 5 | ||||||||||||||||||||
| Mistral Large 3 |
Each cell is one question; deeper colour = higher score. Hover for exact values.