Archived v2.1 result / categories
Substitutions
Ingredient swaps with correct ratios, including allergen-aware alternatives.
Original v2.1 category means (archived, not a ranking)
Sorted by score for readability. Category means were never tested for statistical separation; over as few as a dozen questions, an ordering at this granularity would be noise presented as precision.
- GPT-5.6 Sol Pro95.8
- Grok 4.595.7
- Claude Opus 595.7
- Kimi K395.5
- Claude Fable 595.2
- GPT-5.6 Terra Pro93.1
- Gemini 3.6 Flash91.6
- Mistral Large 391.6
- Claude Sonnet 589.7
- GPT-5.4 Mini88.7
- Qwen 3.7 Max87.1
- DeepSeek V4 Pro83.1
- Llama 4 Maverick78.5
- Gemini 3.1 Pro Preview75.0
Question heatmap (public questions only)
| Model | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 | 012 | 013 | 014 | 015 | 016 | 017 | 018 | 019 | 020 | 021 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol Pro | |||||||||||||||||||||
| Grok 4.5 | |||||||||||||||||||||
| Claude Opus 5 | |||||||||||||||||||||
| Kimi K3 | |||||||||||||||||||||
| Claude Fable 5 | |||||||||||||||||||||
| GPT-5.6 Terra Pro | |||||||||||||||||||||
| Gemini 3.6 Flash | |||||||||||||||||||||
| Mistral Large 3 | |||||||||||||||||||||
| Claude Sonnet 5 | |||||||||||||||||||||
| GPT-5.4 Mini | |||||||||||||||||||||
| Qwen 3.7 Max | |||||||||||||||||||||
| DeepSeek V4 Pro | |||||||||||||||||||||
| Llama 4 Maverick | |||||||||||||||||||||
| Gemini 3.1 Pro Preview |
Each cell is one question; deeper colour = higher score. Hover for exact values.