Archived v2.1 result / categories
Flavour Pairing
Pairing logic, cuisine coherence and balancing dishes.
Original v2.1 category means (archived, not a ranking)
Sorted by score for readability. Category means were never tested for statistical separation; over as few as a dozen questions, an ordering at this granularity would be noise presented as precision.
- GPT-5.4 Mini98.4
- Gemini 3.1 Pro Preview98.4
- Grok 4.597.7
- Mistral Large 395.6
- DeepSeek V4 Pro94.6
- GPT-5.6 Sol Pro93.8
- GPT-5.6 Terra Pro93.6
- Gemini 3.6 Flash90.6
- Kimi K390.5
- Qwen 3.7 Max89.2
- Claude Fable 588.0
- Claude Opus 587.4
- Claude Sonnet 586.8
- Llama 4 Maverick84.8
Question heatmap (public questions only)
| Model | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 | 012 | 013 | 014 | 015 | 016 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.4 Mini | ||||||||||||||||
| Gemini 3.1 Pro Preview | ||||||||||||||||
| Grok 4.5 | ||||||||||||||||
| Mistral Large 3 | ||||||||||||||||
| DeepSeek V4 Pro | ||||||||||||||||
| GPT-5.6 Sol Pro | ||||||||||||||||
| GPT-5.6 Terra Pro | ||||||||||||||||
| Gemini 3.6 Flash | ||||||||||||||||
| Kimi K3 | ||||||||||||||||
| Qwen 3.7 Max | ||||||||||||||||
| Claude Fable 5 | ||||||||||||||||
| Claude Opus 5 | ||||||||||||||||
| Claude Sonnet 5 | ||||||||||||||||
| Llama 4 Maverick |
Each cell is one question; deeper colour = higher score. Hover for exact values.