Archived v2.1 result / categories
Technique
Troubleshooting failures (split sauces, dense bread) and method advice.
Original v2.1 category means (archived, not a ranking)
Sorted by score for readability. Category means were never tested for statistical separation; over as few as a dozen questions, an ordering at this granularity would be noise presented as precision.
- Grok 4.599.4
- GPT-5.6 Terra Pro99.1
- Claude Opus 598.6
- GPT-5.6 Sol Pro98.4
- Kimi K398.4
- Gemini 3.6 Flash96.7
- Claude Fable 595.5
- Qwen 3.7 Max94.1
- GPT-5.4 Mini93.9
- Gemini 3.1 Pro Preview93.4
- Claude Sonnet 593.4
- DeepSeek V4 Pro92.8
- Mistral Large 390.6
- Llama 4 Maverick85.5
Question heatmap (public questions only)
| Model | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 | 012 | 013 | 014 | 015 | 016 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Grok 4.5 | ||||||||||||||||
| GPT-5.6 Terra Pro | ||||||||||||||||
| Claude Opus 5 | ||||||||||||||||
| GPT-5.6 Sol Pro | ||||||||||||||||
| Kimi K3 | ||||||||||||||||
| Gemini 3.6 Flash | ||||||||||||||||
| Claude Fable 5 | ||||||||||||||||
| Qwen 3.7 Max | ||||||||||||||||
| GPT-5.4 Mini | ||||||||||||||||
| Gemini 3.1 Pro Preview | ||||||||||||||||
| Claude Sonnet 5 | ||||||||||||||||
| DeepSeek V4 Pro | ||||||||||||||||
| Mistral Large 3 | ||||||||||||||||
| Llama 4 Maverick |
Each cell is one question; deeper colour = higher score. Hover for exact values.