Archived v2.1 result / categories
Conversions
Volume, weight and temperature conversions across kitchen units and locales.
Original v2.1 category means (archived, not a ranking)
Sorted by score for readability. Category means were never tested for statistical separation; over as few as a dozen questions, an ordering at this granularity would be noise presented as precision.
- GPT-5.4 Mini100.0
- GPT-5.6 Sol Pro100.0
- Grok 4.5100.0
- GPT-5.6 Terra Pro100.0
- Claude Fable 5100.0
- Kimi K3100.0
- Gemini 3.6 Flash100.0
- Claude Opus 5100.0
- Gemini 3.1 Pro Preview100.0
- DeepSeek V4 Pro100.0
- Claude Sonnet 5100.0
- Qwen 3.7 Max100.0
- Mistral Large 394.4
- Llama 4 Maverick88.9
Question heatmap (public questions only)
| Model | 001 | 002 | 003 | 004 | 005 | 006 | 007 | 008 | 009 | 010 | 011 | 012 | 013 | 014 | 015 | 016 | 017 | 018 | 019 | 020 | 021 | 022 | 023 | 024 | 025 | 026 | 027 | 028 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.4 Mini | ||||||||||||||||||||||||||||
| GPT-5.6 Sol Pro | ||||||||||||||||||||||||||||
| Grok 4.5 | ||||||||||||||||||||||||||||
| GPT-5.6 Terra Pro | ||||||||||||||||||||||||||||
| Claude Fable 5 | ||||||||||||||||||||||||||||
| Kimi K3 | ||||||||||||||||||||||||||||
| Gemini 3.6 Flash | ||||||||||||||||||||||||||||
| Claude Opus 5 | ||||||||||||||||||||||||||||
| Gemini 3.1 Pro Preview | ||||||||||||||||||||||||||||
| DeepSeek V4 Pro | ||||||||||||||||||||||||||||
| Claude Sonnet 5 | ||||||||||||||||||||||||||||
| Qwen 3.7 Max | ||||||||||||||||||||||||||||
| Mistral Large 3 | ||||||||||||||||||||||||||||
| Llama 4 Maverick |
Each cell is one question; deeper colour = higher score. Hover for exact values.