CookingBench did not reliably discover the best AI cook—but it exposed how a benchmark can manufacture a convincing model ranking.
A forensic analysis of saturation, concentrated score influence and a semantic grading failure in the archived v2.1 run. No corrected winner is claimed.
Preliminary · retrospective · not peer reviewed