CookingBench
Preserved primary materialImmutable archive

The v2.1 response corpus

Every planned answer from 14 model versions to 184 culinary prompts, retained exactly as returned. The primary material is valuable even though the original scoring is known to be unreliable.

2,576
response artifacts
14
model versions
184
prompts per model
0
empty answer texts
1,998,922
Unicode code points of answer text
2,015,885
UTF-8 answer-text bytes
3
content-filter finish states, disclosed
184/184
coverage for every model version

What this record proves

The archive contains 14 rostered model versions and exactly 2,576 nonempty response artifacts. Each response retains its model, prompt, answer text, provider envelope, token counts, cost, latency and finish state.

It proves corpus completeness and preservation. It does not validate the original scores, prove that the answers are correct or turn an exploratory regrading into confirmatory evidence.

Integrity record

Run id
2026-07-v2.1
Response content digest
cc5b5db1988921298ff72cf034ced530820d18a821009b15f56e9dfee9f40e8b
Git response tree
0e574bdf2ec63ca55067d0dd12bd0756eddec420
Digest scope
Sorted response filenames and content in the immutable v2.1 response set.
Score status
Known unreliable as a ranking; preserved separately from answers.
Audit status
Agent-produced, unblinded classifications; independent human validation pending.
Permitted uses
  • Independent exploratory regrading under a declared scheme.
  • Blinded culinary adjudication and judge-validation studies.
  • Research into scoring, saturation and evaluation failure.
  • Regression cases for repaired graders and new question contracts.
Not supported
  • Citing the original leaderboard as culinary-ability ground truth.
  • Publishing a corrected winner from post-hoc deletions alone.
  • Calling AI text judgement tasted flavour or human preference.
  • Describing the corpus as independent validation of the benchmark.