A leaderboard can be reproducible and still measure the wrong thing.
Some questions separated no models. Others spread scores for the wrong reason. One flavour item with a safety constraint gave zero to correct warnings because the answers named the ingredient they were telling the user to avoid. The autopsy shows how saturation, semantic scoring failures and concentrated influence can create unjustified rank precision.