Benchmark Leaderboards Need Footnotes, Not Just Rankings
In February 2025, Hugging Face re-evaluated 3,751 model submissions on its Open LLM Leaderboard after finding problems in the way answers to the MATH-Hard benchmark were parsed and scored. Some mathematically correct answers had been rejected because their formatting did not match what the evaluator expected. The models had not changed. The questions had not changed. The measurement system had.
After the scoring correction, the average MATH-Hard result increased by 4.66 percentage points, the top-20 ordering on that benchmark was substantially rearranged, and some models moved more than 200 positions in the overall leaderboard.
That episode does not prove that leaderboards are useless. It proves something more consequential: a score belongs to an entire evaluation arrangement, not to the model alone.
A leaderboard rank is the last line of a calculation. Before that line come decisions about which model version was tested, which questions were included, how prompts were written, whether tools were available, how many attempts were allowed, what counted as a correct answer, how invalid responses were handled, which results were excluded, and how uncertainty was calculated. A ranking that hides those decisions compresses a complex measurement process into a deceptively simple ordinal claim.
For a research team, public agency, university, nonprofit, or procurement office, the correct response is neither to trust the ranking nor to dismiss it. The rank should be treated as a pointer to an evidence record.
The rank is a pointer, not a finding.
A defensible leaderboard should therefore carry a compact but substantive footnote: the exact system tested, the construct being measured, the evaluation protocol, the scorer or judge, the uncertainty around the result, the resource conditions under which it was produced, and the risks the benchmark does not measure. Without that record, a ranking may be useful for discovery, but it is not sufficient evidence of fitness, reliability, safety, or institutional suitability.
The ranking is the visible part of the instrument
Leaderboards make comparison possible by holding some parts of an evaluation constant. That is their legitimate value. A shared task, shared metric, and shared reporting format can reveal performance differences that would otherwise be difficult to inspect.
But the apparent simplicity of the final ranking can conceal what was actually compared.
A model name may refer to multiple releases, parameter configurations, quantization levels, system prompts, inference providers, or post-training variants. A benchmark may test factual recall, code generation, mathematical reasoning, instruction following, human preference, or some mixture of those constructs. An evaluation may use exact-match scoring, a rule-based parser, human raters, or another language model acting as judge. Each choice changes the interpretation of the result.
The National Institute of Standards and Technology’s AI Risk Management Framework Playbook makes this dependence explicit. It states that appropriate metrics vary according to the system’s purpose, audience, deployment context, and evaluation needs. It also recommends documenting why metrics were selected, which relevant risks were not measured, whether results generalize beyond the test conditions, and whether the metrics remain effective for their intended use.
A single overall score cannot carry that information. Neither can a model’s position relative to its neighbors.
The first job of a leaderboard footnote is therefore to restore the identity of the evaluated object. It should name the model release, access method, evaluation date, inference configuration, prompt or system-message conditions, tool availability, and any provider-specific settings that materially affect behavior. “Model X ranked third” is not an auditable claim when “Model X” is not a stable experimental object.
This is not clerical detail. Version ambiguity can make a comparison impossible to reproduce and can leave institutions relying on evidence that no longer describes the system available to them.
A benchmark measures a construct, not general competence
The phrase “best model” usually outruns the evidence underneath it.
A mathematics benchmark can support a claim about performance on the mathematical tasks and scoring rules represented in that benchmark. A coding benchmark can support a claim about the included programming problems under the specified execution environment. A human-preference leaderboard can describe which responses participating raters preferred under a particular interface, sampling system, and prompt distribution.
None of these measurements, by themselves, establish that a model is safer, more accurate in another domain, more accessible, less biased, better suited to government records, or more reliable under operational pressure.
The Holistic Evaluation of Language Models project was created partly in response to evaluation practices that reduced models to a small number of benchmark scores. HELM evaluates models across multiple scenarios and metrics, releases prompts and model responses for inspection, and identifies important capabilities and risks that remain unmeasured. Its structure recognizes that model performance is multidimensional and that an evaluation should disclose its omissions rather than allowing an aggregate score to imply completeness.
A leaderboard footnote should consequently state the claim the benchmark is designed to support in one precise sentence. It should also identify the claims it cannot support.
For a public administrator, that distinction is decisive. A ranking on general question answering does not establish that a model can accurately summarize a jurisdiction’s regulations. A preference score does not establish that citations are correct. A coding score does not establish that generated software is secure. A broad multilingual average does not establish reliability in the specific languages used by a community.
The benchmark must be relevant to the contemplated use, not merely prestigious.
The scorer can become the hidden contestant
The Hugging Face correction exposed a recurring evaluation problem: the scoring system can influence the ranking as strongly as the model responses do.
Exact-match metrics can reject semantically correct answers that use unexpected wording or formatting. Flexible parsers can introduce their own interpretation errors. Model-based judges may prefer certain response styles, verbosity levels, structures, or linguistic patterns. Human raters may disagree, misunderstand the task, or apply criteria inconsistently.
The scorer is therefore part of the evaluated system.
A serious footnote should identify whether results were produced through exact matching, executable tests, rule-based parsing, expert review, crowd judgment, or model-based evaluation. It should disclose the scoring implementation and version where possible, along with exclusion rules, failure handling, tie treatment, and any manual corrections.
The same requirement applies when evaluation code changes. A leaderboard should preserve the relationship between each reported result and the evaluator version that produced it. Recomputing historical scores with a revised scorer may improve comparability, but the revision should remain visible. Otherwise, readers cannot tell whether a model changed, a benchmark changed, or the interpretation of a response changed.
MLCommons’ benchmark submission rules illustrate a stronger evidence model. Depending on the benchmark division, submissions are expected to include system metadata, implementation information, code, setup and execution scripts, and result logs. These materials do not make every benchmark conclusion correct, but they make the reported result substantially more inspectable and reproducible.
A ranking without comparable supporting records asks the reader to accept the output of an instrument that cannot be examined.
Human-preference leaderboards measure a real but bounded signal
Human-preference evaluation addresses weaknesses in static benchmarks. It can capture qualities that exact-match tests miss, including usefulness, clarity, tone, and perceived responsiveness across prompts supplied by actual users.
Chatbot Arena was designed around this idea. Users submit prompts, compare anonymous model responses, and select the response they prefer. The original research paper described a large-scale pairwise-comparison system intended to complement static benchmarks that may suffer from saturation, contamination, narrow task coverage, or weak alignment with human judgment.
The current Arena leaderboard also reports more than a bare rank. It displays battle counts, average win rates, pairwise win fractions, and bootstrapped confidence intervals, providing readers with information about sample size and statistical uncertainty.
Those disclosures strengthen the leaderboard. They do not convert preference into universal quality.
Arena results reflect the preferences of participating users, the prompts those users submit, the models shown to them, the sampling and weighting policies in effect, and the criteria participants choose to apply. A user may prefer an answer because it is more accurate, more direct, more detailed, more agreeable, better formatted, or simply closer to the user’s expectations. The vote alone does not reveal which property determined the choice.
The platform’s policies also allow some unreleased models to be tested anonymously and establish minimum-vote conditions before public ranking. Such practices may help developers obtain pre-release feedback, but they also make model status, sampling rules, provisional results, and disclosure timing relevant to interpretation.
These design choices have been contested. A research critique of Arena argued that private testing, selective disclosure, sampling differences, and other structural conditions could affect the meaning of the public rankings. Arena disputed several of the paper’s conclusions while accepting some recommendations for clearer provisional labels, retirement indicators, and methodological disclosure.
The dispute should not be reduced to a verdict about whether Arena is trustworthy. It demonstrates why leaderboard governance belongs in the evidence record. When credible parties disagree about how a ranking was produced or interpreted, the methodology must be visible enough for others to evaluate the disagreement.
The strongest-looking number may still be the wrong proxy
Adding confidence intervals, sample counts, evaluator versions, and prompt records produces a better leaderboard. Yet even a statistically careful ranking can remain unsuitable for an institutional decision.
This is the deeper problem.
Uncertainty estimates describe uncertainty within the measurement design. They do not prove that the measurement design represents the institution’s real task.
A model can lead a benchmark with a narrow confidence interval and still fail on the documents, languages, accessibility requirements, latency constraints, privacy rules, or error costs that matter to a particular organization. Statistical precision cannot repair a mismatch between the tested construct and the intended use.
The most important leaderboard footnote is therefore not a technical parameter. It is a relevance statement.
The publisher should state who the benchmark is for, what decision it is meant to inform, what task population it represents, and which use contexts remain outside its scope. The institution reading the leaderboard must then decide whether its own task falls inside that scope.
For government and public-interest uses, this relevance review should be demanding. The costs of an error may be unevenly distributed. A model that performs well on average may fail systematically on uncommon documents or underrepresented language varieties. A preference benchmark may reward fluent explanations without verifying the controlling source. A benchmark conducted with short, self-contained prompts may say little about performance across long administrative records with conflicting provisions.
These are not reasons to reject benchmark evidence. They are reasons to prevent the benchmark from making a larger claim than it earned.
What the footnote must make visible
The minimum leaderboard footnote begins with the object being ranked. It should identify the exact model or system version, evaluation date, access route, configuration, prompt conditions, tool permissions, and material preprocessing or postprocessing. A model family name is insufficient when the tested release cannot be reconstructed.
It must then name the question being measured. The benchmark publisher should define the target construct, intended users, represented tasks, dataset or prompt population, language coverage, and exclusion criteria. The reader should be able to distinguish a test of mathematical-answer parsing from a test of mathematical reasoning, or a test of user preference from a test of factual correctness.
The protocol must be inspectable. This includes prompts, templates, repetition rules, randomization, sampling, evaluator instructions, scorer code or specification, model-judge identity, handling of refusals and invalid outputs, and conditions under which results are removed or recomputed. Where full release is constrained by security, licensing, privacy, or benchmark-integrity concerns, the restriction and its consequences should be stated rather than hidden.
The footnote should report uncertainty without pretending that rank is more stable than the data allow. Sample sizes, confidence intervals, tie treatment, sensitivity analyses, and meaningful performance tiers are often more informative than a strict first-to-last order. Two systems whose estimated performance is difficult to distinguish should not be presented as though the ordering were certain.
Resource conditions belong in the same record. Cost, latency, hardware, inference provider, context length, tool use, repeated sampling, and other computational conditions can determine whether a result is reproducible or operationally relevant. A system that reaches a score through extensive sampling or external tools is being evaluated under different conditions from one producing a single unaided response.
Finally, the footnote must state what was not measured. Safety, factual grounding, accessibility, multilingual reliability, privacy, robustness, environmental cost, legal suitability, and performance after deployment should not be implied by silence. NIST’s guidance is particularly useful here: unmeasured risks should be documented, not absorbed into a general impression of model quality.
A leaderboard without this information is compression without custody. The result has been separated from the evidence needed to understand who produced it, under what conditions, and for what claim.
How an institution should use the ranking
A small public-interest institution does not need frontier-scale compute to evaluate a leaderboard responsibly. It needs a disciplined rule for deciding what the ranking is allowed to do.
The ranking may begin a review. It should not end one.
An institution can use a leaderboard to identify candidate systems when the benchmark’s construct is plausibly relevant. It should then inspect the footnote and underlying documentation for reproducibility, uncertainty, task coverage, model identity, and unmeasured risks. Only after that review should the institution conduct a bounded confirmation using representative tasks, realistic source materials, appropriate human reviewers, and a simpler baseline.
The decision rule is straightforward: no rank should advance an option unless the supporting record establishes relevance, reproducibility, and bounded uncertainty for the contemplated use.
When the record is incomplete, the proper conclusion is not that the model is poor. It is that the ranking has not supplied enough evidence for that decision.
This distinction protects both institutions and model developers. It prevents a narrow benchmark result from becoming an unsupported institutional endorsement, while allowing well-designed evaluations to contribute useful evidence within their actual scope.
Better disclosure makes leaderboards more useful
The strongest argument for footnotes is not that leaderboards are inherently misleading. It is that disclosure preserves their legitimate value.
HELM demonstrates the value of exposing scenarios, metrics, prompts, responses, and acknowledged evaluation gaps. MLCommons demonstrates the value of connecting reported results to system descriptions, code, execution procedures, and logs. Arena demonstrates how vote counts, pairwise comparisons, confidence intervals, and public methodology can make human-preference rankings more interpretable, even while questions about sampling and governance remain open.
A well-documented leaderboard can help researchers find anomalies, compare methods, identify promising systems, and detect where further testing is warranted. A poorly documented leaderboard encourages readers to substitute position for understanding.
The goal is not to make every leaderboard carry an entire research paper beneath each score. The goal is to ensure that every rank can be traced to enough evidence for a technically competent reader to understand its claim, reproduce or challenge its production, and refuse interpretations the benchmark cannot support.
The institutional standard should be higher than the visual interface
Leaderboards are designed to be scanned. Institutional evidence must be designed to be examined.
That difference matters whenever rankings influence procurement, research agendas, media coverage, funding decisions, or public claims about which systems are “leading.” The cleaner the interface, the easier it becomes to forget the unresolved choices beneath it.
TAIRC’s Open LLM Transparency and Evaluation Frameworks program is explicitly analytical, evaluative, and documentation-focused. Its published program boundary states that TAIRC does not train or fine-tune proprietary systems through this initiative, deploy models in production, certify models as safe, or provide regulatory approval. The program’s proposed frameworks and toolkits are described as planned outputs rather than completed research products.
This article therefore does not rank or endorse any model, and it does not report a completed TAIRC benchmark. It proposes a public-interest publication standard: every ranking should remain attached to the evidence that gives it meaning.
The final question is not, “Which model is number one?”
It is, “Number one at what, measured how, under which conditions, with what uncertainty, and for whose decision?”
Until those questions are answered, the ranking is a starting point. The footnote is where the evidence begins.
Sources and verification
Hugging Face, “Fixing the Open LLM Leaderboard with Math-Verify.” Used to verify the February 2025 re-evaluation of 3,751 submissions, the MATH-Hard scoring problem, the average score change, and reported ranking movements following the evaluator correction.
National Institute of Standards and Technology, AI Risk Management Framework Playbook, Measure function. Used for guidance concerning purpose-specific metrics, intended audiences, evaluation context, generalizability, documentation of metric selection, and disclosure of risks that remain unmeasured.
Percy Liang and coauthors, “Holistic Evaluation of Language Models,” and the official HELM documentation. Used to verify HELM’s multi-scenario, multi-metric approach, its emphasis on transparent evaluation, its release of prompts and model responses, and its acknowledgement of evaluation gaps.
Wei-Lin Chiang and coauthors, “Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.” Used to verify Arena’s pairwise, crowdsourced evaluation design and the reasons its authors give for complementing static benchmarks with live human-preference evidence.
LMArena, public leaderboard, model evaluation policy, and response to “The Leaderboard Illusion.” Used to verify current public disclosures such as vote counts and confidence intervals, policies governing pre-release and publicly ranked models, and Arena’s documented response to methodological criticism.
Evan Frick and coauthors, “The Leaderboard Illusion.” Used only to represent the authors’ methodological criticisms of Arena. Claims disputed by Arena are presented as contested rather than established findings.
MLCommons, MLPerf submission rules. Used to verify documentation requirements involving system metadata, implementation information, code, execution scripts, and result logs in benchmark submissions.
The AI Research Center, Open LLM Transparency and Evaluation Frameworks. Used to verify TAIRC’s published program scope, organizational boundaries, and the status of proposed outputs.


