The Case for Publishing Negative Evaluation Results
An evaluation team tests a language model against a task it was expected to handle. The model underperforms the baseline. The team checks the data, reruns the protocol, examines the scoring code, and reaches the same conclusion.
Then the result disappears.
Nothing was falsified. No chart was altered. No claim was technically withdrawn. The finding simply never enters the public record because it does not support a launch, a paper, a procurement decision, or a story of progress.
That silence changes the evidence base. Future evaluators see the successful trials, the favorable comparisons, and the polished system cards. They do not see the reasonable approaches that failed, the models that broke under ordinary variation, or the improvements that vanished when a stronger baseline was used. The record begins to look more certain than the underlying work ever was.
Negative AI evaluation results should be published when the evaluation was methodologically sound, the tested claim was meaningful, and the report gives readers enough context to interpret the failure. Publication should be delayed or limited only when disclosure would expose a live security vulnerability, violate privacy or licensing obligations, or create a foreseeable risk that cannot yet be mitigated.
The principle is simple: an institution cannot claim to evaluate AI responsibly while publishing only the results that make adoption easier.
A negative result is evidence, not an embarrassment
“Negative result” is an imprecise label. It can describe a failed attempt to outperform a baseline, a hypothesis that was not supported, a replication that did not reproduce an earlier claim, a safety test that revealed unacceptable behavior, or an evaluation that found no reliable difference between systems. It can also describe a result that is mixed across languages, user groups, task types, or deployment conditions.
These findings do not all mean the same thing. A model that performs no better than a simpler baseline is different from a model that fails under adversarial prompting. A null result is different from an inconclusive one. A failed replication is different from a badly designed test. Publishing them under one vague heading would create a new form of opacity.
The useful distinction is between a negative result and an invalid result. A negative result answers a meaningful question with evidence that resists the expected claim. An invalid result cannot support a conclusion because the method, data, implementation, or statistical reasoning is too weak. The first belongs in the evidence base. The second belongs in a methods review, and may still deserve documentation if its failure teaches others how an evaluation can go wrong.
Machine-learning researchers have argued that a culture centered on predictive gains creates incentives to suppress findings that do not show improvement, even when those findings would save others from repeating unproductive work. Their case is not that every failed experiment deserves a paper. It is that methodological value cannot be reduced to whether the final number moved upward. (arXiv)
Selective reporting produces a false map
A public evidence base should tell readers where a claim holds, where it weakens, and where it fails. Selective reporting removes the last two regions.
The effect is familiar across science. Registered Reports were developed so that publication decisions could be made from the importance of the question and the quality of the method before the results were known. In one comparison of psychology papers, 96 percent of conventionally published reports supported the tested hypothesis, compared with 44 percent of Registered Reports. The authors did not treat that gap as proof that every conventional paper was wrong. They treated it as evidence that publication practices can strongly shape which outcomes become visible. (OSF)
AI evaluation has the same structural vulnerability, with an added complication: results are highly conditional. A score can depend on the model snapshot, prompt template, sampling settings, tool access, retrieval layer, test-set construction, scoring parser, language, hardware, and date. A result that disappears is therefore more than a missing number. It may be the only public evidence that a claimed capability collapses under a realistic change in conditions.
A 2024 review of machine-learning methods for fluid-related partial differential equations illustrates the risk. Among papers claiming that machine-learning methods outperformed conventional numerical approaches, the authors reported that 79 percent of those claims relied on weak baselines. They also found evidence of outcome-reporting and publication bias. The lesson is broader than that field: a literature can look consistently successful when unfavorable comparisons and stronger baselines are less likely to appear. (arXiv)
When public administrators, universities, or nonprofit organizations read that literature, they do not experience the bias as an abstract statistical problem. They experience it as a distorted purchasing and governance environment. The apparent consensus may be partly a record of what researchers and vendors chose to show.
AI evaluation is especially vulnerable to success filtering
Traditional experiments usually aim to isolate a defined hypothesis. AI evaluations often mix several questions at once.
A benchmark may be presented as evidence of general capability even though it measures a narrow task. A red-team exercise may generate hundreds of failures but publish only a summarized risk category. A vendor comparison may exclude prompts that produced unstable outputs. A pilot may be described as successful because users completed a workflow, while the review process quietly absorbed the model’s errors. A multilingual test may report an average that conceals severe degradation in lower-resource languages.
Each choice can be defensible. Together, they can turn an evaluation report into a curated surface.
The Model Cards proposal responded to a related problem by calling for documentation of intended uses, evaluation procedures, performance characteristics, and contexts in which a model may be unsuitable. NIST’s Generative AI Profile likewise calls for documentation of knowledge limits, past incidents and failure modes, test and evaluation methods, and the conditions in which human oversight is required. The profile also warns against relying on quantitative metrics without enough contextual analysis. (arXiv)
Neither source says that every raw test output must be released. They support a more demanding idea: evaluation evidence should be sufficient for the people who act on it. If an institution withholds a negative result that would materially change a procurement, deployment, or research decision, the remaining report is incomplete even when every published number is accurate.
The real product of evaluation is a boundary
The usual story treats an evaluation as a gate. The system passes or fails, and publication follows the result.
That is too crude.
A serious evaluation produces a boundary around a claim. It shows the tasks, populations, environments, and conditions in which the available evidence supports reliance. It also shows where the evidence stops. Positive results fill in part of that boundary. Negative results draw the rest.
A negative result is not an empty box. It is a boundary marker.
This is the point that changes the institutional case for publication. The purpose is not to celebrate failure or punish a vendor. The purpose is to prevent an unsupported claim from traveling farther than the evidence beneath it.
For a public agency, that boundary can determine whether a system is limited to low-consequence drafting or allowed into a workflow that affects residents. For a university, it can determine whether a model is suitable for research assistance in one discipline but unreliable in another. For a funder, it can distinguish a project that learns from evidence from one that reports only favorable milestones. For a procurement team, it can reveal that a cheaper baseline performs as well as a complex system, or that a claimed advantage disappears when the test resembles actual use.
The deeper value is institutional memory. Staff change. Vendors update models. Pilot teams dissolve. A negative result that remains only in a meeting, private notebook, or temporary dashboard cannot protect the next decision-maker. Publication converts local disappointment into durable evidence.
What a publishable negative result must contain
A negative result becomes useful only when a reader can reconstruct what failed.
The report should identify the exact claim that was tested. “The model performed poorly” is too weak. A reader needs to know whether the claim concerned factual accuracy, robustness, multilingual consistency, latency, cost, accessibility, safety, or performance against a defined baseline.
The evaluated system must be identifiable. That means the model name or accessible identifier, version or snapshot where available, evaluation date, surrounding system components, prompt or instruction structure, sampling settings, tool access, retrieval configuration, and any material safety filters. Without that record, the result may describe a system that no longer exists.
The method must be inspectable. The report should explain the dataset or test-case source, inclusion and exclusion rules, sample size, task construction, scoring method, human-review procedure, baseline, repeated-run design, and uncertainty. If data or code cannot be released, the reason should be stated and enough procedural detail should remain for an independent reviewer to assess the conclusion.
The failure should be localized. Readers need to know whether the result was broad, concentrated in a subset, triggered by a particular condition, or sensitive to a design choice. A single average can conceal the difference between mild degradation everywhere and complete failure for one affected group.
Alternative explanations must be considered. A negative outcome may come from an unsuitable prompt, contaminated data, a broken parser, an underpowered sample, a mismatched baseline, or an implementation error. A credible report explains which alternatives were checked, which remain possible, and how strongly the evidence supports the conclusion.
The institutional consequence should be explicit. The report should state what changed because of the finding: a claim was narrowed, a use case was rejected, a safeguard was added, a procurement was paused, a benchmark was revised, or no action was taken because the evidence remained inconclusive. This is where evaluation becomes governance.
Current NeurIPS guidance reflects the same emphasis on transparency and reproducibility. Its paper checklist is part of the submission and publication record, asking authors to account for methodological and ethical details rather than leaving them outside the paper’s evidence chain. (NeurIPS)
Publication does not mean indiscriminate disclosure
There are legitimate reasons to control the timing and detail of a negative result.
A red-team evaluation may uncover a live vulnerability that could be exploited before a provider or operator can mitigate it. Publishing exact prompts, attack paths, credentials, or system details immediately may increase risk. In those cases, the correct alternative is coordinated disclosure with documented timelines, responsible parties, mitigation status, and a later public account where feasible. NIST’s vulnerability-disclosure guidance is built around receiving, assessing, managing, and communicating vulnerability reports rather than choosing between permanent secrecy and immediate release. (NIST Computer Security Resource Center)
Privacy creates another boundary. Evaluation data may contain personal information, sensitive records, protected communications, or examples that could identify individuals. The result can often be reported without releasing the underlying material. Aggregation, redaction, synthetic examples, controlled-access artifacts, or a detailed method statement may preserve public value without exposing the data.
Licensing and contractual restrictions also matter, although they should not become a blanket excuse. An institution should distinguish between information it is legally unable to disclose and information it has merely agreed not to disclose for convenience. Where a contract prevents publication of material findings that would affect public risk, the governance failure may have occurred when the contract was signed.
There is also a quality threshold. A team should not rush out a negative claim before checking the evaluation code, baseline, data, and interpretation. Public correction is healthy; preventable error is still preventable. The obligation is to publish sound negative evidence, not to publish every disappointing run.
The minimum public record
An institution does not need a journal-length paper for every negative evaluation. It needs a stable, searchable record with enough information to support scrutiny.
That record should preserve the tested claim, system identity, date, intended use, evaluation conditions, data provenance, baseline, metric, uncertainty, observed failure, subgroup or scenario effects, alternative explanations, limitations, reviewer or owner, resulting decision, disclosure constraints, and revision history. It should link to code, test cases, or controlled artifacts when those can be released responsibly.
The format matters less than the discipline. A concise evaluation card can be sufficient. So can a technical report, benchmark appendix, incident record, or versioned repository entry. What fails is the slide deck that disappears, the dashboard with no archival state, or the prose statement that says a system “did not meet expectations” without recording what those expectations were.
Publication should also be versioned. A later model release may correct the failure. A revised prompt may change the result. A stronger dataset may overturn the original conclusion. The negative result should remain visible with an update, not be deleted as though it never happened. Evidence can age without becoming dishonest.
What changes when institutions publish failure
Publishing negative evaluations changes incentives before the evaluation begins.
Teams become more likely to define claims precisely because the result will remain visible either way. Baselines receive more attention. Evaluation code is treated as part of the institutional record. Product and research teams have less reason to tune the method toward a desired outcome. Procurement officers gain evidence for rejecting exaggerated claims. Funders can see whether a project is learning or merely curating success.
It also changes what counts as progress.
A project that demonstrates that a model cannot reliably support a proposed public-facing use may have produced more public value than a project that reports a modest benchmark gain. The first prevents misplaced reliance. The second may change nothing outside the paper.
This does not mean failure should be rewarded automatically. A badly conceived project remains badly conceived. Repeated negative results can reveal weak planning, poor execution, or a research question that no longer matters. Transparency makes those judgments possible. Silence protects both valuable failures and avoidable ones from examination.
A practical publication rule
The decision can be stated without a complicated scoring system.
Publish the negative result when the tested claim matters, the method is credible, the result would change a reasonable reader’s belief or action, and the report can be released without creating a greater unmitigated harm.
Revise and rerun before publication when the result may be explained by a correctable defect in the evaluation.
Publish as inconclusive when the question matters but the available evidence cannot distinguish among plausible explanations.
Use coordinated or delayed disclosure when immediate detail would create a live security or safety risk.
Do not suppress a sound result because it weakens a preferred narrative, complicates a funding report, undermines a procurement case, or shows that a simpler system was sufficient.
That final category is the one institutions most need to govern. The strongest pressure to hide a result often appears precisely when the result has the greatest decision value.
TAIRC’s position
TAIRC’s Open LLM Transparency and Evaluation Frameworks program places disclosure inside the evaluation method itself, rather than treating it as institutional messaging. Its stated focus is reproducible evaluation, structured reporting, uncertainty disclosure, and public accountability. The program is analytical and documentation-focused; it does not train proprietary frontier-scale models, certify systems, or present planned research as completed work.
Within that boundary, negative-result publication should become part of the evaluation architecture. A future TAIRC reporting framework should require evaluators to preserve unfavorable, null, mixed, and inconclusive outcomes; distinguish valid negative evidence from invalid tests; document security or privacy restrictions; and retain a revision history when later evidence changes the conclusion.
That is a proposed standard, not a claim that TAIRC has already produced or validated such a framework. It follows TAIRC’s broader commitments to test systems before they are used, build oversight into the research process, and make research outputs available for broad public benefit rather than exclusive control.
The result should not disappear
Return to the evaluation team whose model failed to beat the baseline.
If the method was weak, the team should fix it. If the evidence was inconclusive, the report should say so. If disclosure would expose a live vulnerability, the finding should enter a coordinated process. But if the evaluation was sound and the result would matter to another decision-maker, silence is not caution.
It is selective reporting.
The public record of AI performance will never be complete. Models change, contexts differ, and every evaluation leaves something out. That makes the publication of credible negative results more necessary, not less. Without them, institutions see a field composed mainly of successes and are left to discover the boundaries through their own failures.
Responsible evaluation should make those boundaries visible before someone is asked to rely on the system.
The result should not disappear.
This article presents a research and governance position. It is not legal, cybersecurity, procurement, or regulatory advice. Decisions involving sensitive systems, protected data, contractual restrictions, or live vulnerabilities require appropriate domain and legal review.


