A Minimum Viable Evidence Package for Public-Sector AI Procurement
A public agency can receive hundreds of pages from an AI vendor and still lack the evidence needed to make a defensible purchase.
The proposal may include a model card, benchmark scores, a security presentation, an accessibility statement, a demonstration and several pages of responsible-AI principles. Yet the acquisition team may still be unable to answer basic questions: Which model and configuration produced the demonstrated results? Were the tests representative of the agency’s work? What happens when the provider changes the model? Can the agency investigate a failure, reproduce an evaluation or leave the service without losing its data?
That is the difference between documentation and evidence.
A minimum viable evidence package is not the smallest collection of vendor documents an agency can obtain. It is the smallest connected record that allows an independent reviewer to understand the proposed use, examine the material claims, identify who retains authority, and determine whether the system should be accepted, constrained, retested or rejected.
The unit of assurance is not the model in the abstract. It is the model, embedded in a particular system, performing a defined public function under enforceable contractual conditions.
This article uses current U.S. federal guidance as a developed reference point. Federal requirements do not automatically govern state, local, tribal or foreign public bodies, which must apply their own procurement, accessibility, privacy, records, civil-rights and sector-specific rules. The analysis is not legal advice or a certification framework.
TAIRC’s Open LLM Transparency and Evaluation Frameworks program is similarly limited to evaluation, documentation and reproducibility research. It does not certify, endorse or approve AI products.
Minimum viable means decision-complete
There is no universal packet that can establish the acceptability of every AI acquisition. A document-summarization assistant used on public reports does not require the same depth of evidence as a system whose output may materially affect benefits, employment, public safety or access to government services.
The package must therefore be proportionate. But proportionality cannot become permission to omit an entire class of evidence.
At minimum, the record must establish the intended public use, the identity of the system being acquired, the data and technical conditions on which it depends, its performance in the intended environment, the safeguards surrounding its use, and the agency’s rights after award. A higher-consequence use requires deeper testing, stronger independent review, more complete impact analysis and firmer provisions for intervention and recourse.
NIST’s voluntary AI Risk Management Framework reaches the same issue through its Map function: organizations should document the intended purpose, deployment context, affected people, expected benefits, possible harms, knowledge limits and the role of human oversight. The framework is not a certification and does not determine whether a procurement is lawful. Its value is that it forces evaluation to begin with context rather than a generic claim that a model is accurate or trustworthy.
A viable package is therefore organized around claims the agency must be able to defend, not around whatever files the vendor already happens to publish.
The agency must first document its own decision
The first evidence file should be written by the agency, not the vendor.
It should identify the exact task the system would perform, the people who would use it, the individuals or communities who could be affected, the existing process it would supplement or replace, the expected public benefit and the decisions that must remain with accountable officials. It should also identify prohibited uses and the conditions that would require additional approval.
This prevents a familiar procurement failure: evaluating a broadly capable AI product before deciding what the agency is actually buying it to do.
For federal executive agencies, OMB Memorandum M-25-21 defines certain uses as high-impact when AI output serves as a principal basis for decisions or actions with significant legal, material, binding, rights-related or safety consequences. The memorandum requires predeployment testing and a documented impact assessment for covered high-impact uses. That assessment must address intended purpose, expected benefit, data and model suitability, potential effects on privacy and civil rights, reassessment procedures, costs, independent review and signed risk acceptance.
Even where that federal definition does not apply, the underlying discipline remains useful. An acquisition record should explain why AI is being considered, what better performance would mean, how it will be measured against the current process and who has authority to accept the remaining risk.
Without that file, the evaluation has no stable target. A vendor can demonstrate speed, fluency or general benchmark performance while leaving unanswered whether the product improves the agency’s actual work.
The system must be identifiable
The next file should establish what the government is buying.
“Access to an advanced language model” is not a sufficient system description. The record should identify the provider, model family, version or release channel, hosting arrangement, integration layer, retrieval system, agency configuration, material system instructions, external tools, data connectors, subcontractors and any intermediary that can alter the model’s output.
This matters because the model named in a proposal may not be the system that agency employees encounter. A reseller may add moderation controls. An integrator may add retrieval, ranking or classification components. The agency may supply its own documents, prompts and access rules. Each layer can change both performance and risk.
OMB Memorandum M-26-04 recognizes this supply-chain problem in federal LLM procurement. It requires covered agencies to obtain information about models embedded in other products and notes that available information may depend on whether the contractor is the original developer, a reseller, an integrator or another intermediary. The memorandum establishes an initial transparency threshold that includes an acceptable-use policy, available model, system or data cards, end-user resources and a feedback mechanism. It also allows agencies to request additional information relevant to the proposed use.
Those materials are useful. They are not conclusive.
A model card can describe a model. It cannot, by itself, establish how an agency-specific configuration behaves after retrieval components, system instructions, access controls and third-party filters are added. The evidence package should therefore include a dated configuration record that connects the vendor’s general documentation to the exact service being evaluated.
It should also define what constitutes a material change. A new model version, revised system instruction, altered safety filter, changed retrieval source or new subcontractor may invalidate prior testing even when the product name remains unchanged.
Performance evidence must resemble the public task
The performance file is where procurement teams are most likely to mistake a score for an answer.
General benchmarks can reveal something about a model. They rarely establish fitness for a specific government workflow. The agency needs evidence tied to its intended task, representative inputs, operating constraints and failure consequences.
The evaluation record should preserve the tested system version, test dates, prompts or input procedures, test data, baseline process, scoring rules, reviewer instructions, sample sizes where applicable, uncertainty, known exclusions, adverse results and the environment in which the test occurred. It should report failure categories, not only an aggregate success rate.
For a document-synthesis system, that may require separate measurement of unsupported claims, omitted controlling provisions, incorrect citations, failure to distinguish current from superseded authority, and behavior when sources conflict. A single “accuracy” score would conceal the errors most likely to matter.
OMB M-25-22 directs federal agencies, to the greatest extent practicable, to test proposed AI solutions to understand their capabilities and limitations. For ongoing independent evaluation, it says agencies should use agency-defined data that resembles deployment data and is not available to the vendor. Where vendors conduct the testing, results should be detailed enough to be independently verified or reproduced when practicable.
That last condition is critical. A demonstration prepared by the seller can show that the product works under selected conditions. It cannot establish how often it fails outside those conditions.
NIST’s Generative AI Profile adds an important limitation: current predeployment testing may be inadequate, inconsistently applied or mismatched to the eventual deployment context. Predeployment evidence should therefore be treated as an estimate of likely behavior, followed by controlled piloting and continuing measurement—not as permanent proof of reliability.
Independent testing is not automatically superior. An agency can use an unrepresentative test set, select the wrong metric or misunderstand the domain. Independence reduces one conflict of interest; it does not repair a weak evaluation design. The package should identify who designed the test, who executed it, who reviewed it and what expertise each party brought to the work.
The real subject is authority, not accuracy
At this point, the procurement question changes.
A technically capable system may still be unacceptable because the agency has not defined what users may do with its output.
The evidence package must identify whether the system supplies information, drafts material, recommends an action, ranks cases, flags records or contributes to a decision. It should specify when human review is required, what information the reviewer receives, whether the reviewer has enough time and authority to disagree, and how disagreement is recorded.
A “human in the loop” label is not evidence. The agency needs a review procedure.
For federal high-impact uses, M-25-21 requires suitable human oversight and accountability, periodic operator training, monitoring for performance and adverse impacts, and—where appropriate—timely human review and an opportunity to appeal negative effects. It also requires monitoring for changes in the system, its data and its context of use.
The package should therefore include the operational decision map: what the system produces, who sees it, who may act, what independent information the human reviewer receives, what must be documented, and what recourse exists when the result is challenged.
Accessibility belongs in this file as well as in product testing. Section 508 requires federal agencies to ensure that covered information and communication technology they develop, procure, maintain or use is accessible to people with disabilities, subject to applicable exceptions. Simplified purchasing methods do not erase that obligation. An AI interface that cannot be operated with assistive technology, or that presents warnings and citations in inaccessible formats, has not supplied complete evidence of fitness for federal use.
The deeper point is simple: the same model output can be low-risk when used as a private drafting suggestion and dangerous when treated as an official finding. Procurement evidence must describe the authority attached to the output.
Risks must become contract rights
A risk register that does not affect the contract is an observation, not a control.
The evidence package should show how material risks were converted into enforceable terms. Depending on the use, those terms may govern agency data, retention, training on government inputs, intellectual-property rights, portability, subcontractors, testing access, incident reporting, update notification, audit support, records preservation, performance thresholds, remediation, suspension and termination.
M-25-22 directs federal agencies to address intellectual-property rights and lawful use of government data, reduce vendor lock-in, preserve the ability to test and monitor performance, require access needed for independent evaluation, and define protections for data and model portability. It also calls for notification of new AI features and encourages rollback when a new version fails performance standards.
The memorandum also states that contracts should support regular monitoring of performance, risk and effectiveness. It directs agencies to establish operational oversight, assess continuing value and consider sunset criteria when costs, needs, vendor requirements or model performance change.
This is where a minimum evidence package distinguishes a procurement from a trial account. The agency must know not only whether the system appears acceptable today, but whether it has the contractual ability to investigate, constrain or leave the system tomorrow.
That does not require indiscriminate demands for proprietary information. M-26-04 expressly advises agencies, where practicable, to avoid compelling disclosure of sensitive technical material such as model weights. The correct standard is not maximum disclosure. It is sufficient access to evaluate the relevant model-, system- and application-level controls.
Confidential annexes, controlled testing environments, independent assessors and narrowly tailored audit rights can protect legitimate trade secrets while preserving public accountability. “Proprietary” should shape the disclosure mechanism. It should not end the inquiry.
Evidence must remain valid after award
AI procurement cannot end with source selection because the object under contract may change while the contract remains in force.
The lifecycle file should specify how the agency learns about model substitutions, fine-tuning, new system instructions, revised moderation policies, new tools, changed data sources, altered pricing, security incidents and material performance degradation. It should also identify which changes trigger notice, regression testing, impact reassessment or renewed approval.
M-26-04 advises agencies to seek updated disclosures when providers integrate new AI features or components. M-25-22 requires continuing monitoring and contemplates periodic evaluation, rollback, sunsetting and data transfer at closeout.
NIST’s Generative AI Profile similarly recommends retaining histories of testing, evaluation, validation and verification; defining incident responsibilities; establishing third-party incident-response plans; and incorporating performance monitoring and stakeholder feedback into system updates.
The minimum package must therefore contain a change-control rule before the first change occurs.
A static packet can become misleading without containing a single false statement. The model card may still describe the old model. The benchmark may still report the old version. The security review may still cover the underlying cloud service. Yet the system being used by the agency may no longer be the system that was approved.
Evidence has an expiration condition even when it has no printed expiration date.
What should not satisfy the agency
A vendor’s responsible-AI principles should not satisfy the package unless those principles are connected to controls, test results and contractual duties.
A benchmark screenshot should not satisfy it without the benchmark version, model version, evaluation conditions, scoring method and relevance to the agency’s task.
A security authorization should not be treated as evidence that the model is accurate, fair, accessible or appropriate for a particular use. It addresses a different question.
A model card should not substitute for agency testing. A successful demonstration should not substitute for reproducible results. A pilot should not silently become a deployment. Human oversight should not be credited unless the human has the information, competence, time and authority needed to intervene.
GAO’s April 2026 review of 13 AI acquisitions across four federal agencies found that the selected agencies were not yet systematically collecting lessons learned. Officials said their policies did not require it, leaving agencies less prepared to reuse effective provisions concerning issues such as testing and data rights or to avoid earlier mistakes. The finding is a reminder that evidence must include institutional learning, not only product assessment.
GSA’s current AI strategy offers a more developed agency example. It describes use of agency-defined test plans, AI impact statements, real-world context testing, independent evaluation plans, data and model lineage, ongoing monitoring and annual re-registration for covered systems. That is an agency-specific implementation, not proof that every system has met its objectives, but it demonstrates how separate documents can be connected into a lifecycle record.
A decision rule procurement teams can use
An agency has reached the minimum viable evidence threshold when a qualified reviewer who was not dependent on the vendor’s sales presentation can reconstruct the decision from the written record.
The reviewer should be able to identify the exact system and configuration, the public task it is intended to support, the data and external components on which it depends, the conditions under which it was tested, the important failures that were observed, the people who retain authority, the remedies available to affected individuals, the contract rights available to the agency, and the events that require retesting or withdrawal.
If the answer to one of those questions exists only in a meeting, a demonstration or an assurance that the vendor can provide more detail later, the package is not complete.
Smaller agencies may need proportionate methods. They can reuse government-wide resources, collaborate on evaluation infrastructure, rely on carefully governed pilots, narrow the use case or require a vendor to support testing rather than building a large internal laboratory. What they cannot responsibly do is replace evidence with optimism because their evaluation capacity is limited.
Federal agencies themselves report that evolving policy, constrained budgets and insufficient technical resources complicate generative-AI adoption. Those constraints are real. They argue for a smaller and sharper evidence package, not for abandoning the essential questions.
A minimum viable evidence package does not prove that an AI system is safe, unbiased or effective everywhere. It establishes something more modest and more defensible: that the agency knows what it is acquiring, has tested the claims that matter to the intended public use, has preserved human and institutional authority, and can respond when the evidence changes.
The minimum is reached when the government can explain why it bought this system, for this task, under these conditions—and can detect when those conditions no longer hold.
Sources and verification
Office of Management and Budget, Memorandum M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government, verified for acquisition planning, testing, documentation, contract rights, monitoring, portability, rollback and closeout provisions.
Office of Management and Budget, Memorandum M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, verified for high-impact determinations, predeployment testing, impact assessments, independent review, monitoring, human oversight and appeals.
Office of Management and Budget, Memorandum M-26-04, Increasing Public Trust in Artificial Intelligence Through Unbiased AI Principles, verified for federal LLM procurement scope, minimum transparency materials, enhanced disclosures, intermediary-provider considerations and updated disclosures after material changes.
National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework 1.0, verified for intended-use documentation, contextual risk mapping, knowledge limits, human oversight, third-party risk and independent review.
National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, verified for third-party due diligence, predeployment testing limitations, incident response, test-record retention and continuing monitoring.
U.S. Government Accountability Office, Artificial Intelligence: An Accountability Framework for Federal Agencies and Other Entities, verified for the governance, data, performance and monitoring structure used to assess AI accountability.
U.S. Government Accountability Office, Artificial Intelligence Acquisitions: Agencies Should Collect and Apply Lessons Learned to Improve Future Procurements, verified for current federal acquisition practices, documented implementation challenges and the finding concerning systematic collection of lessons learned.
U.S. General Services Administration, AI Strategies and Compliance Plan, verified as an agency-specific example of test plans, impact statements, real-world evaluation, lineage documentation, independent review and continuing oversight.
Section508.gov, IT Accessibility Laws and Policies and Understanding ICT Micro-Purchases, verified for the application of federal ICT accessibility requirements to procurement and micro-purchases.
TAIRC, Research Portfolio and Research Categories and Topics, verified for TAIRC’s evaluation-centered program scope, organizational boundaries and prohibition on certification or endorsement claims.



