Adversarial Evaluation Needs Rules of Engagement
A red team can make an AI system fail in minutes. An institution may then spend weeks discovering that the failure proves almost nothing.
Was the tested model the same version the agency planned to use? Were retrieval, tools, memory, and safety controls enabled? Was the tester authorized to probe connected systems? Did the exercise use production data? Could another evaluator reproduce the result? What decision was the test supposed to inform?
Without answers, a dramatic failure is an anecdote. A clean run is even more dangerous: it can be mistaken for evidence that the system is safe.
Adversarial evaluation therefore needs rules of engagement established before testing begins. These rules should define the decision the exercise will inform, the exact system under examination, the relevant threat model, the authority granted to testers, the permitted and prohibited methods, the evidence that must be retained, the conditions that require a halt or escalation, and the process for remediation, retesting, and disclosure.
The purpose is not to make adversarial testing polite. It is to make its findings interpretable.
Red teaming is still an unsettled practice
The term “AI red teaming” covers several different activities. It may describe humans trying to elicit prohibited behavior, experts examining a narrow domain risk, automated systems generating adversarial inputs, community participants testing culturally specific harms, or mixed campaigns that turn exploratory findings into repeatable evaluations.
NIST describes AI red teaming as an evolving practice, usually conducted in a controlled environment to identify adverse behavior, understand how it can occur, and stress-test safeguards. The agency also cautions that the quality of the findings depends on the expertise and composition of the red team, and that results require further analysis before they are used in governance or risk-management decisions.
That caution matters because “red teamed” can sound far more conclusive than it is.
Michael Feffer and colleagues examined publicly documented generative-AI red-teaming practices and found wide variation in their purposes, targets, threat models, resources, methods, reporting practices, and resulting decisions. Their 2024 AIES paper argued that treating an underspecified red-teaming exercise as a universal answer to AI risk can become security theater.
A rules-of-engagement document addresses that ambiguity at the point where it is still controllable: before the first adversarial interaction.
Start with the decision, not the attack
The first question should not be, “How do we break the model?”
It should be, “What institutional decision will this evidence support?”
An agency considering a limited pilot may need to know whether a system preserves a stated information boundary under plausible misuse. A procurement team may need evidence about whether a vendor’s safeguard claim survives independent testing. A research team may be exploring unfamiliar failure modes so it can design a later benchmark. Those are different purposes. They require different testers, access levels, evidence standards, stopping rules, and conclusions.
A decision-bound evaluation names the intended use of the evidence in advance. It might support a decision to continue testing, impose a deployment restriction, require remediation, reject a vendor claim, or convert a confirmed failure into a regression test. It should also identify who owns that decision. An exercise without a named decision owner can generate a compelling report that nobody is responsible for acting on.
The rules must then identify the artifact being tested. “The model” is rarely specific enough. The relevant unit may include a model version, system instructions, retrieval sources, content filters, user permissions, memory, tool access, interface design, and surrounding application logic. Feffer and colleagues place the artifact, version, guardrails, lifecycle stage, threat model, success criteria, team composition, access, resources, and post-test accountability among the questions that should be resolved around an AI red-teaming exercise.
Where a provider does not disclose a component, the evaluator should record the absence as a limitation. Missing information should never be silently converted into a methodological assumption.
Red teaming is not field testing
NIST’s Assessing Risks and Impacts of AI pilot offers a useful distinction. ARIA 0.1 separated model testing, red teaming, and field testing into different protocols. Model testing used predefined prompts. Red teaming deliberately tried to make applications violate scenario-specific guardrails. Field testing examined interactions intended to resemble realistic use.
In the pilot, 51 red teamers participated between December 2024 and January 2025. They received the exercise goals, definitions and examples of prohibited information, randomized scenarios, and post-task questionnaires. NIST explicitly stated that this red teaming was not designed to mimic real-world use; its purpose was to test whether applications would protect specified information under adversarial pressure.
That boundary prevents a common analytical error. A successful adversarial attack shows that a failure was reachable under the tested conditions. It does not estimate how frequently ordinary users will encounter the failure. A failed attack shows that the evaluators did not reach the targeted behavior through the search process they used. It does not prove that the behavior is unreachable.
When the institutional question concerns prevalence, user behavior, or real-world consequences, adversarial evaluation may be the wrong primary method. Field studies, structured user testing, incident analysis, conventional benchmarks, or other forms of evaluation may be needed.
Borrow the discipline of cybersecurity, not its assumptions
“Rules of engagement” has an established meaning in security testing. NIST Special Publication 800-115, issued in 2008, describes an ROE as the detailed constraints and permissions established before a security test. Its template covers purpose, scope, assumptions, risks, personnel, schedules, authorized locations and access, communications, incident handling, permitted technical activity, data handling, reporting, and accountable signatures.
That publication predates modern generative AI and should not be treated as an AI-specific standard. Its durable contribution is structural: an adversarial test should have explicit authorization, bounded targets, known escalation paths, controlled evidence, and accountable owners.
An AI-specific ROE must go further. It should identify the model or application version, date of access, provider and interface, available configuration, relevant system instructions when known, guardrails, connected tools, retrieval sources, memory state, user role, language, and sampling settings. It should state whether evaluators have black-box, gray-box, or deeper internal access and whether their environment matches the system an institution is actually considering.
The document should also define which systems and data are outside scope. Testers should not be left to infer whether they may interact with production services, submit personal information, access third-party tools, create accounts, change permissions, retain outputs, or exceed normal rate limits. Authorization should be precise enough that a creative evaluator can operate aggressively inside the boundary without improvising the boundary itself.
This is where the cybersecurity analogy reaches its limit. A language-model red team may investigate discriminatory outputs, dangerous assistance, false authority, cultural failure, privacy leakage, tool misuse, or interactions among several system components. The relevant harm may be contextual rather than a conventional software vulnerability. Rules written only by a security team can therefore miss the people, institutions, and social conditions that determine whether an output is harmful.
The evaluator is part of the instrument
Adversarial evaluation is shaped by who performs it.
NIST’s Generative AI Profile states that red-team quality is related to the team’s background and expertise. It recommends domain knowledge and attention to the social and cultural context in which a system may be used.
The threat model should determine the team. Testing a public-information assistant may require expertise in administrative processes, accessibility, source traceability, and the experiences of people who use the relevant service. Testing a scientific assistant requires evaluators capable of judging the domain content, not merely producing unusual prompts. A technically skilled generalist may discover interface weaknesses while missing a subtle but consequential domain error.
Independence also needs definition. An external evaluator may reduce some conflicts of interest, but external status does not guarantee methodological independence. Funding relationships, publication incentives, access restrictions, nondisclosure terms, and reliance on information supplied by the developer can all influence the work. The ROE should disclose who selected the evaluators, who pays them, what access they receive, who controls publication, and how disagreements are adjudicated.
Tester welfare belongs in the same document. A vendor-authored OpenAI paper on external red teaming identifies potential psychological harm from sustained exposure to harmful material, along with information hazards created when evaluators discover techniques that could enable misuse. It recommends informed consent, fair compensation, appropriate support, controlled access, and responsible disclosure. Because the paper describes one developer’s practices, it should be treated as operational evidence rather than an independent standard. The risks it identifies are still directly relevant to campaign design.
NIST’s ARIA pilot had its model-testing, red-teaming, and field-testing activities reviewed as separate protocols by the agency’s Research Protections Office. That does not mean every organizational red team is human-subjects research. It shows that participant oversight should be considered deliberately rather than dismissed because the activity is called testing.
Evidence needs a chain of custody
An adversarial finding should preserve enough context for an independent reviewer to understand what happened and, where lawful and safe, attempt to reproduce it.
That record ordinarily includes the tested system and version, date and environment, evaluator access, relevant configuration, complete interaction history, expected behavior, observed behavior, applicable policy or requirement, reproduction attempts, adjudication method, severity rationale, affected stakeholders, and known limitations. When exact prompts or outputs create security, privacy, or safety risks, the record can use restricted annexes, hashes, controlled repositories, or sanitized examples. The public summary and the underlying evidence do not need identical access rules.
A screenshot of a harmful answer is weak evidence by itself. It may omit preceding turns, system configuration, model version, tool state, or retries that changed the outcome. Conversely, a large count of attempted attacks can create false precision when success criteria, duplicate prompts, evaluator behavior, and sampling conditions are unclear.
The scoring rule should be written before results are known. If a failure requires human judgment, the ROE should specify evaluator training, decision criteria, disagreement handling, and whether adjudicators are blinded to information that could bias the result. “Concerning,” “unsafe,” and “successful attack” are conclusions, not measurement definitions.
NIST recommends formal policies, procedures, standardized measurement protocols, and structured feedback exercises for generative-AI risk measurement. It also places red teaming within a wider lifecycle of oversight, documentation, incident identification, and information sharing.
The essential shift is from an impressive demonstration to an inspectable record.
Rules do not weaken adversarial creativity
A common objection is that detailed rules will constrain testers and make the exercise predictable. That can happen when the ROE dictates every prompt or rewards compliance with a narrow script.
The answer is not to abandon rules. It is to govern the right things.
The ROE should fix the decision, target, authority, safety boundary, evidence standard, and escalation process. Within that envelope, testers can be given broad freedom to explore. Open-ended discovery and procedural discipline are compatible. The former searches for what the institution has not anticipated. The latter preserves enough context to determine what the discovery means.
The deeper trade-off is therefore not creativity versus control. It is inspectability versus unrecorded judgment.
An expert can uncover a real failure and still leave the institution with no defensible basis for estimating its significance. A technically inventive campaign can be institutionally useless when it does not state what was tested, what remained untested, or how the finding changes a decision.
Rules of engagement constrain the claims that can be made from the exercise. They need not constrain the imagination used inside it.
Stopping is part of the test
Adversarial work creates its own risks. Testers may encounter personal information, discover an unreported vulnerability, trigger a connected tool, generate disturbing material, affect a third party, or find that the controlled environment is less isolated than expected.
The time to decide what happens next is before that moment.
The ROE should identify the conditions requiring an immediate halt, temporary pause, restricted evidence channel, security response, ethics review, or senior escalation. It should name the person with authority to stop the exercise and the people who must be contacted. It should also state whether testers may continue exploring a severe finding after initial confirmation or must preserve the evidence and wait for authorization.
NIST’s cybersecurity ROE template requires risks, mitigation measures, incident-response contacts, detailed data-handling requirements, reporting expectations, and signatures from accountable parties. Those controls translate well to AI evaluation when they are adapted to the relevant system and harms.
A test plan without a stop rule quietly delegates institutional risk to whichever evaluator happens to be at the keyboard.
A red-team result is not a safety verdict
Adversarial evaluation can reveal reachable failures, weak assumptions, and promising targets for repeatable testing. It cannot establish that a system is universally safe.
NIST’s 2025 adversarial-machine-learning taxonomy discusses theoretical and empirical limits on mitigations. It concludes that adversarial testing remains useful because it can close attack paths or raise the effort required to exploit them, while warning that organizations—especially those considering high-stakes uses—need measures beyond adversarial testing.
The same limit applies to an apparently successful campaign. A red team’s coverage is bounded by its threat model, expertise, time, access, languages, tools, and imagination. Broader participation may reveal additional failures, but no finite exercise covers every context or future system state.
Results also expire. Models, policies, retrieval collections, interfaces, filters, and connected tools change. OpenAI’s external-red-teaming paper describes findings as point-in-time evidence and recommends turning useful discoveries into repeatable evaluations that can be rerun as systems evolve.
A confirmed adversarial failure should therefore become the beginning of a control cycle. The institution should analyze the cause, decide whether remediation is required, convert the finding into a stable test where possible, verify the mitigation, and rerun the test after material changes. A resolved prompt example is not proof that the broader failure class has disappeared.
A small institution can still produce credible evidence
Frontier-scale red teaming can require extensive access, specialized expertise, substantial compute, and large participant pools. A public-interest organization does not need to imitate that scale to contribute responsibly.
It does need to narrow its claim.
A credible small-institution campaign can focus on one defined system, one intended context, one or two consequential risk hypotheses, and a decision that can actually change. The team can combine a domain specialist, an evaluator familiar with the technology, an independent adjudicator, and an accountable institutional owner. Testing can occur in a controlled environment with synthetic, licensed, or otherwise authorized inputs. Exploratory human work can identify failure patterns; limited automation can then test reproducibility and variation without pretending to cover the entire attack surface.
The resulting report should state exactly what was tested, what access and resources were available, what evidence was found, what remained outside scope, and what decision followed. Sensitive exploit details can be restricted while the public account preserves methodology, limitations, and institutional response.
This approach produces less spectacle. It produces more usable evidence.
TAIRC’s published research portfolio defines its LLM Safety, Evaluation & Reliability program as evaluation-centered and research-bound. It excludes model certification, production deployment, performance optimization, and claims that a model is universally safe. The program’s intended contribution is reproducible methodology and transparent analysis, not operational red teaming or the publication of harmful exploit instructions. This article follows that boundary and does not report a completed TAIRC adversarial evaluation.
Institutions should also treat the framework here as governance guidance rather than legal advice. Authorization, privacy, records retention, contracting, research oversight, employment, cybersecurity, and disclosure obligations should be reviewed for the applicable jurisdiction and setting before testing begins.
The test begins before the first prompt
An adversarial evaluation is ready to start only when the institution can identify the decision at stake, the precise system and version, the threat being examined, the authority granted, the evidence required, the conditions that stop the work, and the person responsible for acting on the result.
If those answers do not exist, the organization is not ready to red-team the system. It is ready to generate stories about it.
The red team finds the failure. The rules of engagement determine whether the institution learns anything from it.
Sources and verification
The National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024. Used for NIST’s description of AI red teaming, evaluator expertise, contextual diversity, governance analysis, standardized measurement, and lifecycle oversight.
The National Institute of Standards and Technology, Assessing Risks and Impacts of AI: Pilot Evaluation Report, NIST AI 700-2, November 2025. Used for the ARIA 0.1 distinction among model testing, red teaming, and field testing; its protocol review; scenario design; participant count; and the stated limits of adversarial testing as a proxy for real-world use.
The National Institute of Standards and Technology, Technical Guide to Information Security Testing and Assessment, Special Publication 800-115, September 2008. Used for the established cybersecurity concept of rules of engagement and its treatment of scope, authorization, personnel, risks, communications, incident handling, data handling, reporting, and accountable signatures.
The National Institute of Standards and Technology, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, March 2025. Used for the limits of adversarial mitigations and the need to combine adversarial testing with broader security and risk-management measures.
Michael Feffer, Anusha Sinha, Wesley H. Deng, Zachary C. Lipton, and Hoda Heidari, Red-Teaming for Generative AI: Silver Bullet or Security Theater?, AIES 2024. Used for the documented variation in red-teaming purposes, artifacts, threat models, methods, resources, reporting practices, and resulting decisions, as well as the authors’ proposed pre-activity, during-activity, and post-activity questions.
OpenAI, OpenAI’s Approach to External Red Teaming for AI Models and Systems, 2025. Used as vendor-authored practice evidence concerning campaign disclosure, point-in-time validity, resource constraints, participant welfare, information hazards, responsible disclosure, and conversion of exploratory findings into repeatable evaluations.
The AI Research Center, TAIRC Research Portfolio and TAIRC Research — Categories & Topics. Used only to verify TAIRC’s stated program scope, intended public benefit, evaluation-centered mandate, and explicit organizational boundaries.



