SignorCrypto note · AI
NIST ARIA: A Practical Guide to AI Evaluations
How model testing, red teaming and user testing build a fuller evidence base

NIST’s ARIA Evaluation Planning Manual, published on September 18, 2026, gives teams a practical starting point for evaluating an AI application. Its core idea is to combine model testing, red teaming and user testing instead of treating one benchmark as proof that a system is trustworthy. The manual is a planning resource, not a universal scorecard: teams must adapt it to the application, users, risks and decision involved.
What NIST ARIA is
ARIA means Assessing Risks and Impacts of AI. The new ARIA Evaluation Planning Manual is NIST Trustworthy and Responsible AI report 200-3. NIST describes it as a first step for developing customised evaluations of AI applications.
The word application matters. The object being evaluated is not only a foundation model, but the model together with instructions, data, interface, safeguards and human use. A system can pass a model test and still fail when users misunderstand it, an adversarial input bypasses a safeguard or a workflow turns an uncertain answer into an irreversible action.
NIST’s AI Risk Management Framework is intended for voluntary use and supports trustworthiness in the design, development, use and evaluation of AI systems. ARIA fits the measurement part of that lifecycle; it is not a safety certification.
The three testing layers
Model testing
Model testing uses defined inputs and measures outputs against a construct or requirement, such as accuracy, robustness, refusal behaviour or relevance. Define the construct before collecting results. “The assistant is good” is not testable; “the assistant extracts invoice fields correctly and cites the approved policy” is closer.
A useful plan records the application and model version, represented task and population, test set, scoring rule, decision threshold and limitations. A benchmark reveals capability under selected conditions, not every way the application will be used.
Red teaming
Red teaming deliberately stresses an application to elicit negative or policy-violating behaviour. It probes adversarial prompts, conflicting instructions, restricted-data requests, misleading context and attempts to bypass safeguards.
Define the risk, attack surface, tester instructions, evidence and failure severity. For an agent with tools, include the surrounding permissions: a refusal in a chat window does not prove that the connected system will refuse before calling a sensitive tool. Classify failures and repeat the tests after model, retrieval, tool or policy changes.
User testing
User testing adds evidence from realistic interaction. It can expose unclear explanations, over-trust, excessive effort, inaccessible design, poor recovery and a mismatch between output and the user’s real decision.
NIST’s earlier ARIA 0.1 pilot report describes model testing, red teaming and field testing across seven AI applications submitted by five organisations. The pilot used dialogue annotation and tester questionnaires. The terminology differs slightly from the 2026 manual’s “user testing,” but the lesson is consistent: human interaction is evidence about impact, not merely a usability afterthought.
Positive feedback is not proof of correctness. Users may like a fast answer that is wrong. Combine user perception with task outcomes and expert review.
A practical ARIA evaluation plan
Start with the decision the evaluation must support: release, procurement, model change, workflow expansion or risk review. Then work backward from the evidence required.
- Define the boundary. Name the model, interface, data, tools, user groups, environment and version. Record what the system may do and what remains under human control.
- Write observable objectives. Ask whether the assistant works on representative cases, escalates when evidence is missing, resists restricted-data attacks and helps the target user without creating new risk.
- Choose complementary methods. Use model testing, red teaming and user testing deliberately. Document what is included and outside scope.
- Set measures in advance. Specify expected outcomes, scoring, annotation guidance, severity categories, sampling, reviewer qualifications and retention rules.
- Preserve disagreement. Automated scoring may call an output correct while users find it unusable; red teamers may uncover a rare but severe failure. Do not hide those differences in one average.
- Report limits and next actions. State what was tested, what was not, the conditions, sample size, failures and remediation plan. Separate measured results from interpretation and uncertainty.
For teams building an internal evaluation backlog, the SignorCrypto Toolkit can be a complementary resource when it directly fits the workflow.
Why ARIA is broader than a benchmark
NIST’s related TEVV-Athlon framework, released as an initial public draft in August 2026, proposes a four-stage method for customised Test, Evaluation, Verification and Validation assessments across language, multimodal and agentic systems. NIST’s AITE programme uses blind data and a sequestered testbed to reduce train/test contamination.
Together, these initiatives point to a more useful standard for AI assurance: explicit objectives, controlled conditions, multiple evidence sources and clear reporting. The more a system affects people or connected systems, the less a leaderboard number can establish on its own.
Verified facts and SignorCrypto analysis
Verified fact: NIST published the ARIA Evaluation Planning Manual on September 18, 2026. Its abstract identifies model testing, red teaming and user testing as the three testing types used to assess trustworthiness.
Verified fact: The ARIA 0.1 pilot involved five organisations and seven AI applications, using model testing, red teaming and field testing with dialogue annotation and tester questionnaires.
SignorCrypto analysis: ARIA’s value is not a new universal metric. It is a disciplined way to connect technical performance, adversarial behaviour and human use before a deployment decision.
Uncertainty: The manual does not replace domain expertise, representative data, privacy safeguards or independent review. Results still depend on the application and test design.
FAQ
Is NIST ARIA a compliance requirement?
No. The NIST AI Risk Management Framework is voluntary, and ARIA is a planning approach. Separate legal, contractual or sector-specific obligations may still apply.
Is model testing enough?
Usually not. It measures defined behaviours under defined conditions, but does not show how the system responds to attacks or how people interpret it in context.
What is the difference between red teaming and user testing?
Red teaming searches intentionally for harmful or policy-violating behaviour. User testing examines realistic interaction, task performance and user impact. They answer different questions.
Sources
- NIST — ARIA Evaluation Planning Manual
- NIST — AI Risk Management Framework
- NIST — ARIA Pilot Evaluation Report
- NIST — TEVV-Athlon Framework
- NIST — AITE
If your team needs to turn AI evaluation evidence into a governed workplace assistant and repeatable business workflow, explore Botchi.