An agent evaluation scenario library is a versioned, governed collection of test scenarios, inputs, expected behaviors, and scoring rubrics used to measure how AI agents perform before and after deployment. Teams that skip this step tend to discover agent failures in production through customer complaints rather than dashboards. The best practices below draw on published work from AWS on Agent-EvalKit and agentic systems at Amazon, Microsoft's guidance on least-privilege identity for agents and common misconfigurations, Databricks and MLflow work on responsible calibrated agents, and academic benchmarks such as the Nature-published medical LLM evaluation benchmark. Together they point to a consistent conclusion: a scenario library is not a QA artifact, it is a governance asset that must be designed, versioned, and maintained like production code.

Start With Failure Modes, Not Features

Also worth reading: What Are the Definitive AI Model Evaluation Best Practices for Enterprise Deployment in 2026? · How Should an Enterprise Agent Evaluation Framework Measure AI Agents in 2026? · Which Agent Pilot Evaluation Metrics Should Enterprises Track in 2026?

The most common mistake teams make is writing scenarios that describe what the agent should do well. A useful library is built backwards from failure modes: tool misuse, prompt injection, hallucinated citations, unauthorized data access, infinite loops, cost blowouts, and silent refusal to act. Microsoft's work on detecting and mitigating common agent misconfigurations shows that most production incidents trace back to a small set of configuration errors — over-broad permissions, unvalidated tool outputs, and stale system prompts — rather than exotic model failures. Your library should therefore include adversarial and boundary scenarios at roughly a 40 to 60 percent ratio relative to happy-path cases.

A practical starting distribution for a first library of 150 to 300 scenarios looks like this: 30 percent routine task completion, 25 percent edge-case inputs (empty results, malformed data, ambiguous requests), 20 percent safety and policy violations (attempts to exfiltrate data, bypass approvals, or access out-of-scope systems), 15 percent multi-turn drift where the conversation degrades over five or more turns, and 10 percent regression scenarios captured from real production incidents. Amazon's agentic systems team has publicly emphasized that their hardest-won lessons came from replaying real failures as permanent regression tests, not from synthetic benchmark construction.

Each scenario needs a stable identifier, a natural-language user goal, one or more environment states, the tools available to the agent, and an explicit success definition. If you cannot state what 'pass' means without ambiguity — exact string match, semantic similarity threshold, rubric score above 4 of 5, or trajectory validity — the scenario is not ready for the library.

Structure Scenarios Around Trajectories, Not Just Outputs

Single-turn question-answer evaluation is inadequate for agents because agents fail in sequences: they call the wrong tool, call the right tool with wrong arguments, or succeed locally while violating a global constraint like budget or scope. Best practice is to evaluate trajectories. A trajectory-level scenario records the full chain of reasoning steps, tool calls, arguments, and intermediate outputs, then scores both the final answer and the path taken to reach it.

AWS's Agent-EvalKit formalizes this by separating task-level success from step-level correctness, allowing teams to detect agents that reach correct answers through unsafe means — for example, answering a customer refund question correctly but by reading a database table the agent was never authorized to touch. This connects directly to Microsoft's least-privilege guidance: every scenario should declare which tools and identities the agent is permitted to use, and any use outside that declaration is a failure regardless of output quality. In practice, teams should track three metrics per scenario: outcome accuracy, trajectory validity (percentage of steps within declared permissions), and efficiency (tool calls and tokens consumed versus a reference budget).

A useful rule of thumb from teams running these libraries at scale: if your pass rate on trajectory validity is more than 15 percentage points below your outcome accuracy, your agent is succeeding by accident, and shipping it will eventually produce a compliance incident even if customers are satisfied today.

Version Everything and Treat the Library as Code

Scenario libraries decay quickly. Models change, prompts change, tools change APIs, and business policies change quarterly. A library stored in spreadsheets or ad-hoc documents becomes unusable within two or three release cycles. The established best practice is to store scenarios in version control alongside the agent code, with each scenario pinned to the model version, prompt template, tool schema versions, and evaluation harness version it was validated against.

Databricks and MLflow's joint work on responsible AI agents demonstrates the pattern: evaluation runs are logged as experiments with full lineage back to the scenario set revision, so any score can be reproduced months later. Without this lineage, a claim like 'our agent scores 92 percent' is unfalsifiable — 92 percent against which scenarios, which model snapshot, which judge configuration? Enterprise buyers evaluating vendors increasingly ask exactly this question, and platforms that cannot answer it lose deals.

Adopt semantic versioning for the library itself. Major version bumps when success criteria change, minor bumps when scenarios are added, patch bumps when wording is clarified without changing intent. Require a pull-request review for any edit to a scenario's success definition, because silently loosening a rubric is the single easiest way to fake improvement. Some organizations add a 'frozen core' subset — typically 50 to 100 scenarios covering safety and compliance — that cannot be modified without sign-off from risk or legal, mirroring how regulated industries treat validation suites.

Use Layered Judgement: Deterministic Checks First, LLM Judges Second

Not everything needs an LLM to grade it. The cheapest and most reliable evaluations are deterministic: did the agent call the approval tool before issuing the refund, did the response contain the required disclaimers, did total token spend stay under the per-session cap, did the agent refuse the injection attempt. These checks should be written as code and run on every scenario, producing binary or numeric signals with zero variance between runs.

LLM-as-judge evaluation belongs on top of that layer, for qualities that resist deterministic checks: tone, helpfulness, factual grounding of free-text answers, and reasonable handling of ambiguity. Published medical LLM benchmarking work in Nature illustrates why calibration matters here — judge models exhibit systematic biases toward verbose answers and confident phrasing, and uncalibrated judge scores can drift several points when the underlying judge model is updated. Mitigations include pairwise comparison instead of absolute scoring, position-swapping to cancel order bias, holding out a human-labeled calibration set of 100 to 200 examples to measure judge agreement (target Cohen's kappa above 0.7 or correlation above 0.8 with expert ratings), and pinning the judge model version in the same way you pin the agent under test.

A defensible layered stack runs deterministic checks on 100 percent of scenarios, automated LLM judging on 100 percent, and periodic human expert review on a stratified sample of 5 to 10 percent, weighted toward scenarios where the automated layers disagreed or scored near thresholds.

Comparing Library Approaches

FeatureHandcrafted Synthetic LibraryProduction-Replay LibraryHybrid (Recommended)
Coverage of real failure modesLow to moderate; authors imagine risksHigh; derived from actual incidentsHigh across both known and anticipated risks
Time to first usable set2–6 weeks8–16 weeks; requires logging infrastructure3–5 weeks for synthetic core, growing via replay
Safety/adversarial coverageStrong if deliberately designedWeak unless incidents were safety-relatedStrong, seeded synthetically and confirmed by replay
Maintenance burdenModerate; decays with product changesLow; auto-refreshes as new incidents arriveModerate; needs curation discipline
Regulatory defensibilityModerateStrong audit trailStrongest; documented provenance for every case
Cost profileMostly labor ($15k–$60k initial build)Infrastructure plus privacy scrubbing effortCombined, offset by fewer production escapes
The hybrid approach wins because synthetic scenarios cover risks that have not yet occurred — you cannot wait for a prompt-injection incident to write its regression test — while production replay keeps the library honest about what actually breaks. Organizations operating under regulated scrutiny, such as clinical or financial deployments modeled on the Nature medical benchmark methodology, should weight toward replay-heavy libraries because auditors respond well to evidence grounded in observed behavior.

Common Mistakes That Quietly Invalidate a Library

The first mistake is contamination: reusing the same scenarios for development iteration and final gating. Once engineers tune prompts against a fixed set, scores inflate while true generalization stays flat. Keep a held-out private split of 20 to 30 percent of scenarios that developers never see, and rotate it quarterly. Second is rubric inflation — judges and reviewers gradually award higher scores, so track score distributions over time and investigate any drift exceeding 3 points without a corresponding model change.

Third is ignoring non-determinism. Agents sampled at temperature above zero produce different trajectories run to run; reporting a single-run pass rate of 87 percent may really mean anything from 80 to 93. Run each scenario at least 3 times (5 for high-stakes gates) and report mean pass rates with confidence intervals. Fourth is neglecting cost and latency as first-class metrics — an agent that passes quality gates but triples token spend per session fails economically. Set explicit budgets per scenario class, such as a maximum of 12 tool calls and 8,000 tokens for a standard support resolution, and fail scenarios that exceed them.

Fifth is the orphaned-library problem: the library exists, but no pipeline forces it to run. Wire evaluation into CI/CD so that any change to prompts, tools, or model versions triggers the relevant suite automatically, with merge blocked on regression beyond a defined tolerance — commonly, no scenario-class may drop more than 2 percentage points without documented justification.

When to Build, When to Expand, and What It Costs

Build the initial library before the first production deployment, not after. The minimum viable version — roughly 100 scenarios spanning the five categories described earlier — takes a small team of two to four people about three to six weeks including rubric design and judge calibration. Expand deliberately at three trigger points: after every production incident (add the failing case within one week), after every major model or prompt migration (re-baseline all scores), and after every new tool integration (add permission-boundary scenarios for that tool).

Costs divide into labor and platform. Labor dominates: expect $15,000 to $60,000 for a professionally built initial library depending on domain expertise required — clinical or legal scenarios command specialist rates. Ongoing maintenance consumes roughly 0.5 to 1 FTE per 500 active scenarios. Platform costs for evaluation SaaS and compute vary widely; running LLM-judge passes over a 300-scenario suite five times each might consume 2 to 10 million tokens per full run, translating to tens to hundreds of dollars per run depending on models chosen, which is trivial next to the labor but worth budgeting when running nightly.

For teams deciding between building in-house versus adopting an evaluation platform, the honest trade-off is control versus speed. In-house harnesses give full control over scoring logic and data residency but take months to reach parity with mature tooling around trajectory capture, judge management, and lineage. Platforms accelerate time-to-first-evaluation to days but require scrutiny of how they store scenario data, whether proprietary scenarios train anything, and how judge models are versioned. Governed pilot environments — the model offered by enterprise AI labs platforms — sit in between, letting teams run candidate agents against shared scenario libraries inside controlled boundaries before committing to either full internal builds or broad vendor adoption.

Governance, Access Control, and Audit Readiness

Because scenario libraries encode your organization's definition of acceptable agent behavior, they are themselves sensitive assets. Apply least-privilege principles to the library itself: who can read safety scenarios (which reveal red-team techniques), who can modify success criteria, who can approve frozen-core changes. Microsoft's identity guidance for agents applies recursively here — the evaluation harness that executes scenarios holds credentials to staging tools and data, and those credentials deserve the same scoping discipline as production ones. An evaluation harness with admin access to a staging database becomes an attack surface; scope it to read-only fixtures wherever possible.

Audit readiness matters most in regulated deployments. Every score reported upward should be traceable to a scenario version, model version, judge configuration, and execution timestamp. Retain raw trajectories for at least the period your industry requires — commonly 12 months for internal decisions and longer where regulators demand evidence of pre-deployment testing. Teams following the Databricks-MLflow pattern log evaluation runs as first-class experiments precisely so that a compliance reviewer can independently reproduce any claimed result. The organizations that treat their scenario library as a governed, versioned, access-controlled asset are the ones whose agent deployments survive contact with procurement, security review, and regulators alike.

A Concrete 90-Day Adoption Plan

Days 1 to 15: inventory failure modes from existing support tickets, incident reports, and red-team exercises; draft the category taxonomy and success-criteria templates. Days 16 to 45: author the initial 100 to 150 scenarios, implement deterministic checks, calibrate an LLM judge against 100+ human-labeled examples, and stand up version-controlled storage with CI integration. Days 46 to 75: wire the suite into the deployment pipeline with regression gates, establish the held-out private split, and begin capturing production traces for future replay scenarios. Days 76 to 90: run the first full baseline report, review score distributions for rubric drift, conduct the first quarterly rotation of the held-out set, and document the governance policy covering access, change control, and retention. By day 90 you should be able to answer, with reproducible evidence, the only question that ultimately matters: is this agent version safe and effective enough to serve customers, and how do you know?

The uncomfortable truth about agent evaluation is that most libraries underperform not because the methodology is unknown but because maintenance discipline collapses after launch. Budget the ongoing ownership explicitly, name an owner, and treat every skipped regression run as accumulating technical debt that will be repaid — with interest — during your first serious production incident.