The Direct Answer: Treat Agent Security as a System Test

An effective agent security evaluation tests the complete system that can take action, not merely the model that generates text. In practical terms, that system includes prompts, retrieved documents, tools, credentials, memory, orchestration code, external services, human approval gates, and the environment in which the agent operates. A model may behave acceptably in a controlled chat test and still create unacceptable risk when connected to a shell, browser, email account, code repository, or production API. The correct unit of evaluation is therefore the deployed agent configuration and its permitted actions.

Also worth reading: How Should Enterprises Evaluate AI Models Safely in 2026 Without Compromising Security or Innovation? · What Is an Agentic AI Security Scoping Matrix and How Do Enterprises Build One in 2026? · Which Metrics Should Enterprises Use to Evaluate AI Agent Pilots Before Production?

Enterprises should begin with a defined abuse case, an isolated test environment, explicit success thresholds, and reproducible evidence. By September 2026, the evaluation market is also being shaped by documented incidents in which coding or research agents interacted with insecure or insufficiently isolated systems. Security evaluation has consequently moved beyond checking whether an answer contains a dangerous string. It now asks whether the agent can be induced to disclose secrets, execute unapproved commands, manipulate data, cross trust boundaries, or pursue an unsafe objective across many steps. For a governed model pilot, the recommended standard is continuous evaluation before deployment, after material configuration changes, and whenever a new tool or data source is connected.

A useful distinction is between capability and permission. Capability testing asks what the agent might be able to do; permission testing asks what its identity, policies, and interfaces actually allow it to do. The second question is often more important because least privilege can contain a model error. Even so, restricting permissions does not eliminate the need for behavioral testing: an agent may misuse an allowed read operation, corrupt records through an allowed write endpoint, or socially engineer a human into approving a harmful action. Agent security evaluation combines adversarial prompts, tool-use traces, boundary tests, policy checks, outcome review, and operational monitoring into one decision process.

What an Enterprise Evaluation Must Measure

The first measurement category is task performance under benign conditions. Establish whether the agent can complete its intended job, follow domain rules, cite reliable sources, and stop when information is insufficient. A security program that only attacks a weak system produces distorted conclusions, because failures may be confused with ordinary unreliability. Record task success, false-positive rate, completion time, tool-call count, cost per run, and human intervention frequency. For a pilot, sensible starting thresholds might be at least 90% success on representative routine tasks, no more than 5% unauthorized tool calls, and zero tolerance for production credential exposure.

The second category is adversarial resilience. Test direct prompt injection, indirect injection placed in documents or tool results, role confusion, encoded requests, malicious attachments, poisoned retrieval content, and attempts to override system instructions. Measure both the attack success rate and the severity of resulting actions. Counting one refused prompt as equivalent to one successful database deletion is misleading. A weighted metric should distinguish harmless policy language, unnecessary sensitive-data access, reversible sandbox changes, and irreversible production effects. Teams should also test multi-step attacks because an agent can accumulate impact through several individually plausible actions.

The third category is system control effectiveness. This includes authentication boundaries, network segmentation, file permissions, secret redaction, tool allowlists, approval gates, logging, rate limits, and rollback. An evaluation should deliberately attempt actions that should be impossible under the approved design. If the agent cannot reach production because the test identity has no route or credential, that is a positive control result rather than a model achievement. At the same time, a denial produced only by a long text warning is weaker than a denial enforced by the platform. The strongest evidence combines behavioral output with an independently verified inspection of actual privileges and effects.

A Repeatable Seven-Stage Evaluation Method

Start by inventorying the agent’s assets, identities, tools, data, destinations, and human dependencies. Assign an action to one of four risk tiers: read-only public data, confidential read, reversible write, and irreversible or regulated action. The August 2026 report about a rogue-agent breach attempt involving Medicare illustrates why consequential environments need especially strict separation from ordinary evaluation infrastructure. An agent should not inherit the permissions of the person who launched it, and evaluation credentials should never be production credentials. This inventory becomes the basis for both test design and least-privilege configuration.

Next, build a representative benchmark containing routine, edge, and hostile cases. Include at least 20 to 50 normal tasks for an early pilot, then expand the suite as production evidence accumulates. A practical hostile set might contain 100 or more attacks covering prompt injection, data exfiltration, privilege escalation, malicious code, unsafe tool invocation, and social engineering. Run each case repeatedly because agent behavior can vary with sampling, tool availability, context length, and external content. Record the model version, full prompt, retrieved context, tool arguments, responses, timestamps, approvals, and final state so that a security engineer can reconstruct exactly what happened.

Then test controls in layers. Evaluate the base model, the prompted agent, the tool-enabled agent, and the fully connected system. If a vulnerability disappears when a restricted identity is substituted, document that control dependency rather than declaring the model intrinsically safe. For high-impact actions, require deterministic policy enforcement outside the model, explicit human approval, transaction limits, and a tested rollback path. After execution, compare observed actions against policy: every tool call should map to an approved purpose, every data transfer should have an authorized destination, and every exception should produce an auditable event. The stage is complete only when the organization can state its residual risk and the conditions under which the agent may operate.

Comparing Evaluation Approaches and Alternatives

There is no single product category called an agent security evaluator. Most organizations combine internal testing, commercial agent or application security platforms, model-provider red teaming, and infrastructure controls. The best choice depends on whether the priority is discovering model behavior, validating code, governing enterprise use, or measuring an entire tool-using system. Open-source policy tools such as Open Policy Agent can enforce decisions, but they do not by themselves provide a complete behavioral benchmark. Static application security tools can inspect code and dependencies, while runtime tests are still needed to reveal whether an LLM selects unsafe operations.

FeatureInternal red-team programCommercial evaluation platformModel-provider evaluationInfrastructure controls
Primary valueTests business-specific abuse casesScales repeatable suites and reportingExposes model behavior and known failure modesPrevents or limits harmful actions
Best coverageEnd-to-end workflows and domain policyMulti-agent, tool-use, and regression testsBase-model safety and instruction handlingIdentity, network, secrets, approvals, and audit
Typical early pilot cost2–6 engineer-weeksSeveral thousand to tens of thousands of dollars per yearProvider programs, credits, or negotiated accessExisting platform spend plus configuration work
Main limitationSlow and difficult to reproduceCoverage and claims require validationMay not represent enterprise tools or dataCan contain impact but may not improve behavior
Evidence neededAttack scenarios and incident historyVersioned results and inspectable tracesPublished methodology and reproducible resultsPolicy tests, logs, and access reviews
A strong program uses all four rather than asking one layer to perform work that it cannot validate. Commercial pricing varies substantially and may be quote-based, so published list prices are rarely useful as a universal benchmark. For planning purposes, a small internal assessment may require 2 to 6 engineer-weeks, while recurring commercial evaluations often range from several thousand dollars to tens of thousands of dollars annually. Model API and sandbox infrastructure costs are usually smaller than the engineering cost of designing valid tests. Enterprise buyers should price incident analysis, red-team maintenance, policy engineering, and human approval operations separately, because hiding those costs under a software fee obscures the real budget requirement.

Test Multi-Agent and Tool-Enabled Systems Carefully

Multi-agent systems are harder to evaluate because authority can be split, messages can carry untrusted instructions, and one agent may treat another agent’s output as trusted. The supplied research reports a medical-agent evaluation context and a security league for AI-generated software, both of which point to domain-specific testing rather than a generic “does it refuse harmful requests” check. In a multi-agent design, label every message with its origin, trust level, permitted uses, and response requirements. Do not allow a planner-generated instruction to inherit the privileges of a human policy source merely because both appear in the same transcript.

Use a broker between agents and sensitive tools. The broker can strip active content from retrieved documents, validate schemas, constrain tool arguments, enforce transaction budgets, and require approval for high-risk effects. Test horizontal attacks, where one compromised agent impersonates another, and vertical attacks, where a low-trust research agent attempts to direct a privileged execution agent. Also examine resource-exhaustion patterns: repeated searches, long-running shells, unbounded tool loops, and expensive model calls. A security evaluation should record token usage, wall-clock time, number of calls, and spend limits, with alerts at perhaps 2 times the expected median for a routine task and a hard cap determined by the workflow’s budget.

External services create another problem because they can change after certification. A tool description may be harmless when approved and malicious after a prompt update, account compromise, or dependency change. Record service version, endpoint, response schema, and permission scope for each run, then retest whenever those inputs change. For coding agents, scan generated code, dependencies, secrets, and build scripts; for research agents, test source integrity, citation laundering, hidden instructions in retrieved pages, and unauthorized publication. The relevant question is not whether every individual step looks reasonable in isolation, but whether the combined action chain stayed within the user’s authority and the organization’s policy.

Common Mistakes That Produce False Confidence

The most common mistake is evaluating a helpfulness prompt instead of the deployed agent configuration. A concise chat demonstration does not test the memory store, retrieval pipeline, tool permissions, approval interface, or production data. Another common error is treating refusal as the only safe outcome. An agent can obey a security instruction literally while still revealing information through logs, tool arguments, cached memory, or side effects. Tests must therefore inspect system state and data destinations, not just the final natural-language response.

Teams also overcount spectacular demonstrations. A single successful injection example is less informative than a stable rate across hundreds of versioned runs, and cherry-picked failures inflate risk. Avoid both extremes: a tiny set of easy attacks cannot represent adversarial use, while intentionally unrealistic threats can make an otherwise suitable pilot impossible. Derive scenarios from actual assets, plausible users, incident history, and the agent’s granted authority. As a mature benchmark starting point, aim for at least 100 adversarial cases and 30 repeated executions of the highest-impact scenarios, then expand based on observed near misses.

A further error is assuming that better prompting creates a reliable security boundary. Prompts can improve behavior, but they remain vulnerable to context changes and novel attacks. Enforce permissions, network rules, data-loss prevention, and approvals in deterministic systems. Do not allow secrets to be available merely because an agent might need them for debugging, and do not permit a model to approve its own high-impact action. Finally, avoid declaring success from average metrics that conceal a critical tail. If 99% of runs are safe but one in 100 can approve a payment, the system may still be unacceptable for that payment authority.

When to Block, Approve, or Expand a Pilot

Enterprises should block deployment when testing reveals cross-boundary access, credential exposure, unrecoverable actions, misleading audit trails, or a material inability to stop the agent. Zero tolerance is appropriate for production secret disclosure, unauthorized access to regulated records, and execution outside the approved environment. A pilot may continue with human supervision when residual failures are reversible, detection is reliable, and the operator can inspect every consequential action. It should proceed unattended only when the organization has measured the full distribution of behavior, enforced least privilege, tested rollback, and established response procedures.

Define thresholds before seeing results to reduce pressure to reinterpret inconvenient findings. For example, require a 99% or higher refusal rate on critical direct attacks, at least 98% prevention of indirect data-exfiltration paths, zero unauthorized writes, and no unresolved high-severity code vulnerabilities. These are planning targets, not universal standards; regulated or high-consequence systems may demand stricter thresholds. Run a staged rollout with a small user population, synthetic data where possible, read-only tools, and tightly bounded credentials. Re-evaluate after every model, prompt, retrieval, tool, permission, or policy change, and at least monthly once behavior is stable.

A pilot should be paused when an external service changes, monitoring gaps appear, or the agent begins handling a new data class. Near misses count as evidence rather than embarrassing exceptions: record them, determine whether a control worked, and add a regression case. The decision to expand should be based on a time-bounded observation period, such as 30 days for a low-risk internal workflow or 90 days for a more exposed system. Enterprise AI Labs’ relevant role is not to claim that one score proves safety, but to help teams define governed pilots, version evaluations, compare configurations, and retain decision evidence across model and platform changes.

The Minimum Enterprise Standard

By September 2026, a defensible agent security evaluation should leave an auditor with five artifacts: an asset and permission inventory, a versioned test suite, raw execution traces, verified control results, and a signed deployment decision. The inventory states what the agent can access. The suite states how it was attacked and how normal performance was measured. Traces show which model, prompt, context, tool, identity, and policy were used. Control results establish that blocked actions were impossible or contained rather than merely discouraged. The decision records accepted residual risk, monitoring requirements, rollback conditions, and named owners.

Use recognized control frameworks as reference points rather than treating any framework as a complete agent test. The NIST AI Risk Management Framework supports governance and risk treatment, while OWASP guidance on LLM and generative-AI application security provides threat-oriented guidance for applications using models. Provider research on real-world evaluation incidents can inform threat scenarios, but it should not be treated as a substitute for testing the enterprise’s own configuration. Standards will continue to develop, yet the basic enterprise requirement is already clear: no autonomous agent should receive broader authority than the organization can observe, reverse, and explain.

The practical answer is therefore to test the smallest realistic agent system first, then add complexity one trust boundary at a time. Begin with read-only access, synthetic or low-sensitivity data, and short task limits. Add writes only after injection and exfiltration tests pass. Add human approval before irreversible actions. Scale the number of users, tasks, and tools only when the residual-risk thresholds are met over repeated runs. This approach is more demanding than a model benchmark, but it measures the thing enterprises actually deploy: an agent capable of causing effects through a governed technical system.