What Are Agent Runtime Security Controls?
Agent runtime security controls are technical and administrative safeguards applied while an AI agent is planning, invoking tools, accessing data, or taking an action. Unlike model evaluation, which examines expected behavior before release, runtime controls govern actual behavior under changing inputs and permissions. The central security boundary is the execution path: the model, orchestrator, tools, credentials, data stores, and external services involved in an agent’s actions. This matters because an agent can produce harmful effects even when its underlying model has passed accuracy, safety, and red-team tests.
Also worth reading: What Is Enterprise Agent Runtime Security and How Should Enterprises Evaluate It in 2026? · How Can Modern Enterprises Implement Agentic Workflow Runtime Governance Effectively? · How Should Enterprises Build Effective AI Governance for Models, Pilots, and Autonomous Agents in 2026?
The controls should answer four operational questions: what may the agent do, which identity may perform the action, what data may it use, and how can risky behavior be detected or stopped before completion. Suitable measures include scoped identities, short-lived authorization, tool allowlists, pre-execution policy checks, input and output filtering, data-loss prevention, transaction limits, approval gates, sandboxing, and immutable audit logs. These controls do not make an agent trustworthy by themselves. They reduce the likelihood and blast radius of prompt injection, credential theft, excessive permissions, data exfiltration, unintended tool use, and unauthorized side effects.
The terminology became more prominent in 2026 as security companies positioned runtime enforcement as a separate control layer. NVIDIA announced an open agent safety platform intended to cover agents from testing through deployment, while reporting also described vendors raising capital around AI-agent runtime security, including an $8 million round for Arrakis and a $4 million round for Kontext Security. These announcements do not prove that the market has settled on one architecture, but they indicate that enterprises now recognize a gap between pre-deployment model testing and production action control.
Why Prompt and Model Testing Are Not Enough
Testing asks whether a model behaves correctly under a defined set of cases. Runtime security asks whether a complete agent remains acceptable when a tool response contains an injection attempt, a user changes an instruction midway through a task, an API returns unexpectedly large data, or one compromised component attempts to escalate privilege. An agent is therefore a distributed system with probabilistic decision-making at its center. Its security depends not only on model behavior but also on orchestration logic, identity configuration, context assembly, retrieval sources, tool semantics, network paths, and downstream authorization.
A survey framed as a review of 247 papers, titled “Agent Security Is a Systems Problem,” reflects this broader view. The number is useful as a measure of research attention, not as proof of consensus or effectiveness. A paper count cannot establish that a proposed defense works against novel attacks, especially because agent behavior changes as tools, models, context windows, and enterprise integrations change. It does, however, support the conclusion that agent security cannot be reduced to a single prompt filter, classifier, or model alignment score.
Runtime threats also cross conventional security boundaries. Traditional application security controls remain necessary, including vulnerability management, API authentication, network segmentation, secrets management, and endpoint detection. Agent workloads add decision-dependent behavior: the system can interpret natural-language instructions and dynamically select actions that were not anticipated in a static application-security rule. Consequently, a technically valid API credential can still represent excessive authority. A permitted email tool can become a spam or data-theft mechanism when policy does not constrain recipients, message content, attachment access, or call frequency. Runtime controls connect model intent to enforceable system policy.
The Main Control Points and Enforcement Model
A defensible design places controls before, during, and after each consequential action. Before execution, a policy engine can inspect the authenticated user, agent identity, task, model, requested tool, parameters, relevant data classifications, session state, and risk score. It can then allow, deny, transform, or route the request for approval. During execution, controls can constrain the tool itself through scoped credentials, input schemas, network rules, filesystem isolation, query limits, and transaction boundaries. After execution, the platform can inspect results for sensitive data or policy violations and record enough evidence to reconstruct the action chain.
Identity is especially important because “the agent called the tool” is not an adequate authorization statement. The enterprise should distinguish the initiating human, the workload operating the agent, the specific agent version, and any delegated service receiving the request. Privileges should be non-human, short-lived, environment-specific, and limited to the minimum resources required for the current task. For example, a reporting agent should not automatically receive unrestricted database administration rights simply because one of its workflows occasionally writes a report. Separate identities and interfaces should be used for reading records, generating content, publishing results, and approving external actions.
Controls may operate synchronously at an execution gateway, synchronously inside tools, asynchronously through monitoring, or through a combination of these methods. Pre-action checks provide the clearest chance to prevent harm, but they cannot stop every malicious result or compromised downstream service, so post-action monitoring remains necessary. High-impact actions should use explicit approval gates: a human should confirm external payments, production changes, bulk exports, privilege changes, deletion requests, customer communications, or access to highly sensitive data. Approval should be contextual and time-bound rather than a permanent checkbox in a prompt. A request approved once should not silently authorize broader future behavior.
Practical Controls for Enterprise Agent Pilots
The first practical step is to inventory what the agent can reach. Build a map of models, tools, identities, data sources, networks, and external actions, then classify each capability by business value and potential harm. Include indirect capabilities such as browser access, shell commands, code execution, package installation, email, file transfer, vector stores, ticketing systems, and customer relationship management tools. Set measurable thresholds before connecting production systems: for instance, no more than 100 records per query, no production write access during the pilot, no arbitrary outbound domains, or mandatory approval for any action affecting more than one customer.
The second step is to define acceptable behavior in testable terms. Convert broad statements such as “protect confidential data” into rules for specific data classes, destinations, and actions. DLP can block secrets, regulated records, or customer identifiers before they cross an agent boundary, while content filters can detect attempts to encode or fragment sensitive values. Tool-level controls should validate arguments against a schema and reject fields that the workflow does not require. A general financial agent may not need account numbers, personal identifiers, or arbitrary memo text, so parameter filtering can remove fields before the model or downstream service processes them.
The third step is to pilot during a limited execution period. The stage, tool registry, credential broker, policy engine, and audit system should be tested with benign and malicious scenarios before business users can invoke the agent. A practical first gate is 30 days with sandboxed tools, synthetic data, and read-only access; a second 30-day stage can add approved internal tools and human confirmation. Teams should record blocked requests, denied tools, approval rates, false positives, policy latency, and attempted policy bypasses. Production access should depend on evidence that the controls work, not merely that the model’s task-completion score remains high.
For governed model pilots and evaluation, the objective should be reproducible testing across both model candidates and runtime configurations. Change two models only when needed, because otherwise teams cannot attribute a failure to the model or the enforcement layer. Version agent instructions, tool schemas, retrieval indexes, policy rules, identities, and evaluation datasets. The most useful pilot report is not a single safety score; it is a breakdown of task success, unauthorized-action rate, sensitive-data exposure, policy-bypass resistance, approval quality, latency, cost, and operator effort.
Comparing Preventive, Detective, and Hybrid Controls
Preventive controls reject or constrain behavior before an action occurs. They are generally more reliable for protecting data and production assets because they can stop a harmful request, but they may miss novel attacks and can create friction if policies are too broad. Detective controls identify suspicious behavior during or after execution. They are valuable for discovering prompt injection, anomalous tool sequences, exfiltration patterns, and compromised agents, but detection alone cannot guarantee that a payment, deletion, or disclosure never occurs.
| Feature | Preventive enforcement | Detective monitoring | Hybrid enforcement |
|---|---|---|---|
| Primary purpose | Block unauthorized or excessive actions | Identify suspicious behavior and support investigation | Block known risks and investigate novel behavior |
| Typical controls | Tool allowlists, schema validation, DLP, scoped credentials, approval gates | Behavioral analytics, anomaly detection, output inspection, audit review | Policy gateway plus telemetry, response automation, and human review |
| Main advantage | Reduces immediate impact | Improves visibility and can reveal new attack patterns | Balances prevention with operational learning |
A hybrid model is usually the more credible enterprise choice, but complexity introduces its own risks. Policy services can become a single point of failure, integrations may create latency, and alert volumes can overwhelm operators. Therefore, teams should begin with a small number of high-consequence policies, test fail-open versus fail-closed behavior, and define who can override a control. Security systems should not silently convert an unavailable approval service into unrestricted execution for a high-risk action. That decision can be appropriate for a low-risk internal query, but not for production writes, bulk exports, privilege changes, or external communications.
Common Mistakes in Agent Runtime Protection
One common mistake is treating prompt injection as a model classification problem alone. Attackers may place instructions in retrieved documents, tool results, web pages, email, code comments, or images, so filtering only the user prompt is insufficient. Another is giving an agent a broad service-account credential and relying on later monitoring. Monitoring may reveal misuse, but the credential has already enabled the misuse; authorization belongs at the tool, data, and transaction layers where it can be enforced deterministically.
Teams also err by applying excessive autonomy before defining rollback. If an agent can change infrastructure, send messages, update records, or move money, the platform needs compensating actions such as cancellation, reversal, quarantine, and incident linkage. A log without a usable rollback mechanism is forensic evidence, not prevention. Similarly, a human approval prompt is ineffective if the approver sees no tool, arguments, data destination, estimated impact, or reason for the action. Approval quality depends on understandable context, not just a modal window.
Security evaluations can become stale because the environment changes. A tool that was safe with ten approved operations may behave differently after a new dependency, permission, data source, or model release is introduced. Teams should schedule re-evaluation after material changes and at least quarterly for active production agents, with immediate retesting after incidents. There is no universal 90-day certification period, so the interval should be based on change frequency and risk. High-privilege, externally accessible agents should be tested more frequently than read-only assistants operating on synthetic data.
Finally, many organizations confuse activity volume with risk reduction. A control that blocks 20,000 actions but allows broad credential export is less valuable than one that prevents the export and logs its context. Measures should be outcome-based: unauthorized action rate, data prevented from leaving approved boundaries, mean time to revoke an identity, percentage of actions reconstructable from logs, and time to contain a compromised session. Raw request counts and dashboard coverage are useful diagnostics but are not sufficient security evidence.
When to Act, and What It May Cost
Enterprises should act before an agent can affect production data or external systems. Read-only experimentation on public, synthetic, or low-sensitivity material can begin with basic logging and sandboxing, but it is not equivalent to a production readiness decision. A practical trigger is the first intended connection to internal records, customer information, privileged APIs, executable code, or irreversible actions. Regulated data, cross-tenant use, autonomous operation, internet-facing access, and multi-agent delegation justify stronger controls and independent review.
There is no standard market price for “agent runtime security controls” because the category includes gateways, identity brokers, DLP, behavior analytics, policy engines, sandboxes, and audit products. Pricing may be based on users, agents, tool calls, actions, protected resources, data volume, or a combination. A small internal pilot may cost less than $10,000 per month when it uses existing cloud infrastructure and commercial security tools, while an enterprise platform can reach tens of thousands or more per month depending on integrations and support. Implementation labor often exceeds the initial subscription, so teams should budget separately for threat modeling, tool adaptation, policy authoring, evaluation, and ongoing operations.
Cost estimates should be tied to a bounded scope rather than a vendor promise. One reasonable planning assumption for a 90-day pilot is $25,000 to $150,000, but the difference is driven mainly by integration count, data sensitivity, existing controls, and whether production systems are touched. Low-risk, read-only pilots can use managed sandboxes and existing identity services; high-risk deployments may require a dedicated policy plane, sensitive-data inspection, custom tool controls, legal review, and 24/7 response. Vendors that cannot state pricing dimensions, data handling terms, model-provider use, deployment options, and integration costs should not be compared on a single annual total.
Budget decisions should also include expected overhead. Approval gates can slow tasks, DLP can flag legitimate information, and tightly scoped tool schemas can require workflow redesign. The relevant comparison is not “secure versus productive” but controlled risk versus unrestricted impact. For most enterprise pilots, a narrow set of tools, a limited set of data, and clear stop conditions produce better evidence than a broad agent connected to the entire enterprise.
How to Evaluate a Runtime Security Platform
Evaluation should begin with the execution model. Determine whether the platform protects model-originated tool calls, agent-to-agent calls, direct API access, browser actions, code execution, and connections made by third-party MCP servers. The answer is important because an attractive dashboard is insufficient if the agent can bypass the protected path and call the same service directly. Ask for an architecture diagram showing identity propagation, policy evaluation, enforcement location, logging, and failure behavior.
The evaluation should then test realistic abuse cases. Include direct prompt injection, indirect injection in retrieved content, malicious tool output, secret exfiltration, cross-tenant access, command injection, excessive tool loops, identity confusion, approval replay, and attempts to invoke disabled tools. Test not only whether malicious prompts are rejected, but whether equivalent actions are rejected when they arrive through a different tool, language, encoding, or orchestration route. Record latency and failure rates because controls that time out may push the application toward an unsafe fallback.
Governance evidence matters as much as technical performance. A platform should support role-based administration, separation of duties, versioned policies, configurable retention, regional data handling, revocation, export of audit records, and documented incident response. Procurement teams should clarify whether telemetry is used to train a vendor’s models, whether prompts and tool results are retained, and who owns policies and evidence. A credible evaluation produces test results, configuration artifacts, policy coverage, known limitations, and an owner for each gap rather than a generic claim that the product is “agent-safe.”
The Recommended Enterprise Operating Model
The strongest operating model makes agent security a shared responsibility. The model provider contributes safe capabilities and documentation; the platform team implements identity, tool boundaries, policy, and monitoring; the application owner defines business impact and approval rules; security operations handles detection and response; and legal or compliance teams establish requirements for data and actions. Shared responsibility does not mean diffuse accountability. Every production action should have a named business owner, while every enforcement point should have a technical owner and a test procedure.
A mature program can proceed in four stages. First, agents remain sandboxed and use synthetic or low-risk data. Second, they receive read-only, least-privilege access with complete tracing. Third, narrowly defined write actions are enabled with approval and rollback. Fourth, higher-volume operation is permitted only after measured control performance, incident exercises, and independent review. This progression avoids a binary choice between unrestricted autonomy and a complete ban. It recognizes that security evidence should accumulate as consequence and exposure increase.
The date context of September 28, 2026 also argues against waiting for a final category definition. Models, tool protocols, identity systems, and vendor offerings continue to change, and the 247-paper research framing is best understood as evidence of an active systems problem rather than a finished playbook. Enterprises should establish a minimum control baseline now, then revise it as agents move from evaluation into governed production use. The aim is not to make every agent safe in the abstract; it is to ensure that each permitted action is attributable, bounded, observable, and stoppable in the actual enterprise environment.