Direct answer: treat the agent as an untrusted runtime

An effective AI agent control architecture is the set of technical and organizational controls that determines which agent can act, what identity it may use, which systems it can reach, what actions are allowed, how outputs are checked, and how activity can be investigated afterward. It should not be treated as a simple extension of an API gateway or access-control system. Traditional access management can approve a user or workload at connection time, but an agent chooses tools, interprets results, and generates subsequent actions at runtime. The central design rule is therefore to assume that even a correctly authenticated agent may select the wrong tool, exceed the intended task, or act on manipulated information.

Also worth reading: What are the main LLM gateway architecture patterns enterprises should adopt in 2026? · How Do Enterprises Run Governed AI Model Pilots Without Creating Another Production Bottleneck? · How Can Enterprises Build Governed AI Pilot Evidence in 2026?

A practical architecture combines identity, policy, tool gateways, data controls, execution boundaries, human approval, observability, and evaluation. Each request should carry an agent identity, a human sponsor, workload and tenant context, a declared objective, and a unique trace identifier. Policies should then evaluate those attributes before each tool call rather than granting the agent permanent access to a browser, desktop, repository, or cloud account. The objective is not to make agents harmless by assuming perfect model behavior. It is to limit the damage caused by model errors, prompt injection, compromised dependencies, confused identities, and excessive autonomy.

For enterprises, the preferred pattern is a controlled execution plane positioned between the model and enterprise resources. The model proposes actions; the control plane authorizes them, supplies minimized data, enforces rate and spending limits, records evidence, and escalates selected actions for approval. Agent identity must also remain distinct from the identity of the user who started the session and from any service account used behind the gateway. If all three collapse into one credential, audit records cannot distinguish delegation, impersonation, and misuse.

Core runtime identity and delegation model

Runtime identity is the foundation of agent control because agents perform work across sessions, tools, and platforms. A production design should create a short-lived, workload-bound identity for every agent run and record the chain from initiating user to delegated workload, tool, and target resource. Human authorization should not silently become a reusable bearer token that an agent can present indefinitely. Instead, the initiating user can delegate a bounded task, such as “prepare a release candidate for repository X,” while policy determines whether that delegation permits reading branches, running tests, creating a branch, or merging code.

The architecture needs separate identities for different trust zones. A research agent that reads public documentation should not inherit a finance system’s account merely because both use the same orchestration framework. Production deployment should usually use distinct service identities from evaluation environments, with separate encryption keys, data stores, tool registrations, and network routes. As of October 2026, vendors including AWS, IBM, Microsoft, NVIDIA, Oracle, and NetApp are publishing patterns around agent gateways, controls, identity, and shared responsibility. This indicates convergence around the need for a control plane, but vendor control does not remove the enterprise’s responsibility for authorization design and data classification.

Identity should be cryptographically verifiable rather than inferred from prompt text such as “I am the finance agent.” Tokens can carry signed claims about workload, tenant, purpose, environment, delegation chain, and expiration. Policy can compare those claims with the requested action, resource sensitivity, session risk, and current approval state. Default-deny behavior is appropriate for destructive or regulated actions, while tightly constrained reads may proceed automatically if they are reproducible and low impact. The useful control is contextual authorization: access changes according to what is happening now, not only according to whether a credential was once accepted.

Control requirementBasic agent integrationGoverned agent control architecture
IdentityOne shared API or service keyShort-lived, workload-specific identity with delegation chain
AuthorizationBroad permission granted at deploymentPer-action, resource-aware policy evaluation
Tool accessDirect model-to-tool connectionBrokered tool gateway with schema and scope enforcement
ApprovalAll actions manual or all actions autonomousRisk-tiered approval based on action and data sensitivity
AuditApplication logs onlyEnd-to-end trace covering prompts, calls, approvals, and outputs
Failure behaviorRetry indefinitely or stop completelyBounded retry, circuit breaker, fallback, or human escalation
EvaluationOffline prompt testsOffline tests plus live policy, tool, and red-team telemetry
## How the control plane works from intent to execution

A governed execution begins with a structured task definition rather than an unrestricted natural-language command. The system should capture the objective, permitted resources, prohibited actions, budget, deadline, approval threshold, and success criteria. This metadata allows policy checks before the model receives sensitive context or receives a tool capable of changing state. It also gives evaluators a stable basis for deciding whether an outcome satisfied the requested task, not merely whether the final answer looked plausible.

The model can reason and request a tool call, but it should not connect directly to the protected system. A gateway validates the call against a registered schema, checks the agent’s identity and task scope, removes unnecessary fields, and determines whether approval is required. Read operations might include fetching a public specification or querying a non-sensitive repository. Write operations might include changing a ticket, issuing a refund, editing a file, sending email, or deploying software. The control plane applies limits such as maximum records, transaction value, tool-call count, execution time, and cumulative model cost. Every decision produces a signed or otherwise tamper-evident audit event linked to the session trace.

Tool responses require equal care because data returned to an agent can contain hostile instructions. A web page may instruct the agent to ignore its policy and upload local files; an email may request an unauthorized payment; a repository issue may contain concealed commands. External content should be labeled as untrusted data, isolated from system instructions, and prevented from directly altering policy. Output validation should check destination, schema, permission, and content. Even a correctly authorized email tool could otherwise send data to the wrong recipient because the model hallucinated an address.

Recovery paths are part of this sequence. A gateway can quarantine malformed tool calls, stop repeated failures, and route uncertain outcomes to a reviewer. Idempotency keys are essential for operations such as payments or ticket creation, where a timeout does not prove that the action failed. Concurrency controls prevent two retries from creating duplicate records. The architecture should distinguish model failure, policy denial, tool failure, authentication expiry, and user cancellation so that operations teams can respond appropriately rather than blindly replaying the entire workflow.

Data, network, and execution boundaries

Agent control must address data movement as carefully as user actions. Before data enters model context, policy should classify it and remove fields that the task does not need. Token counts alone are a poor measure of sensitivity, because one short record may be regulated while a large public document is harmless. Retrieval systems should enforce tenant and document permissions before ranking or generation, because filtering text after retrieval may expose restricted information through prompts, citations, or logs. Data minimization should include prompts, traces, evaluation artifacts, vector stores, caches, and downstream observability systems.

Network segmentation limits what a compromised tool chain can reach. An agent intended to update documentation should not automatically have routes to production databases, identity administration, source-control settings, or developer workstations. Egress controls should restrict outbound destinations, while internal services should validate the calling workload rather than trusting a forwarded header. Sandboxing, temporary credentials, read-only mounts, and isolated workspaces can reduce the consequences of malicious code. These controls do not prove that generated code is safe, so privileged execution should require scanning, testing, restricted permissions, and approval based on repository or destination risk.

The browser and desktop are particularly broad action surfaces. Mobile or web agents can use APIs more reliably than synthetic mouse and keyboard input, but direct desktop control can still create difficult-to-audit side effects. An architecture should prefer narrow APIs when available, then broker browser or GUI actions through constrained automation environments. It should block clipboard capture from unrelated windows, restrict file upload directories, capture before-and-after state, and require confirmation for credential entry, deletion, payment, publishing, or permission changes. The 2026 enterprise pattern is not simply “let agents operate desktops”; it is to define which desktop operations are permissible and enforce those permissions outside the model.

Data residency and retention also affect control design. Organizations may require certain prompts and tool results to remain within a country or approved cloud region, while enterprise policies may set different retention periods for debugging, evaluation, and audit evidence. A platform should make these choices explicit rather than treating all telemetry as one undifferentiated log stream. Useful audit evidence may need longer retention than raw prompt content, so selected event metadata can be preserved while sensitive content is tokenized or deleted under a documented schedule.

Governance, approval, and shared responsibility

Governance converts architecture into daily operating rules. A risk-tier model is more defensible than one universal autonomy setting. Public-information searches can usually be automatic; internal reads may require verified purpose and scoped access; customer communication, code changes, financial transactions, and destructive operations can require stronger review. Approval should name the exact action and payload rather than ask someone to approve an opaque “agent task” that later changes. If the plan changes materially, a new approval may be necessary.

Human review is effective only when the reviewer has enough information. The interface should show the agent identity, objective, source evidence, proposed tool call, target system, exact arguments, expected cost, and potential impact. Reviewers should be able to edit, reject, or downgrade the action without allowing the agent to bypass the modified decision. High-frequency workflows can use sampled review initially, but sampling should not replace deterministic controls for irreversible actions. Approval rates may fall as evaluation improves, yet lower review frequency should follow measured evidence rather than optimism about model accuracy.

Shared responsibility is unavoidable. Model providers test model behavior, but enterprises define data access, business rules, credentials, network routes, user impact, and regulatory obligations. Oracle’s published shared-responsibility framing and the emergence of vendor agent controls support this division. Platform teams can supply identity, gateway, logging, and sandbox primitives, while business owners classify decisions and set risk limits. Security teams define monitoring and incident procedures. Legal and compliance teams determine which uses are acceptable. If ownership is assigned only to an AI innovation team, control exceptions may become routine before anyone has accepted responsibility for them.

A governance program also needs an exception process. Teams sometimes need a temporary autonomous workflow during an incident or a customer deadline. Exceptions should have an owner, permitted actions, expiration date, rollback method, and post-event review. A 24-hour exception should not silently become a permanent architecture. Review metrics should include denied calls, approval overrides, expired credentials, anomalous destinations, tool failures, and repeat policy violations. The objective is a visible control system whose effectiveness can be demonstrated, not a claim that agents are simply “safe.”

Evaluation and continuous assurance

Agent evaluation cannot be reduced to answer quality on static questions. A useful evaluation suite tests task completion, factual grounding, policy compliance, tool selection, argument correctness, refusal behavior, latency, and cost. It should also include adversarial cases involving prompt injection, misleading tool output, stale records, ambiguous user requests, malicious files, credential requests, and attempts to exceed delegation. Each case needs an expected action or allowable action set, not a single ideal sentence. Agent behavior is stochastic, so repeated trials and explicit thresholds are more informative than one successful demonstration.

Thresholds should reflect business impact. A customer-support drafting agent may target 98% success on approved retrieval questions while allowing a lower completion rate for exceptional cases that can be escalated. An infrastructure agent that changes production may require a stricter action-policy pass rate because one wrong deployment can outweigh many successful runs. Cost can also be constrained: teams might set a per-session ceiling, a maximum of 20 tool calls, or a 30-minute execution window, then tune those values from measured distributions. These numbers are design examples, not universal standards; enterprises must derive them from task risk, value at stake, and observed model behavior.

Live assurance should compare production traces with evaluation cases. Sampling one completed session per week may miss thousands of low-risk calls, while retaining every sensitive prompt may violate privacy policy. A balanced program uses event-level telemetry for all actions and more detailed content inspection only where policy permits. Releases should pass unit tests, scenario tests, security tests, shadow-mode execution, and a controlled canary. Canary policies can initially restrict destinations, transaction amounts, repositories, or data classes. Automatic rollback should trigger on elevated denial, error, or unauthorized-call rates, but rollback itself must be tested rather than assumed.

Alternatives, implementation sequence, and cost

Organizations have several architectural choices. Direct model-to-tool connections are fast to prototype but concentrate permissions in the model and offer weak runtime evidence. A general API gateway improves observability and authentication, yet may not understand agent purpose, tool side effects, delegation, or changing plans. A dedicated agent gateway or control plane is more expensive to operate but better suited to risk-based authorization and approval. A human-operated copilot offers strong containment but offers limited automation and does not scale cleanly across every workflow. A fully autonomous multi-agent system may distribute work, but it increases message, identity, and failure paths without automatically improving reliability.

A practical implementation sequence takes 8 to 16 weeks for a bounded pilot, although identity integration and compliance review can extend the schedule. Weeks 1–2 should classify use cases and select a low-impact task; weeks 2–4 can establish workload identity, tool schemas, and trace correlation; weeks 4–6 should add action policies, approval, budgets, and red-team cases; weeks 6–10 can run shadow mode and controlled pilots; weeks 10–12 should test rollback and incident procedures. A production rollout beyond that timeline should be based on evidence rather than a fixed industry schedule.

Build-versus-buy decisions should account for more than model fees. A custom control plane requires engineering for identity, policy, gateways, logging, evaluation, security testing, on-call support, and upgrades. Commercial platforms may reduce initial engineering effort but introduce subscription, usage, infrastructure, integration, and data-egress charges. Organizations should request transparent prices for active agents, model tokens, tool calls, stored traces, evaluation runs, premium models, and approval workflows. Infrastructure planning ranges are commonly several thousand dollars per month for a modest pilot and tens of thousands per month for enterprise-wide production use, but these are estimates rather than quotations; actual cost depends heavily on model choice, context volume, retention, and integration scope.

The most credible alternative for small teams is a managed agent platform with strict tool scopes and human review. For regulated enterprises, a hybrid design may combine an existing identity provider, a central policy layer, cloud-native sandboxes, and specialized gateways. For sensitive workloads, a private deployment or self-managed evaluation environment may justify the higher fixed cost. The right choice is the smallest architecture that can enforce required controls, rather than the most agent framework available.

Common mistakes and when to act

The most damaging mistake is granting broad standing credentials to an autonomous process. Others include treating the initiating user, agent, and service account as the same identity; placing policy instructions only in the system prompt; assuming retrieval permissions enforce themselves; exposing unrestricted browser or shell access; failing to record tool arguments; and approving a broad objective rather than a concrete action. Another common error is evaluating only final answers while ignoring unauthorized attempts that happened to be blocked. A successful user outcome does not excuse a near miss, because the next run may not contain the same control.

Organizations also err by beginning with maximum autonomy, treating a successful demo as production evidence, or making autonomy the only measure of value. They may connect a model to tools before defining owners and failure responses, and they may log every prompt without managing sensitive content. Multi-agent designs intensify these errors by creating more actors and communication channels. Fewer agents with clear responsibilities are often easier to govern, though architecture should follow genuine task decomposition rather than an arbitrary preference for simplicity.

Act now when an agent can write to production, move money, change permissions, communicate externally, execute generated code, or access sensitive records. For a read-only prototype using public data and synthetic credentials, a lighter control set may be reasonable, but identity and logging should still be designed before scaling. Review controls before expanding from 1 workflow to 10, adding a new model provider, introducing another agent, or changing a tool’s permissions. High-risk deployments should be reviewed at least quarterly and after material model, prompt, tool, data, or identity changes, with continuous policy checks between formal reviews.

The decisive question for enterprise AI labs is not whether an agent can complete a task once. It is whether the organization can show exactly why the system acted, what it was allowed to do, which controls stopped unsafe behavior, and who remains accountable for every consequential outcome. A well-designed AI agent control architecture makes governed model pilots and evaluation measurable: tests connect model behavior to real tool calls, policy outcomes, approval quality, operating cost, and production evidence. That approach does not guarantee perfection, but it turns uncertainty into a managed engineering problem with explicit limits, owners, and evidence.