AI Agent Security

AI agent security protects the data and systems an agent can read or change. A useful design assumes the model may misunderstand a request or follow hostile content, then limits the actions that mistake can cause.

Reviewed October 5, 2026. This article covers selected developments from October 5, 2025 through October 5, 2026, followed by an original support-agent example. Dated announcements are separated from ongoing implementation guidance; a standards initiative is not a certification of a product.

TL;DR

These principles align with OWASP's agent-security guidance.

Quick Example

This complete Python 3 program models a read-only support tool. The trusted application supplies the session; the model supplies only a proposed tool call. It uses no model API, database, or network connection.

The program returns one subject and denies both forbidden calls. Its limitations are deliberate: production authentication, persistent storage, rate limits, and audit logging are absent. Its narrow lesson is that changing a model's requested identifier must not change the caller's tenant or authority. OWASP authorization guidance

Core Concepts

Instructions and evidence have different authority

A support ticket might contain a sentence asking the assistant to retrieve another customer's invoice. That sentence belongs to the ticket being processed; it does not grant new access. Filtering suspicious text can help, but the executor must still reject requests outside the authenticated scope. OWASP agent-security guidance

Identity must survive delegation

Record which user or workload initiated a task and which service identity executes it. If a coordinator delegates to a specialist, the handoff must not silently expand access. NIST's February 5, 2026 concept-paper announcement specifically identifies agent identity, authorization, and auditing as areas for investigation. NIST agent identity announcement

Authorization belongs at execution time

Showing a tool in a menu is not authorization to use it on every object. Validate the action, resource, and current permissions at the enforcement point, with denial as the default for unmatched requests. A guessed identifier must receive the same access checks as one returned by search. OWASP authorization guidance

What Changed During the Past Year

Sources: OWASP Agentic Top 10 release, NIST concept-paper announcement, NIST initiative launch, and OWASP Agent Control Standard.

The practical interpretation is to make policy enforcement explicit and testable. These publications have different roles: a risk list, a concept paper, a standards initiative, and a control interface are not interchangeable guarantees.

Design a Support Agent's Action Boundary

Split lookup, drafting, and sending

Consider an agent that answers billing tickets. Give it separate operations to read an authorized ticket, draft a response, and request delivery. Reading a ticket should not implicitly make an arbitrary recipient eligible to receive customer data.

In this example design, the delivery service resolves recipients from the authorized ticket record. The model may draft the body, but cannot replace the account owner or inject a new external destination. If the product supports extra recipients, that becomes a separate, explicitly authorized operation.

Make approvals describe the action

When review is required, display the recipient, message, attachments, and affected account. Associate approval with that exact action and invalidate it after changes. Enforce expiration and one-time use in trusted storage; a model-generated approved: true field supplies no evidence of approval. OWASP transaction-authorization guidance

For the billing example, changing the recipient after review creates a new proposed delivery. Retrying delivery safely requires the delivery service to enforce idempotency with durable deduplication; an operation identifier and a prior-result lookup alone do not prevent concurrent or crash-related duplicates. If the provider cannot resolve an uncertain outcome, pause automatic retries for reconciliation.

Give failures a bounded path

Define what happens after an unavailable lookup, denied action, or uncertain send result. Our example agent stops after two failed lookups and hands the ticket to a person; that is an illustrative product decision, not an industry-wide limit. Keep the threshold in application configuration and include the reason in the ticket's operational history.

Best Practices

Evaluate outcomes, not reassuring language

For the example service, test a ticket that asks to change tenants, a forged delivery approval, a tool timeout, and a cancellation during drafting. Check stored messages and tool events to establish what happened. A final answer saying “I refused” is insufficient if a prohibited request already reached the backend.

Keep a compact audit record

Capture the initiating identity, operation identifier, tool, target identifier, policy decision, outcome, and relevant version information. Exclude secrets and minimize message contents. Restrict access to logs and protect them against tampering and log injection. OWASP logging guidance

Recheck the boundary after changes

Treat a new tool, changed permission scope, or new memory source as a reason to rerun the example's abuse cases. A model-only upgrade may also change which paths are exercised. Keep successful-task checks in the same evaluation so security changes do not silently make the assistant unusable.

Comparison

These controls address different failure points. OWASP describes layered tool, input, execution, and monitoring defenses in its agent-security cheat sheet.

Common Mistakes

Letting the model identify its own tenant

Bad: Accept tenant_id from a model-generated tool call as the authorization scope.

Correct: Derive identity from authenticated application context, then check the requested resource against it.

Treating confidence as permission

Bad: Send a billing message whenever the model reports high confidence.

Correct: Evaluate access and any review requirement independently of the model's confidence.

Testing only the final answer

Bad: Score a run as safe because its final sentence refuses a forbidden request.

Correct: Inspect the tool trace and actual resulting state, including intermediate calls and retries.

FAQ

Does MCP make an agent secure automatically?

No. A protocol connection does not establish the user's authority to act on each application resource. Keep authentication and resource authorization in the integration and executor.

Does every tool call need a person to approve it?

Design the review policy around the task, permissions, and consequences. In this article's example, a scoped ticket lookup runs automatically, while an externally visible delivery follows a separate policy. Both paths still need authorization.

What should a small team test first?

Take its highest-impact tool and exercise one unauthorized identity, one out-of-scope resource, one changed action after approval, and one retry. Record the expected backend state for each case.

Is a standards initiative a production security standard?

No. NIST's February announcement describes a program of standards, protocol, and research work. Evaluate specific published artifacts and their status before claiming conformance. NIST initiative announcement

Related Topics

References