David Knichel, PhD
← All notes

AI · Security

Technical foundations of agentic security

How untrusted content can influence an agent’s actions, and how trust boundaries, permissions, and runtime checks can limit the consequences.

Suppose an assistant compares a supplier’s proposal with an internal cost estimate. It can read documents, retrieve company data, and send email. The user asks for a draft assessment. The supplier controls the proposal, but should have no authority to request internal files or decide where the assessment is sent.

Now suppose the proposal contains instructions that persuade the assistant to send the internal estimate to an external recipient. The email service may execute a valid request using valid credentials. Nevertheless, the system has disclosed information outside the user’s task.

The central issue is whether information from a less trusted source can acquire authority over an agent’s actions. Understanding this requires looking at the model, the software that executes its decisions, and the services it can access together.

This article develops that perspective from Juhee Kim, Wenbo Guo, and Dawn Song’s SoK: Attack and Defense Landscape of Agentic AI Systems, published at USENIX Security 2026. Their systematization reviews 128 papers, with a literature cutoff of October 2025. It organizes existing attacks and defenses; it does not experimentally establish a universally secure agent architecture. The supplier example below is illustrative, and the final design and testing suggestions are engineering deductions from the paper.

System overview

From external content to real actions

User task & application policyDefine the permitted data, operations, and destinations.

These permissions constrain execution. Retrieved content cannot expand them.

01 / EntryExternal content

Documents, web pages, tool descriptions

May contain attacker-controlled instructions.
02 / DecisionModel & context

Interpret inputs, plan, propose a tool call

Failure: treating external content as authority.
03 / ActionExecution layer

Check the operation, data, and destination

If permitted: call the tool or service.

Tool results return to the model’s context. Memory can carry their influence into later tasks.

Supplier example

Proposal requests disclosure → model proposes an email → a draft-only policy rejects sending.

Figure 1. External content supplies evidence, not permission. Source handling limits how it influences decisions; authorization checks and restricted tools limit which decisions become actions. The checks must cover every route with an external effect. This simplified view illustrates the article’s example, rather than a complete defense architecture.

1. What changes when a model can act?

A language model produces an output from its current input, or context. In an agent, that output can specify a tool call: a request to read a file, query a database, send a message, or execute code. The surrounding software validates and executes the call, then returns the result to the model. The model can use that result to decide what happens next.

This creates a feedback loop. A tool result is both information about the environment and a possible influence on subsequent decisions. Persistent memory extends that influence to later tasks. In systems with several agents, one agent’s output may also become another agent’s input.

The paper therefore treats an agent as a combined system of models, workflow logic, memory, tools, and an external environment. Security depends on the connections between these components. A model’s incorrect decision can become a real operation when the execution layer accepts it. Paper, §3.

Data and authority

The supplier’s proposal is relevant evidence for the assessment. It has no authority to change the user’s task. The assistant should extract prices and delivery conditions from it, while treating any request to disclose internal information as untrusted content.

A trust boundary separates components or data sources with different authority. Here, the boundary lies between supplier-controlled content and the trusted instructions and policies governing the task. The failure occurs when content crosses that boundary as an instruction that the system follows.

An instruction in a system prompt can describe this boundary. Enforcement also depends on what the surrounding software permits. If the model proposes an unauthorized email, the execution layer still has an opportunity to reject it.

2. Following an attack through the system

To make the assumptions explicit, consider an external attacker who can alter the supplier document. The attacker cannot modify the user’s request, change the assistant’s code, or directly access the internal estimate. The assistant, however, can read that estimate and has an email tool whose permissions are broader than this task requires.

A possible attack proceeds in four steps:

  1. Entry. The assistant retrieves the proposal. Alongside the commercial terms, the document contains an instruction to transmit internal information as part of a supposed verification procedure.
  2. Influence. The model treats that instruction as a necessary step in completing the assessment. The external document has changed the plan.
  3. Action proposal. The model requests an email containing the internal estimate and specifies an external recipient.
  4. Execution. The surrounding software checks that the email request is well formed and that its credentials permit sending, then executes it without checking the task’s authorization.

The attack vector is indirect prompt injection: instructions enter through material retrieved while performing a legitimate task. The behavioral failure is following those instructions. The consequence is disclosure of confidential information. These describe different stages, and defenses can interrupt different stages. Paper, §§4.1–4.3.

No defect in the email API is required. The attacker has induced the assistant to use its existing authority for an unauthorized purpose. Possession of a credential establishes that the service will accept certain requests. It does not establish that a particular request belongs to the user’s task.

The example also shows why checking only the final answer is insufficient. The assistant could report that the assessment is ready after the disclosure has already occurred. A preventive check must run before the operation with the unwanted effect.

3. Which design choices determine exposure?

The paper identifies seven dimensions of agent design. They describe separate choices, rather than successive stages on a single scale. An agent can use a fixed workflow while retaining broad access to sensitive data, or plan dynamically while operating with very limited permissions.

The following questions translate those dimensions into the supplier example. Paper, §3.2 and Table 1.

DimensionQuestion for the assistant
Input trustDoes it read only selected documents, or follow arbitrary links supplied by their authors?
Access sensitivityCan it retrieve only the relevant estimate, or search all internal files?
WorkflowDoes software define the steps, or can the model add new operations after reading the proposal?
ActionCan it return a draft, or also send messages and modify records?
MemoryIs information retained only for this task, or reused across later assessments?
ToolAre the available tools reviewed in advance, or can the assistant discover and adopt new ones?
User interfaceDoes the result display plain text, or load remote images and interactive content?

These choices affect different parts of the attack. Reading arbitrary external content creates more opportunities for attacker input to enter. Access to more confidential records increases the information potentially exposed. Sending messages creates a route for disclosure. Persistent memory can extend the effect beyond one execution.

Consequently, “the agent uses a secure model” leaves most of the relevant design questions unanswered. Two applications using the same model can have substantially different exposure and consequences.

4. Why prompt injection is only part of the problem

The paper distinguishes external attackers, attackers who can submit inputs directly, and attackers with access to internal components. These capabilities must be stated when evaluating an attack. Directly editing an agent’s memory assumes more access than placing a malicious paragraph in a document it might retrieve. Paper, §4.1.

Several mechanisms extend beyond the document instruction in our example.

Tools can introduce both instructions and executable behavior

A tool description helps the model decide when and how to use a tool. A malicious provider can place misleading instructions in that description. The tool implementation can also perform harmful operations independently of how accurately the model follows its instructions. The paper groups these possibilities under tool poisoning.

Reviewing the description and constraining the executable code address different risks. A familiar tool name does not establish either property. Similarly, authenticating the provider identifies who supplied the tool; it does not prove that the tool is appropriate for a particular task.

Memory can preserve an earlier compromise

Suppose the assistant saves the supplier’s claimed “verification procedure” as guidance for future assessments. A later task may retrieve that guidance without retrieving the original proposal. The attacker-controlled instruction has gained persistence through the normal memory-writing process.

This illustrates why memory needs a record of where information came from and limits on how it may be used. Moving a statement into internal storage does not make its source more authoritative. This route differs from the paper’s internal-adversary case in which an attacker can modify memory directly.

Ordinary data and ordinary software still matter

An attacker may change a payment destination or supply a malicious package without writing an instruction to the model. A model may also invent a package name that an attacker subsequently registers. In these cases, relevant checks concern data validity, dependency integrity, and execution permissions. A detector looking only for instruction-like text does not cover the whole problem.

The interface itself can create another path for disclosure. If generated output embeds a remote image whose URL contains sensitive information, rendering that image may send the information to an external server. The model need not call an explicit “send data” tool. The application’s renderer performs the request. Paper, §4.2.

The consequences span confidentiality, integrity, and availability: exposing data, making unauthorized changes, or exhausting resources through repeated calls. The paper also discusses contextual security, which concerns what information is admitted into the agent’s context and what authority it receives relative to the intended task. This makes the handling of instructions and retrieved data an explicit security concern. Paper, §5.1.

5. Where defenses intervene

Section 5 of the paper covers runtime protection, architectural protections, identity and access management, and component hardening. Their roles become clearer when each is tied to a specific failure in the example.

Runtime checks: examine inputs and proposed actions

An input guardrail could flag the suspicious instruction before the model processes the proposal. An output guardrail could examine the proposed email before it is sent. For an agent, “output” includes tool calls and their arguments, as well as the answer shown to the user.

Checks may use rules, classifiers, or a combination. A rule can reject every email request during a draft-only task. A classifier can assess whether a less obviously suspicious action is consistent with the task. The latter requires a judgment about meaning and can make mistakes.

The distinction matters for assurance. A prompt-injection detector can have false negatives, where an attack passes, and false positives, where useful content is rejected. A rule that disables sending for this task does not depend on recognizing the attacker’s wording. Its protection does depend on every sending route passing through the enforcing software. Paper, §§5.2.1–5.2.2.

Information flow: track what may influence what

Information flow control associates data with security labels and restricts how it can move through the system. Two properties are relevant here:

  • The internal estimate is confidential. Data derived from it should reach only permitted destinations.
  • The supplier’s proposal is untrusted as a source of instructions. It should not determine permissions or authorize a new recipient.

These are different properties. Public data can be untrusted, and confidential data can originate from a trusted internal source.

To protect derived information, labels must follow transformations. Summarizing an internal estimate does not automatically make the summary public. A conservative implementation might label the whole resulting assessment as confidential and restrict its destinations accordingly.

The paper surveys approaches that use tracked variables and separated model roles, as well as approaches that estimate how inputs influence outputs. It also identifies a practical difficulty: propagating restrictions too broadly can prevent useful work. Relaxing a label therefore needs an explicit policy for when release is permitted. A model’s unsupported claim that a summary is safe is insufficient evidence for that decision. Paper, §5.2.3.

Architecture: separate processing from authority

Privilege separation gives different components only the permissions each requires. A document-processing component could extract commercial terms while having no access to internal files or email. A separate component could compare those terms with the permitted estimate and return a draft through a restricted interface.

The interfaces between components remain part of the security boundary. If a less privileged component returns arbitrary text and a more privileged model treats that text as new instructions, splitting the system into two agents has not removed the original problem. Restrictions must also govern what the receiving component can do with the result.

The paper discusses architectures that separate planning from processing tool results, together with formal verification of specified properties. Any such guarantee has a scope: the modeled operations, the policy, and the assumptions about the implementation. Verifying a rule about permitted actions does not establish that every generated statement is correct. Paper, §5.3.

Identity, credentials, and component integrity

An agent needs an identifiable authority under which it accesses services. That authority can be limited to a task, particular resources, permitted operations, and a validity period. For the assessment, reading one internal estimate requires much less authority than the user’s full account provides.

Credentials should be handled by trusted execution components wherever possible, rather than included in model-visible text. Keeping a credential out of the context reduces exposure, but does not prevent misuse of a tool that already holds it. The tool still needs to check the requested operation.

Model training that improves instruction handling and review of tool code and descriptions provide additional protection. They address the reliability and integrity of individual components. The architectural and authorization checks still determine the consequences when a component behaves incorrectly. Paper, §§5.4–5.5.

6. Applying the principles to the example

The paper emphasizes least privilege, complementary defenses, and complete mediation: checking every access to a protected resource. Applied to the draft assessment, these principles suggest the following design. This is an illustrative implementation strategy, not an architecture evaluated by the authors.

  1. Fix the task’s authority. Record that the task permits reading the selected proposal and estimate, and returning a draft to the authenticated user. Supplier content cannot expand these permissions.
  2. Limit access at the service boundary. Enforce access to the selected internal record and remove email sending from this task. The underlying services must enforce these limits even if the model requests something else.
  3. Constrain external communication. Control outbound network requests, including requests from generated previews. A restriction on the email tool alone leaves other possible disclosure routes.
  4. Preserve source and sensitivity information. Retain the origin of extracted facts and treat the assessment as containing internal information. Supplier statements remain evidence to assess, rather than durable operating instructions.
  5. Bound and record execution. Enforce limits on calls, elapsed time, and resource use. Record tool requests and policy decisions with enough detail to investigate failures, while protecting sensitive log contents.

This design still allows the model to misunderstand the proposal or draft an inaccurate assessment. Those are relevant failures to evaluate. The permissions restrict which external effects those failures can cause.

If the user subsequently asks to send the assessment, that creates a new authorization decision. The system should establish the permitted recipient and content before executing the send. Where human approval is required, the approval should show the actual recipient, body, and attachments, and apply to that specific operation. A generic request to “continue” provides too little information for an informed decision.

The paper treats human review as one defense and notes its limits: frequent prompts can cause decision fatigue. It also discusses monitoring across complete executions, since a harmful sequence may consist of individually ordinary actions. Neither human review nor monitoring removes the need for enforceable permissions. Paper, §§5.2.4–5.2.5 and §5.6.

7. What a security evaluation must establish

A useful evaluation begins with an explicit attacker capability and a measurable unwanted effect. In our example, the attacker controls the proposal, and success means that protected information reaches an unauthorized recipient. A refusal in the final answer does not establish that no earlier tool call disclosed it.

For the illustrative design, the following tests follow from the identified boundaries:

  • Task authorization: submit a valid email request during the draft-only task and verify that the execution layer rejects it, independently of how the model produced it.
  • Alternative disclosure routes: check whether browser requests, remote previews, or other tools can transmit protected information despite the email restriction.
  • Persistence: check whether attacker-controlled material saved during one task can acquire authority in a later task.
  • Resource limits: verify that repeated retrievals or tool calls stop at the enforced budget.
  • Useful operation: verify that ordinary proposals still produce usable assessments and measure how often benign work is blocked.

These tests cover different properties. Model-based detection needs evaluation against varied and adaptive inputs. Programmatic policy enforcement needs tests of the rule, its implementation, and routes that might bypass it. Repeated executions are relevant when model decisions vary, but a passing sample does not prove the absence of an exploitable sequence.

The paper provides a framework for organizing this analysis, with open challenges in information flow, isolation, policy specification, and practical deployment. Its survey does not provide a common experimental comparison establishing which defense is best for every agent.

For a particular application, the decisive questions are concrete: what can the attacker control, which decisions can that input influence, and which effects remain possible under the enforced permissions? Answering those questions connects model behavior to the security properties of the system that actually executes the task.

Reference and scope

Juhee Kim, Wenbo Guo, and Dawn Song. SoK: Attack and Defense Landscape of Agentic AI Systems. 35th USENIX Security Symposium, 2026, pp. 6047–6066. The discussion follows the supplied paper, particularly its design dimensions (§3), attack landscape (§4), and defense analysis (§5). Its literature review covers 2023 through October 2025. The worked example and evaluation checklist are this article’s engineering interpretation of those findings.

← Back to all notesDiscuss this note ↗