Skip to content
Cloud Security DeskSearch
Menu

Technical guideAI systems

Design AI agents that contain indirect prompt injection

Filters lower the odds that an agent obeys injected text. Architecture decides what an obeying agent can reach. Compare six published patterns and CaMeL by what each removes, what it costs and what it still misses.

Published
Sources checked
Next review
Reading time
17 minutes
Coverage
Google DeepMind · AgentDojo · NIST · OWASP · Microsoft
A white document with a small amber note folded into its corner travels along a path toward two workspaces divided by a solid wall. The upper workspace receives a separate blue request and holds instruction cards; the document is routed into the lower workspace, which is fully enclosed.
Conceptual illustration: untrusted text, and the instruction hidden inside it, goes to a sealed workspace, while planning happens in a separate space that sees only the trusted request.

An architecture guide for engineers building tool-using LLM agents that read untrusted content, based on the 2025 design-patterns paper, the CaMeL paper and its code, AgentDojo, adaptive-attack research and NIST and OWASP guidance reviewed in October 2026. It explains each pattern, charts CaMeL's reported utility and attack results, and gives a pattern decision tree, implementation pitfalls and a test plan.

At a glance

Key findings

  • Detection and prompt-level defenses lower the rate of successful injection but do not limit what a hijacked model can do; adaptive attacks pushed spotlighting and prompt sandwiching above 95 percent attack success on AgentDojo. [5][10]
  • The June 2025 design-patterns paper names six patterns that share one rule: after a model reads untrusted input, that input must not be able to trigger a consequential action. It argues the case through ten case studies, not benchmark measurements. [1]
  • CaMeL fixes the program from the trusted request before any data is read, hides tool outputs from the planner, and checks provenance and reader tags against Python policies before each tool call. [2]
  • In CaMeL's AgentDojo tables, benign utility fell by 3.1 to 32.0 points across six 2025 models, while successful attacks fell to between 0 and 11 of 949, with the remainder attributed to cases outside its threat model. [2]
  • None of the published patterns stop text-to-text manipulation, side channels or a user approving a bad action, so pair the architecture with sink-level policy, output validation and adaptive testing. [1][2]

Filtering lowers the odds, architecture sets the blast radius

The published architectures that contain indirect prompt injection rest on one rule: once a model has read text an attacker could have written, that model must not be able to choose a consequential action. The June 2025 paper Design Patterns for Securing LLM Agents against Prompt Injections names six ways to enforce the rule: action-selector, plan-then-execute, LLM map-reduce, dual LLM, code-then-execute and context-minimization. CaMeL, from researchers at Google, Google DeepMind and ETH Zurich, is a detailed published implementation of the dual LLM and code-then-execute ideas, and it adds provenance tags on every value and a policy check before every tool call. [1][2]

Each pattern buys containment by removing capability on purpose. An action selector never sees tool results, so it cannot reason over them. A planner that commits to its steps before reading data cannot carry out a request such as "do the actions listed in this email", because the steps live inside untrusted content. CaMeL's own AgentDojo tables show the price: across six 2025 models, benign task utility fell by between 3.1 points (o4 Mini High) and 32.0 points (Gemini 2.5 Pro) compared with each provider's native tool calling, while successful attacks dropped to between 0 and 11 out of 949. [2]

Filters and prompt techniques work on a different variable. Spotlighting marks or encodes untrusted text so the model can tell it apart from instructions, and its authors report attack success falling from above 50 percent to below 2 percent on GPT-family models in their experiments. [5] In October 2025, Nasr, Carlini and colleagues ran search-based adaptive attacks against spotlighting and prompt sandwiching on AgentDojo and report attack success above 95 percent for both, where the benchmark's static attacks had shown rates as low as 1 percent; human red-teamers produced 265 and 178 successful attacks against them. The same study reports above 90 percent success against the Protect AI detector, PromptGuard and Model Armor. [10]

Standards bodies say the same thing in plainer terms. NIST AI 100-2 E2025 states that current mitigations do not offer full protection against all attacker techniques and suggests designing systems on the assumption that injection is possible whenever a model is exposed to untrusted input, for example by using multiple LLMs with different permissions. OWASP's LLM01:2025 entry says it is unclear whether fool-proof prevention methods exist. [6][7] Keep detectors in the stack, since they remove crude attacks and produce useful telemetry, but treat them as a layer that lowers the rate. The architecture decides what a successful injection can reach.

Name the untrusted inputs and the consequential sinks

Pattern selection starts with two lists. The first holds every source of text that a third party can author or alter and that can reach a model's context: fetched web pages, inbound email and calendar invitations, documents shared into a drive, support tickets, product reviews, repository files and issue comments, responses from other agents, and anything stored in memory that was derived from those. NIST's taxonomy describes indirect prompt injection as an attack enabled by resource control, where the attacker manipulates resources the system reads without ever interacting with the application, and notes that in many cases the person harmed is the primary user. [6]

The second list holds consequential sinks: actions with side effects and any path that carries data outward. Sending email, inviting external participants, moving money, writing or deleting files, pushing commits, installing packages and fetching a URL all qualify, since a query string is an exfiltration channel. The design-patterns paper extends the list to the agent's own output, which must not exfiltrate data through embedded links or plant instructions that steer later behavior. [1]

A quick screen helps. Simon Willison's lethal trifecta names three capabilities: access to private data, exposure to untrusted content, and the ability to communicate externally. His point is that when one agent combines all three, an attacker can trick it into sending the private data out. [9] If all three meet inside a single model context in your design, a filter cannot close the gap and you need one of the patterns below.

Write down your assumption about the user's own prompt. CaMeL assumes the prompt is trusted, that the user is not pasting text from an untrusted source, and that any memory has not been compromised. [2] The design-patterns paper deliberately draws no distinction between direct and indirect injection, and offers context-minimization for applications where the person typing can be hostile, such as a public customer service bot. [1] That one assumption decides whether context-minimization belongs in the design.

Hypothetical inventory for an email and calendar assistant, showing the trust label and the sinks each source could reach in a naive single-model design.
ItemTrustReaches in a naive design
User's typed requestTrusted by assumptionPlanning and every tool
Inbound email body and attachmentsUntrustedSend, reply, delete, forward
Calendar invitation textUntrustedEvent creation with external guests
Shared document contentsUntrustedAttachments sent to any recipient
Directory lookup of colleaguesTrusted if from your identity systemRecipient fields
Memory summary of past threadsUntrusted unless admission is controlledFuture planning

Six patterns that constrain influence

The design-patterns paper was posted to arXiv on June 10, 2025 and revised to its third version on June 27, 2025. Its 14 authors are listed alphabetically, from Invariant Labs, IBM, EPFL, ETH Zurich, Swisscom, Google, Microsoft, the ETH AI Center, the appliedAI Institute for Europe and Kyutai, with Florian Tramèr as corresponding author. It is a design analysis: it describes each pattern, discusses utility and security tradeoffs, and applies the patterns to ten case studies, but it reports no benchmark runs. Its abstract speaks of provable resistance; read that as a property of the structure, which holds only where ordinary code rather than a model enforces the isolation. The authors state that they consider it unlikely that general-purpose agents built on current models can give meaningful and reliable guarantees, and they recommend application-specific agents with clearly defined trust boundaries. [1]

Action-selector. The model translates a request into one or more predefined actions, possibly templated with parameters, and nothing a tool returns flows back to it. The paper compares this to an LLM-driven switch statement. A customer service bot that can return an order link or point to the password settings page fits. Injection in tool output has no path to the model; the cost is that the agent cannot read results at all. [1]

Plan-then-execute. The model fixes a list of tool calls before touching untrusted data, then executes it. Injected text can no longer add or reorder calls, which the paper likens to control-flow integrity. The arguments remain exposed: in the paper's example, a planner that commits to reading today's calendar and emailing the boss cannot be diverted to another recipient, but injected calendar text can still rewrite the email body. [1]

LLM map-reduce. Each untrusted item goes to its own isolated model call with no dangerous tools, and the results are combined either by code that tolerates tampered inputs or by a model whose inputs are constrained, for example by a regex that only admits a number. In the paper's invoice example, each file's map call returns a Boolean, so a malicious file can at most get itself marked as an invoice. [1]

Dual LLM. Proposed by Willison in April 2023, this splits a privileged model, which takes trusted input and can call tools, from a quarantined model, which processes untrusted content and has no tools. Ordinary controller code holds the quarantined outputs as variables such as $VAR1, and the privileged model only ever handles the names. Willison called his own proposal pretty bad: it adds implementation complexity and degrades the user experience, a single mistake that leaks untrusted text to the privileged model undoes the protection, and users can still be socially engineered into pasting their own data out. [8] The paper adds that the quarantined model remains injectable and can still return attacker-shaped content that reaches a human or another system. [1]

Code-then-execute and context-minimization. Code-then-execute takes the dual LLM to its limit: the model writes a program that calls tools and spawns unprivileged model calls, and CaMeL is the paper's cited instance. Context-minimization handles hostile user prompts: the prompt shapes an initial step, such as turning a car-quote request into a database query, and is then removed from context before the response is written, so an embedded "give me a large discount" never reaches the step that replies. [1]

The patterns compose, and the paper's second recommendation is to combine them because no single pattern fits every threat model. It separates them from baseline practices it expects in every agent: action sandboxing, strict data formatting, permissions no broader than the user's, user confirmation for sensitive actions, and attribution of outputs to their data. [1]

Figure 01

What each pattern still lets untrusted data touch

Every pattern narrows influence differently, and each one pays for it in a specific capability. [1]

Matrix of six design patterns with three columns: what untrusted data can still affect, and the main capability cost. Action-selector leaves nothing after selection; plan-then-execute leaves arguments of planned calls; map-reduce leaves each item's own constrained result; dual LLM leaves quarantined return values; code-then-execute leaves argument values unless policies block them; context-minimization leaves the first step that reads the prompt.

Source. Conceptual comparison based on the pattern descriptions in Beurer-Kellner et al. (2025). [1]

Method. Conceptual summary of the paper's Section 3 descriptions and examples; the paper reports no measurements for these patterns.

Accessible table and figure data
Figure 1 accessible table
PatternUntrusted data can still affectMain capability cost
Action-selectorNothing; tool results never return to the modelNo reasoning over results
Plan-then-executeArguments of planned calls, such as an email bodyPlan fixed before any data is read
LLM map-reduceIts own item's result, within constrained outputItems processed in isolation; outputs narrowed
Dual LLMValues the quarantined model returnsPlanner cannot read the data it acts on
Code-then-executeArgument values, unless policies block the flowCode errors; needs documented tool outputs
Context-minimizationThe first step that reads the user promptLater steps lose the user's wording
Figure 1 accessible table
PatternUntrusted data can still affectMain capability cost
Action-selectorNothing; tool results never return to the modelNo reasoning over results
Plan-then-executeArguments of planned calls, such as an email bodyPlan fixed before any data is read
LLM map-reduceIts own item's result, within constrained outputItems processed in isolation; outputs narrowed
Dual LLMValues the quarantined model returnsPlanner cannot read the data it acts on
Code-then-executeArgument values, unless policies block the flowCode errors; needs documented tool outputs
Context-minimizationThe first step that reads the user promptLater steps lose the user's wording

What CaMeL adds with capabilities and policies

CaMeL was first posted on March 24, 2025, and its second version, dated June 24, 2025, carries the results used here. A privileged model (the P-LLM) sees only the user's request and writes code in a restricted Python subset. A custom interpreter built on Python's ast module runs that code, calls tools and stores their outputs in variables the P-LLM never reads. When a step needs to understand untrusted text, the code calls a quarantined model (the Q-LLM), which has no tools and must answer in a schema the P-LLM declared as a Pydantic model. CaMeL adds a have_enough_information field to every schema; when the Q-LLM sets it to false, the interpreter raises NotEnoughInformationError, and the Q-LLM cannot say what is missing, because that message would itself be an injection channel. [2]

Isolation alone fixes the control flow, but a fixed plan can still send the right document to the wrong person if the address came from an attacker. CaMeL closes part of that gap with capabilities, tags on each value that record where it came from and who may read it. Provenance is the user (literals in the generated code), CaMeL itself (derived values) or a specific tool, optionally with an inner source such as an email's sender. Readers are either public or a set of users. Policies are Python functions that receive a tool name and its arguments and return allowed or denied with a reason. The paper's calendar policy allows an event if the participant list came straight from the user, and otherwise requires every participant to be able to read the event details already. The banking send_money policy is stricter and requires both the recipient and the amount to come from the user with no untrusted ancestor. In a deployed system the authors expect a denial to become a confirmation prompt rather than a hard stop. [2]

Dependencies propagate through the interpreter. A NORMAL mode tracks direct data dependencies; a STRICT mode also makes every value assigned inside an if or for block depend on the condition or iterable, which closes a class of leaks where a tool call reveals a private value without directly consuming it. The interpreter retries code generation up to 10 times on errors, and redacts any error message that depends on untrusted data before showing it to the P-LLM. [2]

Separate what is proven from what is measured. By construction, untrusted data cannot add tool calls to a program written before the data was read, and a tool call whose arguments carry disallowed provenance is blocked by deterministic code. Both statements hold only if the interpreter is correct and the policies say what you meant; the authors list formal verification of the interpreter as future work, and the released repository warns that it is a research artifact that likely contains bugs and might not be fully secure. [2][3] Utility and attack counts are empirical results on one benchmark. Microsoft's Fides planner explores the same direction with confidentiality and integrity labels, deterministic policy enforcement and primitives for selectively hiding information, also evaluated on AgentDojo. [11]

Example fragment for Python 3.10 or later: a hypothetical send_email policy written in the style CaMeL describes. The Value type and helpers are illustrative and are not CaMeL's API.
from dataclasses import dataclass


@dataclass(frozen=True)
class Value:
    raw: object
    sources: frozenset          # for example {"user"} or {"tool:read_email"}
    readers: frozenset | None   # None means public


def from_user_only(value: Value) -> bool:
    return value.sources <= {"user"}


def readable_by(value: Value, people: set) -> bool:
    return value.readers is None or people <= value.readers


def send_email_policy(recipients: Value, body: Value) -> tuple[bool, str]:
    if from_user_only(recipients):
        return True, "recipients came from the user's request"
    if not readable_by(body, set(recipients.raw)):
        return False, "body is not readable by every recipient"
    return True, "body is already shared with these recipients"
Figure 02

The planner never reads what the reader reads

Only the trusted request reaches the planner; untrusted text goes to a tool-less reader that returns typed values held by the interpreter. [2]

Illustration. A trusted request enters a planner box, which writes code for an interpreter. The interpreter calls a read tool that returns a document with a hidden amber note. The document goes to a quarantined reader with no tools, which returns a typed value stored as a variable. A policy gate checks the variable's tags before the send tool runs. A barrier separates the planner from the untrusted side.

Source. Conceptual illustration based on the CaMeL design section and the dual LLM pattern. [2][8]

Method. Conceptual hand-authored illustration. It simplifies CaMeL to one read, one parse and one send.

Accessible table and figure data
Figure 2 accessible table
ElementWhat it represents
Trusted requestThe user's prompt, the only input the planner sees
PlannerPrivileged model that writes the program
InterpreterOrdinary code that runs the program and tracks tags
Amber noteInjected instruction hidden in a tool result
Quarantined readerTool-less model that parses untrusted text
Typed valueSchema-checked output held as a variable
Policy gateCheck of provenance and readers before a tool call
Figure 2 accessible table
ElementWhat it represents
Trusted requestThe user's prompt, the only input the planner sees
PlannerPrivileged model that writes the program
InterpreterOrdinary code that runs the program and tracks tags
Amber noteInjected instruction hidden in a tool result
Quarantined readerTool-less model that parses untrusted text
Typed valueSchema-checked output held as a variable
Policy gateCheck of provenance and readers before a tool call

What the defenses cost in utility

AgentDojo, the benchmark behind these numbers, was introduced in 2024 with 97 realistic tasks and 629 security test cases, and is built to be extended with new tasks, defenses and adaptive attacks. [4] CaMeL evaluates on its Workspace, Banking, Travel and Slack suites and counts 949 attacks in total. [2] The paper does not name a benchmark version in its text; the released code loads AgentDojo benchmark version v1.2 and uses AgentDojo's important_instructions attack. [3]

The comparison that matters most for this article is the one where only the defense changes. With Claude 3.5 Sonnet as the model, the paper's Tables 5 and 6 report overall utility without and with attacks for the undefended model, three AgentDojo defenses and CaMeL. Spotlighting and prompt sandwiching cost nothing measurable: spotlighting scored 92.78 percent against the undefended model's 90.72 percent. The tool filter, which commits to a tool subset up front, fell to 73.20 percent. CaMeL fell furthest, to 63.92 percent. [2]

Most of CaMeL's gap sits in one suite. On Travel it solved 25.00 percent of tasks against 75.00 percent undefended, which the authors trace to tools whose output structure is poorly documented: the P-LLM cannot see outputs, so it cannot write code that parses them. They report Travel utility for the Claude Sonnet family rising from 25 percent (3.5) to 55 percent (3.7) to 75 percent (4) with no change to CaMeL or its prompt, as newer models learned to hand unparseable data to the Q-LLM. Two other failure modes are structural: tasks where the data itself dictates the action, and Q-LLM calls that lack the context they need. [2]

Treat suite-level differences with care. The paper reports an interval with each value, and for single suites they reach plus or minus 21.47 points in Table 5. The overall figures are steadier. Under attack, the undefended model lost 3.89 points of utility (90.72 to 86.83 percent), while CaMeL's attacked runs scored slightly higher than its benign runs (66.70 against 63.92 percent), which the paper presents as CaMeL preserving utility better under attack. [2] Cost is not only accuracy: the authors measure roughly two to three times the tokens of native tool calling for the median task with Claude 3.5 Sonnet, and the text and appendix tables give slightly different medians. [2]

Figure 03

Heuristic defenses kept utility, isolation paid for it

With Claude 3.5 Sonnet, spotlighting matched the undefended model while CaMeL solved 63.92 percent of tasks against 90.72 percent. [2]

Grouped bar chart of AgentDojo overall utility with Claude 3.5 Sonnet. Without attack and under attack: undefended 90.72 and 86.83 percent; spotlighting 92.78 and 89.36; prompt sandwiching 89.69 and 84.19; tool filter 73.20 and 72.71; CaMeL 63.92 and 66.70.

Source. Debenedetti et al., Defeating Prompt Injections by Design, arXiv version 2, Table 5 (utility) and Table 6 (utility under attack), Claude 3.5 Sonnet on AgentDojo. [2]

Method. Overall values copied from Tables 5 and 6 without transformation; the paper's plus or minus intervals are omitted from the plot. Released code loads AgentDojo benchmark v1.2 with the important_instructions attack. [3]

Accessible table and figure data
Figure 3 accessible table
ConfigurationUtility without attack (%)Utility under attack (%)
Undefended model90.7286.83
Spotlighting92.7889.36
Prompt sandwiching89.6984.19
Tool filter73.272.71
CaMeL63.9266.7
Figure 3 accessible table
ConfigurationUtility without attack (%)Utility under attack (%)
Undefended model90.7286.83
Spotlighting92.7889.36
Prompt sandwiching89.6984.19
Tool filter73.272.71
CaMeL63.9266.7

What the attack counts show and what they do not

Table 7 of the CaMeL paper gives successful attacks out of 949 for the same Claude 3.5 Sonnet configurations: 44 undefended, 24 with prompt sandwiching, 8 with the tool filter, 5 with spotlighting, 1 for CaMeL without policies and 0 with policies. [2] The authors caution that Claude 3.5 Sonnet was already fairly resistant to AgentDojo's default attack and that adaptive prompts have been shown to weaken it sharply, so the heuristic defenses' small counts may not hold. [2] The adaptive-attack study later measured exactly that for spotlighting and sandwiching. [10]

The newer-model tables tell a consistent story. Isolation removes most attacks for every model, while the utility cost varies from small to severe. Gemini 2.5 Flash went from 297 successful attacks to 1, but its benign utility dropped from 55.7 to 35.1 percent. o3 at high reasoning effort kept most of its utility and went from 11 attacks to 0. [2]

Read the residual counts against the threat model. Without policies, the paper finds two injection tasks still scoring as successes. One asks the agent to make a payment based on instructions in a document, which means the user delegated the decision to data. The other succeeds when the user asks about hotel reviews and the code prints them, showing the injected recommendation verbatim to the user. The authors place both outside CaMeL's threat model and say the second is not a prompt injection at all, since AgentDojo scores it by whether the hotel name appears in the output. Claude 4 Sonnet's 11 remaining successes with CaMeL all sit in the Travel suite. [2]

Three limits apply to every number here. The attacks are a fixed template, not an adaptive adversary. Counts come from one benchmark's tasks, so they say nothing about your tools or data. And they measure whether the attacker's goal was reached, not how often a defended system would ask a human to approve something. On that last point, Table 8 reports that CaMeL's policies denied a tool call on 33.87 percent of benign tasks in NORMAL mode and 53.23 percent in STRICT mode, and every denial in production would be a confirmation prompt. [2]

Overall utility without attack and successful attacks out of 949, native tool calling versus CaMeL with policies, transcribed from Tables 2 and 4 of the CaMeL paper (arXiv version 2). Claude 4 Sonnet with reasoning is marked as in the paper. [2]
ModelUtility nativeUtility CaMeLAttacks nativeAttacks CaMeL
Claude 4 Sonnet86.6%74.2%7511
Claude 4 Sonnet (reasoning)83.5%70.1%7511
Gemini 2.5 Flash55.7%35.1%2971
Gemini 2.5 Pro73.2%41.2%1630
o3 High84.5%77.3%110
o4 Mini High79.4%76.3%21
Figure 04

Successful attacks out of 949 by defense

CaMeL with policies recorded zero successful attacks; the next lowest was spotlighting with five. [2]

Bar chart of successful attacks out of 949 on AgentDojo with Claude 3.5 Sonnet: undefended 44, prompt sandwiching 24, tool filter 8, spotlighting 5, CaMeL without policies 1, CaMeL with policies 0.

Source. Debenedetti et al., Defeating Prompt Injections by Design, arXiv version 2, Table 7, Claude 3.5 Sonnet on AgentDojo. [2]

Method. Overall counts copied from Table 7 without transformation; the paper's plus or minus values are omitted. Counts reflect AgentDojo's fixed important_instructions attack, not adaptive attacks. [3]

Accessible table and figure data
Figure 4 accessible table
ConfigurationSuccessful attacks (of 949)
Undefended model44
Prompt sandwiching24
Tool filter8
Spotlighting5
CaMeL without policies1
CaMeL with policies0
Figure 4 accessible table
ConfigurationSuccessful attacks (of 949)
Undefended model44
Prompt sandwiching24
Tool filter8
Spotlighting5
CaMeL without policies1
CaMeL with policies0

Choosing a pattern for a real workflow

Pick the most restrictive pattern that still serves the tasks you must support, and decide per workflow rather than per product. The questions in the decision tree follow the order in which each pattern gives up capability. The first question is whether a fixed menu of actions covers the need. The second, and the one that rules out most designs, is whether untrusted content must decide which actions run. If it must, none of the published patterns contains it, and CaMeL's authors describe this failure as inherent to the dual LLM approach; the realistic options are to limit the agent to drafting output for a person, or to require approval with enough attribution that the person can see where each instruction came from. [1][2]

Three hypothetical workflows show how the answers land. A public support bot that looks up orders and resets preferences fits action-selector plus context-minimization, which is how the paper treats its customer service case study. An accounts team that wants invoices found and summarized fits map-reduce: one isolated call per file returning a Boolean or a few typed fields, and a final step that never sees raw file text. An executive assistant that reads mail, schedules meetings and replies fits code-then-execute with provenance policies on recipients and attachments, plus confirmation for anything a policy denies; the paper's email and calendar case study weighs exactly these options. [1]

The paper's case studies also show where patterns run out. For a software engineering agent that reads third-party documentation, the safest design it describes routes untrusted docs through a quarantined model that emits only a strictly formatted API description, with limits such as method names of at most 30 characters, accepting that the agent loses examples and prose. For a medical diagnosis intermediary, it removes the patient's prompt from context before the model words the doctor's or retrieval system's answer, and goes further with a structured symptoms summary that leaves no open-ended text for an injection to ride in. [1] When a workflow keeps landing on approval prompts, that is a sign to narrow the task rather than to tune the prompt.

Figure 05

Choose the most restrictive pattern the workflow tolerates

Ask first whether untrusted content must choose actions; if it must, none of the published patterns contains it. [1][2]

Decision tree with five questions: whether a fixed menu of actions covers every request; whether untrusted content must decide which actions run; whether each untrusted item can be processed independently; whether untrusted values can reach consequential tool arguments; and whether the user prompt can be hostile.

Source. Conceptual decision aid based on the design-patterns paper and the CaMeL threat model and failure analysis. [1][2]

Method. Conceptual ordering of published patterns by the capability each removes; not derived from measurements.

Accessible table and figure data
Figure 5 accessible table
QuestionYes, thenNo, then
Can a fixed menu of actions serve every request?use action-selectorask whether data must choose actions
Must untrusted content decide which actions run?draft for a person or require attributed approvalask how untrusted items are processed
Can each untrusted item be processed on its own?use map-reduce with constrained outputsask where untrusted values flow
Can untrusted values reach consequential tool arguments?use dual LLM or code-then-execute with policiesuse plan-then-execute
Can the user prompt itself be hostile?strip the prompt before the reply steprecord the trusted-prompt assumption
Figure 5 accessible table
QuestionYes, thenNo, then
Can a fixed menu of actions serve every request?use action-selectorask whether data must choose actions
Must untrusted content decide which actions run?draft for a person or require attributed approvalask how untrusted items are processed
Can each untrusted item be processed on its own?use map-reduce with constrained outputsask where untrusted values flow
Can untrusted values reach consequential tool arguments?use dual LLM or code-then-execute with policiesuse plan-then-execute
Can the user prompt itself be hostile?strip the prompt before the reply steprecord the trusted-prompt assumption

Implementation details that break the guarantees

The patterns are structural claims, so they fail through plumbing. Most breakages put untrusted bytes back where the design promised they would not go, or let untrusted data choose a branch. Review an implementation against this list before trusting the label on its architecture diagram.

  • Quarantined output pasted back into the planner's prompt as text. That recreates the feedback loop the dual LLM pattern exists to cut; the planner must handle references, and the controller substitutes values only at the tool boundary. [1][8]
  • Free-text fields flowing into consequential arguments. Constrain Q-LLM output with schemas, enumerations and length or format checks, as the map-reduce pattern's regex example does, and validate in code before use. [1]
  • Error messages that quote untrusted data. A stack trace containing a document's text is a prompt. CaMeL redacts exception content that depends on untrusted values before the P-LLM sees it. [2]
  • Generated code that loops over tools and dispatches on data. The CaMeL authors show models writing code that picks an action based on an email's contents, which turns data flow back into control flow and defeats isolation without policies. [2]
  • Side channels. CaMeL documents leaks through indirect dependencies, through exceptions that reveal one bit when a tool call does or does not happen, and through timing; STRICT mode and the absence of the time module from the interpreter address the examples shown, and the authors do not rule out others. [2]
  • Tools that do not label their outputs. Capabilities only help when tools attach provenance and readers. The authors note that third-party tools rarely do, and that utility suffers when they are missing. [2]
  • Output shown to the user. Text-to-text manipulation and phishing are explicit non-goals for CaMeL; its provenance graph can still label content as untrusted in the interface. [2] The quarantined model in a dual LLM design can still write attacker-shaped prose to a human. [1]

Testing and residual risk

Test the structure first, then the behavior. Structural tests are cheap and deterministic: assert that the planner's context never contains tool output bytes, that every Q-LLM response validates against its schema before use, that provenance survives each transformation the interpreter supports, and that each sink policy denies the cases it should. Write those as unit and property tests against the orchestrator, because they check the claim the architecture makes rather than the model's mood on a given day.

Behavioral testing needs two suites run against the same build: your own utility tasks, so the capability cost is measured on work you care about, and an adversarial suite. AgentDojo is designed to be extended with new tasks, defenses and adaptive attacks, which makes it a reasonable harness for an internal suite. [4] Static attack strings are not enough. The adaptive-attack study bypassed 12 recent defenses, most with success above 90 percent, and its authors argue that defenses must be evaluated against attackers who tune their strategy to the defense. [10] That study skipped plan-then-execute designs such as the tool filter and CaMeL because their control flow does not depend on the injected text, and it notes that such defenses apply to a limited set of tasks and noticeably reduce utility. Your adversarial suite should therefore focus on the paths that remain: argument values inside allowed flows, content shown to people, and policy gaps. [10]

Residual risk stays even with a perfect build. Injected text can still shape content inside an allowed flow, such as the body of an email to the right recipient. Users can approve the wrong prompt. Side channels may exist that no one has found. CaMeL's authors answer their own question of whether prompt injection is solved with a plain no, and compare their position to control-flow integrity, which was later bypassed by return-oriented programming that chains individually allowed steps. [2] Record these residuals in the release decision rather than leaving them implicit.

The order of operations that follows from the evidence: inventory untrusted sources and consequential sinks; choose the most restrictive pattern each workflow tolerates; enforce isolation in orchestrator code, never in a system prompt; add provenance-aware policies at the sinks that remain reachable; keep filters and spotlighting as a rate-reducing layer; then test structure deterministically and behavior adaptively, and repeat the adversarial suite whenever the model, tools or policies change.

Method and provenance

Source-led analysis of the design-patterns and CaMeL papers, their released code, the AgentDojo and spotlighting papers, an adaptive-attack study, NIST and OWASP guidance, with original diagrams and explicitly hypothetical examples. Chart values were transcribed from the CaMeL paper's LaTeX source and checked against its arXiv HTML rendering. Sources were reviewed on October 7, 2026.

No agent, benchmark run or live environment was used. Reported results are bounded to the cited papers, models and AgentDojo configuration; they do not predict results on other tools, data or adaptive attackers.

AI assistance. AI assisted research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.

Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.

References

  1. Design Patterns for Securing LLM Agents against Prompt Injections (arXiv 2506.08837, version 3) arXiv (Beurer-Kellner et al.). Published . Accessed .
  2. Defeating Prompt Injections by Design (CaMeL, arXiv 2503.18813, version 2) arXiv (Debenedetti et al.). Published . Accessed .
  3. CaMeL research artifact repository (README and main.py) Google Research on GitHub. Accessed .
  4. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents arXiv (Debenedetti et al.). Published . Accessed .
  5. Defending Against Indirect Prompt Injection Attacks With Spotlighting arXiv (Hines et al.). Published . Accessed .
  6. LLM01:2025 Prompt Injection OWASP GenAI Security Project. Accessed .
  7. The Dual LLM pattern for building AI assistants that can resist prompt injection Simon Willison. Published . Accessed .
  8. The lethal trifecta for AI agents: private data, untrusted content, and external communication Simon Willison. Published . Accessed .
  9. Securing AI Agents with Information-Flow Control (Fides) arXiv (Costa et al.). Published . Accessed .