
An architecture guide for engineers building tool-using LLM agents that read untrusted content, based on the 2025 design-patterns paper, the CaMeL paper and its code, AgentDojo, adaptive-attack research and NIST and OWASP guidance reviewed in October 2026. It explains each pattern, charts CaMeL's reported utility and attack results, and gives a pattern decision tree, implementation pitfalls and a test plan.
At a glance
Key findings
- Detection and prompt-level defenses lower the rate of successful injection but do not limit what a hijacked model can do; adaptive attacks pushed spotlighting and prompt sandwiching above 95 percent attack success on AgentDojo. [5][10]
- The June 2025 design-patterns paper names six patterns that share one rule: after a model reads untrusted input, that input must not be able to trigger a consequential action. It argues the case through ten case studies, not benchmark measurements. [1]
- CaMeL fixes the program from the trusted request before any data is read, hides tool outputs from the planner, and checks provenance and reader tags against Python policies before each tool call. [2]
- In CaMeL's AgentDojo tables, benign utility fell by 3.1 to 32.0 points across six 2025 models, while successful attacks fell to between 0 and 11 of 949, with the remainder attributed to cases outside its threat model. [2]
- None of the published patterns stop text-to-text manipulation, side channels or a user approving a bad action, so pair the architecture with sink-level policy, output validation and adaptive testing. [1][2]
Filtering lowers the odds, architecture sets the blast radius
The published architectures that contain indirect prompt injection rest on one rule: once a model has read text an attacker could have written, that model must not be able to choose a consequential action. The June 2025 paper Design Patterns for Securing LLM Agents against Prompt Injections names six ways to enforce the rule: action-selector, plan-then-execute, LLM map-reduce, dual LLM, code-then-execute and context-minimization. CaMeL, from researchers at Google, Google DeepMind and ETH Zurich, is a detailed published implementation of the dual LLM and code-then-execute ideas, and it adds provenance tags on every value and a policy check before every tool call. [1][2]
Each pattern buys containment by removing capability on purpose. An action selector never sees tool results, so it cannot reason over them. A planner that commits to its steps before reading data cannot carry out a request such as "do the actions listed in this email", because the steps live inside untrusted content. CaMeL's own AgentDojo tables show the price: across six 2025 models, benign task utility fell by between 3.1 points (o4 Mini High) and 32.0 points (Gemini 2.5 Pro) compared with each provider's native tool calling, while successful attacks dropped to between 0 and 11 out of 949. [2]
Filters and prompt techniques work on a different variable. Spotlighting marks or encodes untrusted text so the model can tell it apart from instructions, and its authors report attack success falling from above 50 percent to below 2 percent on GPT-family models in their experiments. [5] In October 2025, Nasr, Carlini and colleagues ran search-based adaptive attacks against spotlighting and prompt sandwiching on AgentDojo and report attack success above 95 percent for both, where the benchmark's static attacks had shown rates as low as 1 percent; human red-teamers produced 265 and 178 successful attacks against them. The same study reports above 90 percent success against the Protect AI detector, PromptGuard and Model Armor. [10]
Standards bodies say the same thing in plainer terms. NIST AI 100-2 E2025 states that current mitigations do not offer full protection against all attacker techniques and suggests designing systems on the assumption that injection is possible whenever a model is exposed to untrusted input, for example by using multiple LLMs with different permissions. OWASP's LLM01:2025 entry says it is unclear whether fool-proof prevention methods exist. [6][7] Keep detectors in the stack, since they remove crude attacks and produce useful telemetry, but treat them as a layer that lowers the rate. The architecture decides what a successful injection can reach.
Name the untrusted inputs and the consequential sinks
Pattern selection starts with two lists. The first holds every source of text that a third party can author or alter and that can reach a model's context: fetched web pages, inbound email and calendar invitations, documents shared into a drive, support tickets, product reviews, repository files and issue comments, responses from other agents, and anything stored in memory that was derived from those. NIST's taxonomy describes indirect prompt injection as an attack enabled by resource control, where the attacker manipulates resources the system reads without ever interacting with the application, and notes that in many cases the person harmed is the primary user. [6]
The second list holds consequential sinks: actions with side effects and any path that carries data outward. Sending email, inviting external participants, moving money, writing or deleting files, pushing commits, installing packages and fetching a URL all qualify, since a query string is an exfiltration channel. The design-patterns paper extends the list to the agent's own output, which must not exfiltrate data through embedded links or plant instructions that steer later behavior. [1]
A quick screen helps. Simon Willison's lethal trifecta names three capabilities: access to private data, exposure to untrusted content, and the ability to communicate externally. His point is that when one agent combines all three, an attacker can trick it into sending the private data out. [9] If all three meet inside a single model context in your design, a filter cannot close the gap and you need one of the patterns below.
Write down your assumption about the user's own prompt. CaMeL assumes the prompt is trusted, that the user is not pasting text from an untrusted source, and that any memory has not been compromised. [2] The design-patterns paper deliberately draws no distinction between direct and indirect injection, and offers context-minimization for applications where the person typing can be hostile, such as a public customer service bot. [1] That one assumption decides whether context-minimization belongs in the design.
| Item | Trust | Reaches in a naive design |
|---|---|---|
| User's typed request | Trusted by assumption | Planning and every tool |
| Inbound email body and attachments | Untrusted | Send, reply, delete, forward |
| Calendar invitation text | Untrusted | Event creation with external guests |
| Shared document contents | Untrusted | Attachments sent to any recipient |
| Directory lookup of colleagues | Trusted if from your identity system | Recipient fields |
| Memory summary of past threads | Untrusted unless admission is controlled | Future planning |
Six patterns that constrain influence
The design-patterns paper was posted to arXiv on June 10, 2025 and revised to its third version on June 27, 2025. Its 14 authors are listed alphabetically, from Invariant Labs, IBM, EPFL, ETH Zurich, Swisscom, Google, Microsoft, the ETH AI Center, the appliedAI Institute for Europe and Kyutai, with Florian Tramèr as corresponding author. It is a design analysis: it describes each pattern, discusses utility and security tradeoffs, and applies the patterns to ten case studies, but it reports no benchmark runs. Its abstract speaks of provable resistance; read that as a property of the structure, which holds only where ordinary code rather than a model enforces the isolation. The authors state that they consider it unlikely that general-purpose agents built on current models can give meaningful and reliable guarantees, and they recommend application-specific agents with clearly defined trust boundaries. [1]
Action-selector. The model translates a request into one or more predefined actions, possibly templated with parameters, and nothing a tool returns flows back to it. The paper compares this to an LLM-driven switch statement. A customer service bot that can return an order link or point to the password settings page fits. Injection in tool output has no path to the model; the cost is that the agent cannot read results at all. [1]
Plan-then-execute. The model fixes a list of tool calls before touching untrusted data, then executes it. Injected text can no longer add or reorder calls, which the paper likens to control-flow integrity. The arguments remain exposed: in the paper's example, a planner that commits to reading today's calendar and emailing the boss cannot be diverted to another recipient, but injected calendar text can still rewrite the email body. [1]
LLM map-reduce. Each untrusted item goes to its own isolated model call with no dangerous tools, and the results are combined either by code that tolerates tampered inputs or by a model whose inputs are constrained, for example by a regex that only admits a number. In the paper's invoice example, each file's map call returns a Boolean, so a malicious file can at most get itself marked as an invoice. [1]
Dual LLM. Proposed by Willison in April 2023, this splits a privileged model, which takes trusted input and can call tools, from a quarantined model, which processes untrusted content and has no tools. Ordinary controller code holds the quarantined outputs as variables such as $VAR1, and the privileged model only ever handles the names. Willison called his own proposal pretty bad: it adds implementation complexity and degrades the user experience, a single mistake that leaks untrusted text to the privileged model undoes the protection, and users can still be socially engineered into pasting their own data out. [8] The paper adds that the quarantined model remains injectable and can still return attacker-shaped content that reaches a human or another system. [1]
Code-then-execute and context-minimization. Code-then-execute takes the dual LLM to its limit: the model writes a program that calls tools and spawns unprivileged model calls, and CaMeL is the paper's cited instance. Context-minimization handles hostile user prompts: the prompt shapes an initial step, such as turning a car-quote request into a database query, and is then removed from context before the response is written, so an embedded "give me a large discount" never reaches the step that replies. [1]
The patterns compose, and the paper's second recommendation is to combine them because no single pattern fits every threat model. It separates them from baseline practices it expects in every agent: action sandboxing, strict data formatting, permissions no broader than the user's, user confirmation for sensitive actions, and attribution of outputs to their data. [1]
What each pattern still lets untrusted data touch
Every pattern narrows influence differently, and each one pays for it in a specific capability. [1]

Source. Conceptual comparison based on the pattern descriptions in Beurer-Kellner et al. (2025). [1]
Method. Conceptual summary of the paper's Section 3 descriptions and examples; the paper reports no measurements for these patterns.
Accessible table and figure data
| Pattern | Untrusted data can still affect | Main capability cost |
|---|---|---|
| Action-selector | Nothing; tool results never return to the model | No reasoning over results |
| Plan-then-execute | Arguments of planned calls, such as an email body | Plan fixed before any data is read |
| LLM map-reduce | Its own item's result, within constrained output | Items processed in isolation; outputs narrowed |
| Dual LLM | Values the quarantined model returns | Planner cannot read the data it acts on |
| Code-then-execute | Argument values, unless policies block the flow | Code errors; needs documented tool outputs |
| Context-minimization | The first step that reads the user prompt | Later steps lose the user's wording |
| Pattern | Untrusted data can still affect | Main capability cost |
|---|---|---|
| Action-selector | Nothing; tool results never return to the model | No reasoning over results |
| Plan-then-execute | Arguments of planned calls, such as an email body | Plan fixed before any data is read |
| LLM map-reduce | Its own item's result, within constrained output | Items processed in isolation; outputs narrowed |
| Dual LLM | Values the quarantined model returns | Planner cannot read the data it acts on |
| Code-then-execute | Argument values, unless policies block the flow | Code errors; needs documented tool outputs |
| Context-minimization | The first step that reads the user prompt | Later steps lose the user's wording |
What CaMeL adds with capabilities and policies
CaMeL was first posted on March 24, 2025, and its second version, dated June 24, 2025, carries the results used here. A privileged model (the P-LLM) sees only the user's request and writes code in a restricted Python subset. A custom interpreter built on Python's ast module runs that code, calls tools and stores their outputs in variables the P-LLM never reads. When a step needs to understand untrusted text, the code calls a quarantined model (the Q-LLM), which has no tools and must answer in a schema the P-LLM declared as a Pydantic model. CaMeL adds a have_enough_information field to every schema; when the Q-LLM sets it to false, the interpreter raises NotEnoughInformationError, and the Q-LLM cannot say what is missing, because that message would itself be an injection channel. [2]
Isolation alone fixes the control flow, but a fixed plan can still send the right document to the wrong person if the address came from an attacker. CaMeL closes part of that gap with capabilities, tags on each value that record where it came from and who may read it. Provenance is the user (literals in the generated code), CaMeL itself (derived values) or a specific tool, optionally with an inner source such as an email's sender. Readers are either public or a set of users. Policies are Python functions that receive a tool name and its arguments and return allowed or denied with a reason. The paper's calendar policy allows an event if the participant list came straight from the user, and otherwise requires every participant to be able to read the event details already. The banking send_money policy is stricter and requires both the recipient and the amount to come from the user with no untrusted ancestor. In a deployed system the authors expect a denial to become a confirmation prompt rather than a hard stop. [2]
Dependencies propagate through the interpreter. A NORMAL mode tracks direct data dependencies; a STRICT mode also makes every value assigned inside an if or for block depend on the condition or iterable, which closes a class of leaks where a tool call reveals a private value without directly consuming it. The interpreter retries code generation up to 10 times on errors, and redacts any error message that depends on untrusted data before showing it to the P-LLM. [2]
Separate what is proven from what is measured. By construction, untrusted data cannot add tool calls to a program written before the data was read, and a tool call whose arguments carry disallowed provenance is blocked by deterministic code. Both statements hold only if the interpreter is correct and the policies say what you meant; the authors list formal verification of the interpreter as future work, and the released repository warns that it is a research artifact that likely contains bugs and might not be fully secure. [2][3] Utility and attack counts are empirical results on one benchmark. Microsoft's Fides planner explores the same direction with confidentiality and integrity labels, deterministic policy enforcement and primitives for selectively hiding information, also evaluated on AgentDojo. [11]
from dataclasses import dataclass
@dataclass(frozen=True)
class Value:
raw: object
sources: frozenset # for example {"user"} or {"tool:read_email"}
readers: frozenset | None # None means public
def from_user_only(value: Value) -> bool:
return value.sources <= {"user"}
def readable_by(value: Value, people: set) -> bool:
return value.readers is None or people <= value.readers
def send_email_policy(recipients: Value, body: Value) -> tuple[bool, str]:
if from_user_only(recipients):
return True, "recipients came from the user's request"
if not readable_by(body, set(recipients.raw)):
return False, "body is not readable by every recipient"
return True, "body is already shared with these recipients"
The planner never reads what the reader reads
Only the trusted request reaches the planner; untrusted text goes to a tool-less reader that returns typed values held by the interpreter. [2]

Source. Conceptual illustration based on the CaMeL design section and the dual LLM pattern. [2][8]
Method. Conceptual hand-authored illustration. It simplifies CaMeL to one read, one parse and one send.
Accessible table and figure data
| Element | What it represents |
|---|---|
| Trusted request | The user's prompt, the only input the planner sees |
| Planner | Privileged model that writes the program |
| Interpreter | Ordinary code that runs the program and tracks tags |
| Amber note | Injected instruction hidden in a tool result |
| Quarantined reader | Tool-less model that parses untrusted text |
| Typed value | Schema-checked output held as a variable |
| Policy gate | Check of provenance and readers before a tool call |
| Element | What it represents |
|---|---|
| Trusted request | The user's prompt, the only input the planner sees |
| Planner | Privileged model that writes the program |
| Interpreter | Ordinary code that runs the program and tracks tags |
| Amber note | Injected instruction hidden in a tool result |
| Quarantined reader | Tool-less model that parses untrusted text |
| Typed value | Schema-checked output held as a variable |
| Policy gate | Check of provenance and readers before a tool call |
What the defenses cost in utility
AgentDojo, the benchmark behind these numbers, was introduced in 2024 with 97 realistic tasks and 629 security test cases, and is built to be extended with new tasks, defenses and adaptive attacks. [4] CaMeL evaluates on its Workspace, Banking, Travel and Slack suites and counts 949 attacks in total. [2] The paper does not name a benchmark version in its text; the released code loads AgentDojo benchmark version v1.2 and uses AgentDojo's important_instructions attack. [3]
The comparison that matters most for this article is the one where only the defense changes. With Claude 3.5 Sonnet as the model, the paper's Tables 5 and 6 report overall utility without and with attacks for the undefended model, three AgentDojo defenses and CaMeL. Spotlighting and prompt sandwiching cost nothing measurable: spotlighting scored 92.78 percent against the undefended model's 90.72 percent. The tool filter, which commits to a tool subset up front, fell to 73.20 percent. CaMeL fell furthest, to 63.92 percent. [2]
Most of CaMeL's gap sits in one suite. On Travel it solved 25.00 percent of tasks against 75.00 percent undefended, which the authors trace to tools whose output structure is poorly documented: the P-LLM cannot see outputs, so it cannot write code that parses them. They report Travel utility for the Claude Sonnet family rising from 25 percent (3.5) to 55 percent (3.7) to 75 percent (4) with no change to CaMeL or its prompt, as newer models learned to hand unparseable data to the Q-LLM. Two other failure modes are structural: tasks where the data itself dictates the action, and Q-LLM calls that lack the context they need. [2]
Treat suite-level differences with care. The paper reports an interval with each value, and for single suites they reach plus or minus 21.47 points in Table 5. The overall figures are steadier. Under attack, the undefended model lost 3.89 points of utility (90.72 to 86.83 percent), while CaMeL's attacked runs scored slightly higher than its benign runs (66.70 against 63.92 percent), which the paper presents as CaMeL preserving utility better under attack. [2] Cost is not only accuracy: the authors measure roughly two to three times the tokens of native tool calling for the median task with Claude 3.5 Sonnet, and the text and appendix tables give slightly different medians. [2]
Heuristic defenses kept utility, isolation paid for it
With Claude 3.5 Sonnet, spotlighting matched the undefended model while CaMeL solved 63.92 percent of tasks against 90.72 percent. [2]

Source. Debenedetti et al., Defeating Prompt Injections by Design, arXiv version 2, Table 5 (utility) and Table 6 (utility under attack), Claude 3.5 Sonnet on AgentDojo. [2]
Method. Overall values copied from Tables 5 and 6 without transformation; the paper's plus or minus intervals are omitted from the plot. Released code loads AgentDojo benchmark v1.2 with the important_instructions attack. [3]
Accessible table and figure data
| Configuration | Utility without attack (%) | Utility under attack (%) |
|---|---|---|
| Undefended model | 90.72 | 86.83 |
| Spotlighting | 92.78 | 89.36 |
| Prompt sandwiching | 89.69 | 84.19 |
| Tool filter | 73.2 | 72.71 |
| CaMeL | 63.92 | 66.7 |
| Configuration | Utility without attack (%) | Utility under attack (%) |
|---|---|---|
| Undefended model | 90.72 | 86.83 |
| Spotlighting | 92.78 | 89.36 |
| Prompt sandwiching | 89.69 | 84.19 |
| Tool filter | 73.2 | 72.71 |
| CaMeL | 63.92 | 66.7 |
What the attack counts show and what they do not
Table 7 of the CaMeL paper gives successful attacks out of 949 for the same Claude 3.5 Sonnet configurations: 44 undefended, 24 with prompt sandwiching, 8 with the tool filter, 5 with spotlighting, 1 for CaMeL without policies and 0 with policies. [2] The authors caution that Claude 3.5 Sonnet was already fairly resistant to AgentDojo's default attack and that adaptive prompts have been shown to weaken it sharply, so the heuristic defenses' small counts may not hold. [2] The adaptive-attack study later measured exactly that for spotlighting and sandwiching. [10]
The newer-model tables tell a consistent story. Isolation removes most attacks for every model, while the utility cost varies from small to severe. Gemini 2.5 Flash went from 297 successful attacks to 1, but its benign utility dropped from 55.7 to 35.1 percent. o3 at high reasoning effort kept most of its utility and went from 11 attacks to 0. [2]
Read the residual counts against the threat model. Without policies, the paper finds two injection tasks still scoring as successes. One asks the agent to make a payment based on instructions in a document, which means the user delegated the decision to data. The other succeeds when the user asks about hotel reviews and the code prints them, showing the injected recommendation verbatim to the user. The authors place both outside CaMeL's threat model and say the second is not a prompt injection at all, since AgentDojo scores it by whether the hotel name appears in the output. Claude 4 Sonnet's 11 remaining successes with CaMeL all sit in the Travel suite. [2]
Three limits apply to every number here. The attacks are a fixed template, not an adaptive adversary. Counts come from one benchmark's tasks, so they say nothing about your tools or data. And they measure whether the attacker's goal was reached, not how often a defended system would ask a human to approve something. On that last point, Table 8 reports that CaMeL's policies denied a tool call on 33.87 percent of benign tasks in NORMAL mode and 53.23 percent in STRICT mode, and every denial in production would be a confirmation prompt. [2]
| Model | Utility native | Utility CaMeL | Attacks native | Attacks CaMeL |
|---|---|---|---|---|
| Claude 4 Sonnet | 86.6% | 74.2% | 75 | 11 |
| Claude 4 Sonnet (reasoning) | 83.5% | 70.1% | 75 | 11 |
| Gemini 2.5 Flash | 55.7% | 35.1% | 297 | 1 |
| Gemini 2.5 Pro | 73.2% | 41.2% | 163 | 0 |
| o3 High | 84.5% | 77.3% | 11 | 0 |
| o4 Mini High | 79.4% | 76.3% | 2 | 1 |
Successful attacks out of 949 by defense
CaMeL with policies recorded zero successful attacks; the next lowest was spotlighting with five. [2]

Source. Debenedetti et al., Defeating Prompt Injections by Design, arXiv version 2, Table 7, Claude 3.5 Sonnet on AgentDojo. [2]
Method. Overall counts copied from Table 7 without transformation; the paper's plus or minus values are omitted. Counts reflect AgentDojo's fixed important_instructions attack, not adaptive attacks. [3]
Accessible table and figure data
| Configuration | Successful attacks (of 949) |
|---|---|
| Undefended model | 44 |
| Prompt sandwiching | 24 |
| Tool filter | 8 |
| Spotlighting | 5 |
| CaMeL without policies | 1 |
| CaMeL with policies | 0 |
| Configuration | Successful attacks (of 949) |
|---|---|
| Undefended model | 44 |
| Prompt sandwiching | 24 |
| Tool filter | 8 |
| Spotlighting | 5 |
| CaMeL without policies | 1 |
| CaMeL with policies | 0 |
Choosing a pattern for a real workflow
Pick the most restrictive pattern that still serves the tasks you must support, and decide per workflow rather than per product. The questions in the decision tree follow the order in which each pattern gives up capability. The first question is whether a fixed menu of actions covers the need. The second, and the one that rules out most designs, is whether untrusted content must decide which actions run. If it must, none of the published patterns contains it, and CaMeL's authors describe this failure as inherent to the dual LLM approach; the realistic options are to limit the agent to drafting output for a person, or to require approval with enough attribution that the person can see where each instruction came from. [1][2]
Three hypothetical workflows show how the answers land. A public support bot that looks up orders and resets preferences fits action-selector plus context-minimization, which is how the paper treats its customer service case study. An accounts team that wants invoices found and summarized fits map-reduce: one isolated call per file returning a Boolean or a few typed fields, and a final step that never sees raw file text. An executive assistant that reads mail, schedules meetings and replies fits code-then-execute with provenance policies on recipients and attachments, plus confirmation for anything a policy denies; the paper's email and calendar case study weighs exactly these options. [1]
The paper's case studies also show where patterns run out. For a software engineering agent that reads third-party documentation, the safest design it describes routes untrusted docs through a quarantined model that emits only a strictly formatted API description, with limits such as method names of at most 30 characters, accepting that the agent loses examples and prose. For a medical diagnosis intermediary, it removes the patient's prompt from context before the model words the doctor's or retrieval system's answer, and goes further with a structured symptoms summary that leaves no open-ended text for an injection to ride in. [1] When a workflow keeps landing on approval prompts, that is a sign to narrow the task rather than to tune the prompt.
Choose the most restrictive pattern the workflow tolerates
Ask first whether untrusted content must choose actions; if it must, none of the published patterns contains it. [1][2]

Source. Conceptual decision aid based on the design-patterns paper and the CaMeL threat model and failure analysis. [1][2]
Method. Conceptual ordering of published patterns by the capability each removes; not derived from measurements.
Accessible table and figure data
| Question | Yes, then | No, then |
|---|---|---|
| Can a fixed menu of actions serve every request? | use action-selector | ask whether data must choose actions |
| Must untrusted content decide which actions run? | draft for a person or require attributed approval | ask how untrusted items are processed |
| Can each untrusted item be processed on its own? | use map-reduce with constrained outputs | ask where untrusted values flow |
| Can untrusted values reach consequential tool arguments? | use dual LLM or code-then-execute with policies | use plan-then-execute |
| Can the user prompt itself be hostile? | strip the prompt before the reply step | record the trusted-prompt assumption |
| Question | Yes, then | No, then |
|---|---|---|
| Can a fixed menu of actions serve every request? | use action-selector | ask whether data must choose actions |
| Must untrusted content decide which actions run? | draft for a person or require attributed approval | ask how untrusted items are processed |
| Can each untrusted item be processed on its own? | use map-reduce with constrained outputs | ask where untrusted values flow |
| Can untrusted values reach consequential tool arguments? | use dual LLM or code-then-execute with policies | use plan-then-execute |
| Can the user prompt itself be hostile? | strip the prompt before the reply step | record the trusted-prompt assumption |
Implementation details that break the guarantees
The patterns are structural claims, so they fail through plumbing. Most breakages put untrusted bytes back where the design promised they would not go, or let untrusted data choose a branch. Review an implementation against this list before trusting the label on its architecture diagram.
- Quarantined output pasted back into the planner's prompt as text. That recreates the feedback loop the dual LLM pattern exists to cut; the planner must handle references, and the controller substitutes values only at the tool boundary. [1][8]
- Free-text fields flowing into consequential arguments. Constrain Q-LLM output with schemas, enumerations and length or format checks, as the map-reduce pattern's regex example does, and validate in code before use. [1]
- Error messages that quote untrusted data. A stack trace containing a document's text is a prompt. CaMeL redacts exception content that depends on untrusted values before the P-LLM sees it. [2]
- Generated code that loops over tools and dispatches on data. The CaMeL authors show models writing code that picks an action based on an email's contents, which turns data flow back into control flow and defeats isolation without policies. [2]
- Side channels. CaMeL documents leaks through indirect dependencies, through exceptions that reveal one bit when a tool call does or does not happen, and through timing; STRICT mode and the absence of the
timemodule from the interpreter address the examples shown, and the authors do not rule out others. [2] - Tools that do not label their outputs. Capabilities only help when tools attach provenance and readers. The authors note that third-party tools rarely do, and that utility suffers when they are missing. [2]
- Output shown to the user. Text-to-text manipulation and phishing are explicit non-goals for CaMeL; its provenance graph can still label content as untrusted in the interface. [2] The quarantined model in a dual LLM design can still write attacker-shaped prose to a human. [1]
Testing and residual risk
Test the structure first, then the behavior. Structural tests are cheap and deterministic: assert that the planner's context never contains tool output bytes, that every Q-LLM response validates against its schema before use, that provenance survives each transformation the interpreter supports, and that each sink policy denies the cases it should. Write those as unit and property tests against the orchestrator, because they check the claim the architecture makes rather than the model's mood on a given day.
Behavioral testing needs two suites run against the same build: your own utility tasks, so the capability cost is measured on work you care about, and an adversarial suite. AgentDojo is designed to be extended with new tasks, defenses and adaptive attacks, which makes it a reasonable harness for an internal suite. [4] Static attack strings are not enough. The adaptive-attack study bypassed 12 recent defenses, most with success above 90 percent, and its authors argue that defenses must be evaluated against attackers who tune their strategy to the defense. [10] That study skipped plan-then-execute designs such as the tool filter and CaMeL because their control flow does not depend on the injected text, and it notes that such defenses apply to a limited set of tasks and noticeably reduce utility. Your adversarial suite should therefore focus on the paths that remain: argument values inside allowed flows, content shown to people, and policy gaps. [10]
Residual risk stays even with a perfect build. Injected text can still shape content inside an allowed flow, such as the body of an email to the right recipient. Users can approve the wrong prompt. Side channels may exist that no one has found. CaMeL's authors answer their own question of whether prompt injection is solved with a plain no, and compare their position to control-flow integrity, which was later bypassed by return-oriented programming that chains individually allowed steps. [2] Record these residuals in the release decision rather than leaving them implicit.
The order of operations that follows from the evidence: inventory untrusted sources and consequential sinks; choose the most restrictive pattern each workflow tolerates; enforce isolation in orchestrator code, never in a system prompt; add provenance-aware policies at the sinks that remain reachable; keep filters and spotlighting as a rate-reducing layer; then test structure deterministically and behavior adaptively, and repeat the adversarial suite whenever the model, tools or policies change.
Method and provenance
Source-led analysis of the design-patterns and CaMeL papers, their released code, the AgentDojo and spotlighting papers, an adaptive-attack study, NIST and OWASP guidance, with original diagrams and explicitly hypothetical examples. Chart values were transcribed from the CaMeL paper's LaTeX source and checked against its arXiv HTML rendering. Sources were reviewed on October 7, 2026.
No agent, benchmark run or live environment was used. Reported results are bounded to the cited papers, models and AgentDojo configuration; they do not predict results on other tools, data or adaptive attackers.
AI assistance. AI assisted research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Design Patterns for Securing LLM Agents against Prompt Injections (arXiv 2506.08837, version 3) arXiv (Beurer-Kellner et al.). Published . Accessed .
- Defeating Prompt Injections by Design (CaMeL, arXiv 2503.18813, version 2) arXiv (Debenedetti et al.). Published . Accessed .
- CaMeL research artifact repository (README and main.py) Google Research on GitHub. Accessed .
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents arXiv (Debenedetti et al.). Published . Accessed .
- Defending Against Indirect Prompt Injection Attacks With Spotlighting arXiv (Hines et al.). Published . Accessed .
- LLM01:2025 Prompt Injection OWASP GenAI Security Project. Accessed .
- The Dual LLM pattern for building AI assistants that can resist prompt injection Simon Willison. Published . Accessed .
- The lethal trifecta for AI agents: private data, untrusted content, and external communication Simon Willison. Published . Accessed .
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections arXiv (Nasr et al.). Published . Accessed .
- Securing AI Agents with Information-Flow Control (Fides) arXiv (Costa et al.). Published . Accessed .