Design the action boundary for prompt-injection failure
Prompt injection can influence an AI system through untrusted input or retrieved content. Combine prevention and detection with permissions, approval and disclosure controls that constrain the consequences if an attack succeeds.
Trace the interfaces between applications and records before changing a system.
A prompt-injection review needs to examine the full path from untrusted content to output and action. Input checks and model instructions can contribute, but the assessment also needs to test what authority is available if the model follows malicious content.
Models can receive role-labelled messages and other mechanisms intended to distinguish instruction authority. Those mechanisms do not establish reliable isolation from malicious text in a document, message or tool result. Treat the trust boundary as something to test, with independent controls around consequential actions.
Assess prompt injection through prevention, detection and containment. The architecture should constrain the data and actions exposed when a defence fails, while model-level mitigations reduce the chance of failure. Avoid treating either layer as a complete answer.
Why the filtering answer keeps failing
OWASP has kept prompt injection at the top of its list for language model applications across both editions of the Top 10 for LLM Applications, and its own guidance is careful to describe defence in depth rather than a control that closes the class. That framing is correct and it is usually read too optimistically. Layered mitigation reduces the rate. It does not change the property, because the property is that natural language has no syntax for “treat the following as data only”.
Test both direct attacks in user messages and indirect instructions in material the system processes. Documents, web pages, invitations and support tickets can carry the latter. Select cases that exercise the actual retrieval, tool and output paths of the proposed service.
Microsoft’s 2025 EchoLeak issue in Microsoft 365 Copilot is the clean example, because it required nothing of the victim at all: a crafted email arriving in the mailbox was enough for the assistant, doing its ordinary job of reading context, to be induced to exfiltrate data it was legitimately entitled to see. The vulnerability record sits in Microsoft’s update guide as CVE-2025-32711. The assistant was not compromised. It was used, at its full existing privilege, by someone who was not its user.
That pattern has a name in security that predates all of this. It is a confused deputy: a component with legitimate authority, persuaded to exercise it on behalf of someone who has none.
The four things worth doing
- Control 01 Privilege The system acts with the requesting user’s entitlements, never with a service account that can see everything.
- Control 02 Separation Untrusted content is fetched, summarised and quarantined by a component that holds no tools, before it reaches one that does.
- Control 03 Irreversibility Every action the system can take is classified as reversible or not, and the irreversible ones require a human who can see what they are approving.
- Control 04 Egress Where output can go is constrained by policy — allowed destinations, no arbitrary URLs, no rendering of attacker-supplied links.
Constrain permissions to the task and preserve the requesting actor’s identity where appropriate. User-scoped access limits the available data, but leaking that data can still be a serious incident. Service identities need equally explicit scopes, ownership and delegation records.
Review the channels through which data can leave, including tool destinations, rendered links, image fetches and webhooks. Restrict and test them. Egress controls reduce some exfiltration paths, but do not address every harmful effect an injected instruction may cause.
The question to ask about every tool
If a stranger could choose when this tool runs and with what arguments, what is the worst outcome?
- Assess retrieval scope and disclosure controls before exposure. A read-only tool can still disclose sensitive information or consume excessive resources.
- Expose with a rate limit, an audit record and an owner. Draft, tag, schedule. The kind of action a person can undo the next morning.
- Do not expose it to a path that reads untrusted content. Put a person on it. Payments, sends, deletions, permission changes, external posts.
Apply the tool question to the actual action inventory. Check whether untrusted content can influence an effectful call and which receiving-system controls constrain it. The assessment may identify a missing approval, excessive permission or an output channel requiring a different design.
What to require before an agent reads anything a stranger wrote
What this does not tell you
No control listed here establishes complete protection. Report what the tests exercise, what attacks were resisted and what remains exposed. Prevention can reduce attack success, while containment can limit consequences. Both need reassessment when tools, models or data paths change.
It also does not mean assistants over untrusted content should not be built. They should — that is most of the useful work. It means the design review has to ask what the system can do rather than what it can be told, and those are different questions with different answers.
The architect approving the tool inventory should review the worst credible misuse of each action and the controls that prevent or contain it. Test that boundary with untrusted inputs and preserve the findings. Approval should state the allowed operating scope and remaining risks.
Permission checks surround the model
- Retrieve
- An unauthorised resource is denied before retrieval.
- Context
- Only permitted passages enter the necessary context.
- Disclose
- The answer is checked for permitted disclosure; otherwise review or withhold.
Representative structure from the design pack, not evidence of a measured deployment.