Request a scoping call Contact
← Research

Design the action boundary for prompt-injection failure

Prompt injection can influence an AI system through untrusted input or retrieved content. Combine prevention and detection with permissions, approval and disclosure controls that constrain the consequences if an attack succeeds.

Architecture / Conceptual study
Make the connections explicit.
  1. Applications
  2. Interfaces
  3. Records

Trace the interfaces between applications and records before changing a system.

A prompt-injection review needs to examine the full path from untrusted content to output and action. Input checks and model instructions can contribute, but the assessment also needs to test what authority is available if the model follows malicious content.

Models can receive role-labelled messages and other mechanisms intended to distinguish instruction authority. Those mechanisms do not establish reliable isolation from malicious text in a document, message or tool result. Treat the trust boundary as something to test, with independent controls around consequential actions.

Assess prompt injection through prevention, detection and containment. The architecture should constrain the data and actions exposed when a defence fails, while model-level mitigations reduce the chance of failure. Avoid treating either layer as a complete answer.

Why the filtering answer keeps failing

OWASP has kept prompt injection at the top of its list for language model applications across both editions of the Top 10 for LLM Applications, and its own guidance is careful to describe defence in depth rather than a control that closes the class. That framing is correct and it is usually read too optimistically. Layered mitigation reduces the rate. It does not change the property, because the property is that natural language has no syntax for “treat the following as data only”.

Test both direct attacks in user messages and indirect instructions in material the system processes. Documents, web pages, invitations and support tickets can carry the latter. Select cases that exercise the actual retrieval, tool and output paths of the proposed service.

Microsoft’s 2025 EchoLeak issue in Microsoft 365 Copilot is the clean example, because it required nothing of the victim at all: a crafted email arriving in the mailbox was enough for the assistant, doing its ordinary job of reading context, to be induced to exfiltrate data it was legitimately entitled to see. The vulnerability record sits in Microsoft’s update guide as CVE-2025-32711. The assistant was not compromised. It was used, at its full existing privilege, by someone who was not its user.

That pattern has a name in security that predates all of this. It is a confused deputy: a component with legitimate authority, persuaded to exercise it on behalf of someone who has none.

The four things worth doing

Architecture controls that complement model-level defences Fig. 01
  1. Control 01 Privilege The system acts with the requesting user’s entitlements, never with a service account that can see everything.
  2. Control 02 Separation Untrusted content is fetched, summarised and quarantined by a component that holds no tools, before it reaches one that does.
  3. Control 03 Irreversibility Every action the system can take is classified as reversible or not, and the irreversible ones require a human who can see what they are approving.
  4. Control 04 Egress Where output can go is constrained by policy — allowed destinations, no arbitrary URLs, no rendering of attacker-supplied links.

Constrain permissions to the task and preserve the requesting actor’s identity where appropriate. User-scoped access limits the available data, but leaking that data can still be a serious incident. Service identities need equally explicit scopes, ownership and delegation records.

Review the channels through which data can leave, including tool destinations, rendered links, image fetches and webhooks. Restrict and test them. Egress controls reduce some exfiltration paths, but do not address every harmful effect an injected instruction may cause.

The question to ask about every tool

How to decide whether a tool can be exposed to untrusted content Fig. 02

If a stranger could choose when this tool runs and with what arguments, what is the worst outcome?

  • Nothing leaves and nothing changes Assess retrieval scope and disclosure controls before exposure. A read-only tool can still disclose sensitive information or consume excessive resources.
  • Something changes, but it can be reversed and it is logged Expose with a rate limit, an audit record and an owner. Draft, tag, schedule. The kind of action a person can undo the next morning.
  • Something leaves the boundary, or cannot be undone Do not expose it to a path that reads untrusted content. Put a person on it. Payments, sends, deletions, permission changes, external posts.

Apply the tool question to the actual action inventory. Check whether untrusted content can influence an effectful call and which receiving-system controls constrain it. The assessment may identify a missing approval, excessive permission or an output channel requiring a different design.

What to require before an agent reads anything a stranger wrote

What this does not tell you

No control listed here establishes complete protection. Report what the tests exercise, what attacks were resisted and what remains exposed. Prevention can reduce attack success, while containment can limit consequences. Both need reassessment when tools, models or data paths change.

It also does not mean assistants over untrusted content should not be built. They should — that is most of the useful work. It means the design review has to ask what the system can do rather than what it can be told, and those are different questions with different answers.

The architect approving the tool inventory should review the worst credible misuse of each action and the controls that prevent or contain it. Test that boundary with untrusted inputs and preserve the findings. Approval should state the allowed operating scope and remaining risks.

illustrative diagram

Permission checks surround the model

Permission checks surround the modelRetrieve: An unauthorised resource is denied before retrieval. Context: Only permitted passages enter the necessary context. Disclose: The answer is checked for permitted disclosure; otherwise review or withhold.NoYesYesNoUser requestAuthorised for resource?Deny retrievalRetrieve permitted passagesSelect necessary contextModel endpointAnswer disclosure permitted?Return sourced answerReview or withhold
Permission checks surround the modelRetrieve: An unauthorised resource is denied before retrieval. Context: Only permitted passages enter the necessary context. Disclose: The answer is checked for permitted disclosure; otherwise review or withhold.YesUser requestResource permissionNo: deny retrievalYes: continueRetrieve permittedpassagesSelect necessarycontextModel endpointAnswer disclosureYes: sourced answerNo: review or withhold
Retrieve
An unauthorised resource is denied before retrieval.
Context
Only permitted passages enter the necessary context.
Disclose
The answer is checked for permitted disclosure; otherwise review or withhold.

Representative structure from the design pack, not evidence of a measured deployment.

Filed under · Architecture · Security · Agents · Architecture Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.