Review tool definitions as a supply-chain dependency
Tool definitions are supplier-controlled inputs to an agent’s context. Review their provenance, permissions and changes as part of the software supply chain, and test whether untrusted text can influence an authorised action.
Trace the interfaces between applications and records before changing a system.
Connecting an agent to a ticketing platform, warehouse or calendar introduces more than an API endpoint. The integration also supplies tool names, descriptions and schemas that the model reads when choosing an action. Those inputs need an owner and a review boundary.
A tool description can carry malicious instructions alongside legitimate usage guidance. The model may follow that content even though it should have lower trust than the application’s instructions. The practical risk is that supplier-controlled text influences a component able to request actions with delegated permissions.
Treat tool definitions as a supply-chain dependency. Record their source, approved version and permissions, and detect material changes. Natural-language descriptions are not ordinary executable code, but their influence on tool selection can still create an operational attack path.
The attack class, named
Security research distinguishes several related patterns. They overlap in practice, but naming the mechanism helps specify which trust boundary and control a test should exercise.
- 01 Tool poisoning Instructions hidden in a tool description or schema the model reads.
- 02 Tool shadowing A server redefines or overrides a tool another server provides.
- 03 Rug pull A definition that is benign at review time and changes after approval.
- 04 Confused deputy The agent uses authority it legitimately holds on behalf of whoever supplied the text.
The third stage is the one that defeats a one-off security review, and it is the reason this is a supply chain problem rather than a code review problem. A remote tool server can serve a different definition tomorrow than it served on the day you assessed it. Unless the definitions are pinned and their changes are detected, the assessment describes a moment that has passed.
Microsoft’s 2026 MCP security review describes risks across the protocol ecosystem. The MCPTox benchmark evaluates tool poisoning against real tool-server environments. These sources establish relevant attack mechanisms and test settings, rather than the prevalence of compromise in a particular organisation.
Why this is different from an ordinary dependency
Review both the text’s influence on the model and the permissions available when the model acts. Familiar dependency controls remain useful, but they need to account for this additional instruction boundary.
A compromised software dependency can also misuse the permissions of its process. Tool poisoning reaches that authority through a different route: manipulating the model’s interpretation of descriptions or returned content. The agent’s context should therefore preserve trust distinctions, and receiving tools should enforce permissions independently of the model.
Maintain an inventory of connected servers, approved definitions and responsible owners. Versioning, change review and monitoring help establish what the agent was presented with at the time. Where a service cannot be pinned, define how changes are detected and which changes block further use pending review.
An output shape does not grant authority
Hypothetical assistant proposes sending a customer reminder.
Swipe or scroll for the full diagram →
action: send_reminder; account: A-12; reviewed_version: 4; route: review. This is a displayed record, not executable code.
- Eligibility
- The receiving tool checks actor permission and the reviewed account version.
- Confidence
- No calibrated probability is claimed in this fixture. A model score alone cannot grant permission.
- Bounded route
- Only an authorised, current-state decision reaches the bounded application action.
- Abstention
- Missing permission, version conflict or insufficient evidence returns review or abstention.
- Evidence limit
- Valid shape is not proof of factual truth, policy compliance or a safe business decision.
Constructed routing fixture. No model is called and no external action is performed.
Text and typed outputs both need policy and state checks before any effect.
Reviewed 2026-10-02
What to require before connecting anything
Choose a definition-update policy that the environment can enforce. Pinning can support reproducibility, while dynamic services may require comparison with an approved snapshot. In either case, a changed description should not silently inherit approval granted to a different version.
Use permissions appropriate to the requesting task and preserve the relationship between the user, agent and receiving service. Delegated access can limit the available scope, but it does not establish safe use of everything within that scope. Service identities also need constrained permissions, clear ownership and records of the task they represent.
The procurement version of the question
For a purchased agent product, the equivalent questions are contractual rather than technical, and there are four: which tool servers does this connect to, who controls their definitions, how am I notified when a definition changes, and what does the product do when a tool returns content that contains instructions.
Ask the supplier to explain how returned content is marked as untrusted and how it is prevented from granting new authority. Separating content processing from effectful tools can reduce exposure, but its effectiveness depends on the full data flow. Require a demonstrated failure test and stated residual risks rather than one preferred architecture described as a universal answer.
What this does not tell you
These controls reduce exposure and constrain possible effects. They do not establish that a model will always distinguish instructions from data. Assess prevention, detection and containment together, and report the scope of any attack testing rather than treating a filter result as complete protection.
Tool access can make an agent useful for operational work. Introduce it through the organisation’s dependency and access-control processes, with tests matched to the consequences of each action. Read-only retrieval and irreversible effects may need different approval and recovery rules.
The supply-chain owner can begin by reviewing the connected-server inventory. For each connection, establish who controls the definitions, how changes are handled and what actions the agent can request. Missing answers identify the next assessment task.
A typed decision still crosses gates
- Schema
- Malformed output is rejected.
- Policy
- Out-of-scope or uncertain decisions reach human review or abstention.
- Effect
- Record and assess the actual outcome of a bounded application action.
Representative structure from the design pack, not evidence of a measured deployment.