Start a conversation Contact
← Research

Trace the shared dependency behind every AI fallback

A second model endpoint does not keep a business service running if both routes depend on the same identity, retrieval or cloud service. Map those dependencies and test the minimum service before accepting a fallback plan.

A service review shows two model suppliers in the continuity plan. If the first endpoint fails, traffic moves to the second. The plan receives a green mark. Then the team rehearses an outage in its own identity service. Neither route can retrieve customer records, because both enter through the same gateway. The second contract did not create a second way to complete the work.

This is a constructed test, not a claim about a particular supplier. It names the question the service owner must settle: which dependencies can stop both routes at once, and what work can still be completed when they do? Provider count is a poor substitute for an answer. The unit of resilience is the business service, from request to completed outcome.

The earlier article on capital-cycle risk asks whether a system can survive a change in inference price. The article on model retirement asks how to replace a version on a supplier’s schedule. This decision concerns a different failure: two apparently available choices can stop together during an operational disruption. A portable prompt and a model evaluation set help with migration, but they do not show that the rest of the service can continue.

Count failure domains, not contracts

The Bank of England’s analysis of AI in the financial system warns that reliance on a small number of AI service providers can create systemic operational risk, especially when rapid migration is not feasible. It describes a scenario in which an outage of vendor-provided models interrupts vital customer-facing services. That is an assessment of possible financial system effects, not evidence that any named enterprise’s fallback will fail.

Draw the path for a real piece of work. A customer request reaches an application, an identity service authenticates the worker, a retrieval system supplies permitted records, a gateway calls a model, and a queue hands the result to a person or another system. Mark the owner and location of each dependency. Then draw the alternate route, including the steps that stay the same. If both routes require the same identity provider, document store, region, network connection or approval queue, that component remains a single failure domain even though the model names differ.

Ask suppliers about the layers the contract does not display. Two products may run on different infrastructure, or may rely on a common cloud, identity or data provider. The buyer should seek evidence rather than infer independence from branding. Where a supplier will not disclose an upstream dependency, record the uncertainty and test the outage scenarios that can be simulated at the buyer’s boundary. A contract for an alternate endpoint is useful only when the route to it can be invoked under the disruption it is meant to cover.

The July 2026 announcement of UK oversight for designated critical third parties states that the regime complements existing duties for regulated firms to manage their own third-party arrangements and contingency plans. Oversight of a supplier therefore does not answer whether a particular customer’s manual or alternate route works. Applicability and legal interpretation belong with the firm’s counsel. The practical question for any buyer is still what happens to its service when a shared dependency is absent.

Rehearse the minimum service

Start with the outcome the organisation must preserve. It might be the ability to receive and triage a customer request, send a time-sensitive notice or complete a review with human judgement. An AI feature can be unavailable while the business service continues. Conversely, a healthy model endpoint is not useful if the request, permitted evidence or final handoff cannot reach it.

For each important workflow, name the minimum acceptable mode and the person who can invoke it. That may be an independently hosted model, a reduced capability using fewer data sources, a human queue or a temporary pause on a non-essential feature. State the trade-off in accuracy, throughput, staff load and customer communication. A manual route is not a fallback until someone has shown that it can take work from the failed digital queue and return completed cases to the normal record.

The Bank’s operational resilience guidance asks the firms it regulates to identify important business services, set impact tolerances, map their dependencies and test whether the services can continue through severe but plausible disruption. Those are sector-specific expectations, not a rule imposed on every AI buyer. The method is still a useful way to avoid mistaking a spare supplier for a tested service outcome.

The fallback test ends when work is completed Fig. 01
Baseline
Completed cases, elapsed time, review effort and customer communication under the normal route.
Outcome
Cases completed to an acceptable standard while a model provider and then a shared upstream dependency are unavailable.
Guardrails
Permission boundaries, unreviewed effects, stranded queue items and staff capacity during the degraded mode.
Decision rule
Accept the fallback only if the minimum service stays within the organisation's stated tolerance in both disruptions.

Test the routes separately. First remove the primary model endpoint and observe whether the second route completes work. Then remove one shared dependency that the map exposed, such as the gateway or retrieval source, and observe whether the minimum service survives. Reconcile the queue afterwards so a recovered system does not repeat an action a human completed during the outage. A rehearsal that only returns a successful model response stops before the decision that matters.

The result should change spending. If the shared gateway is the dominant failure, buying a third model endpoint is a poor use of a resilience budget. If the manual route cannot absorb demand, training and process capacity may be more useful than another contract. If the second provider truly has a separate path and can meet the service tolerance, keep it and test the switch regularly. The map makes those choices visible to the service owner and finance team.

What the map cannot prove

An internal drill cannot establish every supplier’s hidden infrastructure or predict a sector-wide event. Supplier disclosures may be incomplete, and simulated load may differ from demand during a real incident. Record those limits beside the test result. The evidence supports a decision about the specified scenarios and current configuration, which must be revisited when a provider or internal architecture changes.

The service owner should sign off on the work the organisation can still finish, the dependencies that can stop it and the next investment needed to close the gap. A list of model vendors belongs inside that decision, where it can be tested against the service it is meant to protect.

What this does not tell you

Non-negotiable wherever the piece touches regulation, and reliably the most credible paragraph in anything published here. Say what is outside the claim.

Close on the consequence, and name the person who decides differently now.

Filed under · Governance · Supplier concentration · Operational resilience · AI services Inference Institute · 28 Sept 2026

Related engagement

The decision behind this article

A structured assessment of one AI system covering risks, impacts, controls and residual risk.

Explore AI Risk & Impact Assessment →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.