Put a decision gate on live AI testing
A financial-services AI trial should enter real-world conditions with an agreed outcome, customer guardrails and authority to stop. A mature proof of concept is the start of that decision, not the evidence that settles it.
A financial-services team has a proof of concept that sorts customer messages well in a test set. The proposed next step is to expose it to live traffic. The project board sees a model score and a demonstration. It has not agreed which customers will see the system, which errors require intervention or who can suspend it. The question is not whether the demonstration was convincing. It is whether the organisation can learn safely from real use.
Put a decision gate between proof of concept and live testing. The gate names the intended customer outcome, the cohort, the evidence to collect, the guardrails and the person who can stop the test. Without it, a pilot can run for months and still leave the board unable to decide whether to expand, redesign or retire the service.
What the FCA is actually testing
The FCA describes AI Live Testing as a programme for firms with mature proofs of concept ready for controlled market conditions. It distinguishes an AI system from an AI model and considers deployment context, governance, human involvement, evaluation, and input and output controls. Its programme moves through discovery, framework validation and AI system testing. Those are features of the FCA initiative, not a general-purpose approval pathway for every firm’s product.
The FCA’s AI sandbox serves an earlier kind of work: experimenting and developing AI propositions in a controlled environment. The two routes illustrate a maturity question a firm can ask independently of applying to either programme. Are the product and controls sufficiently defined to learn from real-world conditions, or is the organisation still discovering whether the use case is viable?
For the customer-message system, a sandbox may help establish that the model can process the data and support the desired task. Live testing asks whether the entire service behaves acceptably with actual queues, staff, customers and exceptions. That includes which messages it does not handle, how it passes a case to a human and what customers experience while an error is corrected.
Write a release decision before launch
The decision owner should specify the cohort and exposure boundary. Start with the kinds of messages, channels and customer groups for which the assessment is valid. Exclusions are as important as inclusions: a tool tested on routine questions has not thereby been tested on complaints, vulnerable customers or urgent loss reports.
Next, record the incumbent process. Measure response time, accuracy, repeat contacts, complaint escalation and manual effort before the AI path changes them. Choose one primary outcome and guardrails that represent customer harm, not only operational savings. Keep a comparison for long enough to observe the relevant outcomes, and preserve the ability to attribute a decision to the system version and human action that produced it.
Name stop triggers that a frontline team can execute. A concerning error pattern, missing trace, drift in the case mix or failure of the handoff may require a pause before the next governance meeting. Give a named role the technical ability to switch to the manual route. Rehearse that route once before the test begins.
- Baseline
- Response time, correct routing, repeat contacts and manual effort for the same message classes under the current process.
- Outcome
- Faster resolution of routine messages at comparable or better routing accuracy and customer experience.
- Guardrails
- Missed complaints or urgent cases, harm to vulnerable customers, unsupported answers, missing traces and failed human handoffs.
- Decision rule
- Start with routine messages only. Pause on a guardrail breach, investigate the trace and resume only after the owner accepts the remedy and a retest.
Make the first live slice small enough to inspect. A team could use the AI path for one routine message class while sending complaints and urgent loss reports through the existing route. Sample both automated and escalated cases each day. Record the reason for every handoff and any case that should have been handed off but was not. If the case mix changes, the results from the original slice should not be used to justify extending it without a new assessment.
The pause route needs its own test. Who can disable the AI path outside office hours? Does the customer message reach a staffed queue immediately, or sit in a failed integration? Can the team identify which customers received an affected answer before the switch? A drill may reveal that the switch exists but the trace cannot be retrieved. That is a release blocker even if the model’s test score remains high.
At the gate, distinguish a controlled continuation from expansion. The board may have enough evidence to keep learning with the same cohort while lacking evidence for new customer groups or message types. Document that boundary, the next observation period and the person accountable for deciding again.
At the review gate, present both numbers and cases. Aggregate performance may hide a small group experiencing poor outcomes. A collection of anecdotes may not show whether the service improved. The board needs enough of each to decide what is known, where it is uncertain and which condition must change before wider release.
What this does not establish
Participation in an FCA programme does not certify a firm’s AI system or remove its regulatory responsibilities. Nor does this article say every firm must join one. The particular obligations depend on the service and the firm. Legal interpretation stays with counsel and the relevant compliance team.
The accountable executive should be able to sign a test plan with a credible stop decision, rather than merely approve another pilot. That is how live use becomes evidence for deployment instead of a prolonged demonstration.