Request a scoping call Contact
← Research

Hold out the customer before funding the rollout

A model tested on fresh records from familiar customers may still be untested on new accounts. Keep customer groups separate when the spending decision depends on that transfer, and report performance for the population the rollout will actually serve.

A service director is buying a system that routes incoming customer requests. The supplier has trained it on historical tickets and tested it on tickets kept out of training. The result looks strong enough to support a smaller manual triage queue. The proposed rollout, however, includes newly acquired accounts. Many customers in the supplier’s test also appeared in its training data.

That overlap can change what the result means. Familiar customers bring familiar products, wording and recurring problems. Recognising those patterns may be useful for serving them again, while saying little about the accounts the business is about to onboard. Funding a staffing change on the wrong population can leave new customers waiting and experienced staff repairing misrouted work.

When the business case depends on serving unseen customers, the evaluation must include customers absent from training. The data owner should define that boundary before the supplier produces another aggregate score. The decision is whether the evidence supports expansion to new accounts, continued use with established accounts, or more observation before either commitment.

Name what will be new

A ticket can be new while almost everything that makes it predictable is familiar. Another message in the same conversation is the clearest case. A separate request from the same account can also carry a recognisable product configuration or a recurring service problem. Removing the account-number column does not remove every signal associated with the account.

The scikit-learn guide to grouped validation explains how dependent observations from the same group can cross a conventional split. To assess performance on unseen groups, its grouped splitters keep a group’s records entirely on one side of each training and validation comparison. This is a method for constructing the comparison. It does not choose the right business grouping or prove that the test represents future demand.

The mechanism has been studied in a different setting. Saeb and colleagues compared record-wise and subject-wise validation using activity-recognition data and simulation. They found that mixing subjects across training and test could produce optimistic estimates for a model intended to work on new subjects. The study concerns sensor and clinical prediction settings. Applying its identity mechanism to customer tickets is an inference, not a reported customer-service result.

Start with a sentence the service director can approve: the system will route future requests from existing accounts, first requests from new accounts, or both. Those are different acceptance populations. Preserve separate results when the rollout includes both, along with their expected shares of incoming work. A large established-account workload should not conceal a weak result for the smaller population driving the expansion decision.

Make the split fail a simple rehearsal

Before running another model, test the dataset assignment with a constructed fixture. Give fictional accounts named Alder, Birch, Cedar and Dune an earlier and a later ticket. Within this deliberately simplified fixture, each account always uses the same queue. Give a lookup rule the earlier tickets and ask it to route the later ones by remembering the account’s previous queue.

The rule recognises every account in that test. Move Cedar and Dune entirely out of training and it must abstain on their tickets. Nothing about the routing rule improved or deteriorated. The test changed from repeat business to unseen accounts. This is a demonstration of the evidence boundary, not a measurement of a learned model or a prediction of its accuracy.

Run an overlap check on the real evaluation extract next. Compare account keys, conversation keys and source-document identifiers across the split. Review aliases introduced by CRM migrations and merged accounts. A new identifier can still refer to the same customer. Keep uncertain matches visible instead of claiming independence from a clean-looking key comparison.

Choose the grouping at the level the deployment needs to cross. An account boundary tests new accounts. A corporate-family boundary may be necessary when subsidiaries share distinctive configurations and the intended claim concerns new customer groups. Holding out an entire service centre asks a broader question again. Record why the chosen boundary matches the spending decision rather than selecting whichever grouping yields the most favourable score.

Keep identity and time in the same acceptance plan

Grouping is not a substitute for chronology. The earlier article on historical feature availability asks what the system could know at prediction time. An account-disjoint test could still train on records collected after the test period. Conversely, a correctly timed test can contain familiar accounts throughout. A deployment to new accounts next quarter needs evidence about both changes.

Ask for a dated training cut-off and a later evaluation cohort whose account membership is explicit. Keep all information used to fit the candidate inside the permitted training boundary, including learned preprocessing and feature selection. The scikit-learn leakage guidance explains why fitting those transformations before splitting can expose the test data even when the final model sees only its training rows.

Compare the candidate with the incumbent routing process on the same eligible cases. Report misrouting, manual handling and unresolved requests separately for established and unseen accounts. Preserve the model version, prediction time, available inputs, account assignment and adjudicated destination. Where a request can legitimately reach several teams, keep that ambiguity in the review record rather than manufacture a single clean answer.

Evidence for expanding customer coverage Fig. 01
Baseline
The incumbent routing process on the same later-period requests, separated into established and unseen accounts.
Outcome
Correct routing and manual triage effort for the population the proposed rollout will add.
Guardrails
Misrouted urgent work, unresolved requests, repeat contacts, account overlap and uneven results across customer groups.
Decision rule
Expand to new accounts only when their own comparison meets the agreed service and operating-cost thresholds. Retain a narrower deployment when the evidence supports established accounts alone.

Buy the coverage the evidence supports

A grouped test is not automatically the best estimate for every service. The published perspectives accompanying Saeb’s paper examine distribution mismatch, variability and the limits of treating subject-wise validation as a universal remedy. Holding out atypical groups can create a comparison unlike the population actually expected. Too few independent accounts can also leave a result too uncertain for a staffing decision.

The practical response is to describe the target intake, inspect how the held-out accounts differ and report uncertainty at the account level. More tickets from the same accounts do not establish wider customer coverage. Nor does a lower score after regrouping prove that identity overlap caused the whole difference. The population and amount of training evidence may have changed too.

Offline routing quality still cannot establish realised savings. A bounded service trial must show what happens to completed work, staffing effort and customer outcomes. Keep the manual route available while that evidence develops.

Before signing the rollout, the service director should ask which customers the result can speak for. If the answer is existing accounts, fund that scope on its own merits. Let the new-account commitment wait for evidence drawn across the boundary the business is paying the system to cross.

Filed under · Data · Evaluation · Customer data · Generalisation Inference Institute · 02 Oct 2026

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.