Start a conversation Contact
← Research

Randomise the capacity pool before testing the AI

An AI booking or allocation policy can lift its test group's results by taking capacity from the control group. Choose an experiment boundary that contains that competition before treating the measured gain as a reason to roll out.

A service business tests an AI assistant that recommends appointment times. Customers are randomly assigned to the new assistant or the existing booking page. Conversion rises for the assistant’s customers. The commercial director prepares to buy more licences, while the operations director notices that the same appointment book is still full and customers using the old page are finding fewer convenient slots.

The model may have improved how quickly its customers secure scarce capacity without increasing the work the business can sell. That distinction decides whether the rollout pays for itself. When treatment changes what remains available to control, the experiment must account for the shared resource. Random customer assignment alone cannot establish the effect of giving the new policy to everybody.

This is a constructed situation, not a client result. The buyer decision is whether to accept a customer-level conversion test or fund an experiment at the level of the appointment pool. It applies equally to AI policies allocating stock, delivery capacity or specialist staff when their users compete for the same limited resource.

A holdout can inherit the treatment

The earlier personalisation article asks buyers to retain a customer holdout and measure incremental margin. That remains useful when one customer’s treatment leaves another’s opportunity substantially unchanged. Here the treatment can alter that opportunity directly. A faster booking consumes a slot. A different dispatch policy changes which worker is available for the next request.

Researchers tested this measurement problem in an Airbnb pricing experiment. They compared individual listing assignment with assignment of groups of similar listings. Their joint analysis found that grouping reduced interference bias in the estimated booking effect. The intervention changed platform fees, rather than introducing an AI assistant. It provides evidence that experiment boundaries can change the estimated effect in a connected market, not a forecast of gains for another business. Holtz and colleagues, published in Management Science

Interference can also spread benefits into control. DoorDash’s engineering account describes customers sharing a delivery fleet: changing demand for some customers can affect service for the rest. Its team used randomised regional time windows to compare dispatch policies. This is an account of that team’s implementation, not a universal recipe for enterprise experiments. DoorDash’s switchback design

The direction therefore needs investigation. Competition can make a treatment look better by worsening control. Shared improvements can make a useful policy look weaker because control benefits too. A narrow confidence interval does not settle either problem if the comparison answers the wrong rollout question.

Follow a booking through the shared calendar

Before choosing a test, run a small deterministic exercise with the booking team. Give the assistant and incumbent page access to the same test calendar. Record the slots visible to an untreated customer. Let an assisted customer reserve a desirable slot, then repeat the untreated customer’s search. The change in available choices demonstrates a path by which treatment can affect control. It does not measure how often that path matters in production.

Repeat with separate calendars and with ample spare capacity. The purpose is to distinguish the existence of competition from its likely commercial size. Review historical availability, abandoned searches and requests that moved to another branch or date. An apparently separate branch is a poor boundary if customers routinely substitute across it or staff are moved between branches.

Write the intended outcome before selecting the randomisation unit. For this business it could be contribution from completed appointments across a capacity pool, with cancellations, waiting time and unmet demand alongside it. Count customers who could not book. Measuring only successful bookings discards the people most likely to reveal displacement.

An AI policy can still create value with unchanged appointment volume. It might reduce cancellations, improve the mix of work or reduce handling cost. Those are separate hypotheses with their own records. The exercise prevents a shift in who gets a slot from being sold as extra capacity, while leaving those other benefits open to measurement.

Choose a boundary the service can defend

Where substitution is mostly local, consider assigning entire capacity pools to the candidate or incumbent policy. Grouping branches by administrative region is insufficient on its own. Use search, booking and staff-allocation records to show why competition mostly stays inside the proposed group, and record the cross-boundary traffic that remains. The analyst must account for assignment by group when calculating uncertainty.

Where a resource can return to a comparable state, a switchback can be useful: apply one policy to the whole pool during a randomised time block, then use the other in later assigned blocks. Agree the schedule before seeing outcomes and account for recurring demand patterns. A simple before-and-after comparison can confuse the policy with a staffing change or a seasonal peak.

The difficult condition is persistence. An appointment booked during treatment may occupy capacity far into a control period. Switching the interface does not reset the calendar. Research on switchback design and analysis explicitly models how long a treatment continues to affect later outcomes. Its design results depend on those assumptions. They do not supply a safe switching interval for an appointment business.

For long-lived reservations, separate pools may be more credible than rapid switching. If the business cannot identify enough independent pools, or cannot wait for effects to clear, the buyer may need a narrower claim or a different trial. A larger count of individual customer visits cannot repair a missing independent comparison.

Test the policy across the resource it reallocates Fig. 01
Baseline
The incumbent booking policy, with demand, available appointments, completed work and cancellations recorded by capacity pool.
Outcome
Contribution from completed appointments across each assigned pool and observation window, including operating costs.
Guardrails
Unmet demand, waiting time, cancellations, staff load, cross-pool substitution and reservations carried into later periods.
Decision rule
Expand only when a comparison that accounts for shared capacity supports an economically worthwhile gain without breaching service limits. Redesign the test if displacement or carryover makes the rollout effect unresolved.

Put the experiment design in the buying decision

The evidence pack should contain the resource map, assignment rule, treatment exposure, available capacity at each decision, completed outcomes and unresolved spillovers. Ask the analyst to estimate uncertainty using the actual assignment structure and assess whether the proposed run can distinguish a worthwhile change. Agree how to handle outages and incomplete periods before launch.

The cited evidence concerns marketplace experiments and statistical design. Applying it to an enterprise booking assistant is an inference from the shared capacity mechanism. Neither a rehearsal nor a favourable trial proves a lasting return after competitors, customers or staffing respond to a full rollout.

The commercial director should approve an experiment whose boundary matches the spending decision. If the business wants more valuable completed work from its calendar, the comparison must follow that calendar. Otherwise it risks paying to redistribute the same appointments while calling the redistribution growth.

Filed under · Method · Experiments · Interference · Capacity Inference Institute · 29 Sept 2026

Related engagement

The decision behind this article

Define the problem, requirements, feasibility, options, costs and delivery approach.

Explore AI Discovery & Scoping →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.