Put the AI copilot beside the newest service agents first
A service copilot creates value by shortening the path from new starter to capable agent. Start where experience is scarce, measure resolved work and protect quality before extending it to everybody.
A customer-service director has three related problems. The queue is growing, new agents take time to become effective, and the most experienced people spend part of each day answering questions for everyone else. The proposed AI business case usually begins with a fourth problem: how many agents the technology might replace.
That framing misses the more immediate source of value. The useful system sits inside the conversation, retrieves what the organisation already knows, and suggests a diagnosis or response while the agent remains responsible for the customer. Its first job is to make scarce experience available at the moment a newer colleague needs it.
Put the AI copilot beside the newest service agents first. The commercial case is a shorter path to competent work, more issues resolved with the team already employed, and fewer avoidable escalations. A broad rollout can hide that effect because experienced agents have less to gain and may have something to lose.
The value is in the experience curve
Support work contains more judgement than its scripts suggest. An experienced agent recognises the symptom behind an imprecise question, knows which diagnostic step to ask for next, finds the relevant product instruction and phrases the answer in a way that keeps the conversation moving. Much of that knowledge is visible in past conversations but absent from the formal playbook.
A copilot can retrieve and recombine those patterns. It does not need to replace the agent to change the economics. If a new starter resolves more work without lowering quality, capacity rises. If fewer conversations reach a supervisor, the most experienced people recover time. If competence arrives earlier, the organisation spends less of each hiring cycle below its required service level.
The strongest field evidence supports that mechanism. Brynjolfsson, Li and Raymond studied a staggered deployment of a generative AI assistant across 5,172 customer-support agents. Access increased issues resolved per hour by 15 per cent on average, with the largest gains among less experienced and lower-skilled workers. Agents with two months of tenure and AI assistance performed as well as or better than untreated agents with more than six months of tenure. The most experienced agents saw small gains in speed and small declines in quality, which is why the average is not a rollout plan. Generative AI at Work, March 2026
The system in that study monitored chats and suggested responses. Agents could accept, edit or ignore them. The result was therefore produced by a redesigned human workflow, not by a chatbot taking the customer relationship away from the service team.
| Before assistance | With assistance designed into the workflow |
|---|---|
| A new agent searches several systems while the customer waits | The copilot retrieves a relevant procedure and proposes the next diagnostic step |
| An experienced colleague is interrupted for a familiar question | The agent sees a suggestion drawn from approved knowledge and prior successful patterns |
| Coaching happens after a supervisor reviews a poor conversation | Guidance appears during the conversation, when it can still change the outcome |
| Every agent receives the same rollout and the same target | Access and expectations differ by tenure, issue type and measured benefit |
The last row protects the business case. A tool that materially helps a new starter and distracts an expert can look mediocre when their results are averaged together. Segmenting the rollout is part of discovering the value, not an analysis performed after the pilot.
Start with a queue, not a licence count
Choose one service queue where the organisation can observe outcomes. Prefer a queue with repeatable issue families, usable knowledge, enough completed conversations to establish a baseline and a genuine experience gap between new and established agents. Avoid beginning with the most emotionally charged or irreversible decisions merely because they appear valuable.
Map the conversation before selecting a model. Identify where an agent diagnoses the issue, retrieves information, drafts a response, takes an action and closes the case. The copilot should enter at a named step. A suggestion that appears outside the system where the agent works creates another window to read and another record to reconcile.
- 01 Baseline Measure the existing queue by tenure, issue type, resolution, time and escalation.
- 02 Assist Give one defined cohort suggestions inside the live workflow while retaining agent control.
- 03 Compare Measure resolved work and quality against a credible concurrent or phased control.
- 04 Decide Expand, narrow, redesign or stop according to the result for each cohort.
The baseline prevents saved time from becoming a story nobody can verify. The control prevents a quieter week or a change in issue mix from being credited to the technology. Random assignment is strongest where it is practical. A phased rollout can still be useful when the comparison accounts for tenure, queue and calendar effects.
Measure capacity and the cost of obtaining it
Issues resolved per hour is the obvious capacity measure. It is insufficient on its own. A fast conversation that causes a repeat contact has moved work rather than removed it. A suggestion that closes a ticket while leaving the customer wrong has made the dashboard better and the service worse.
Track resolution, handle time, repeat contact, escalation, abandonment and customer-rated quality together. Add time to proficiency for new starters and retention by cohort if the pilot lasts long enough to support either measure. Record whether suggestions were accepted, edited or ignored. That tells the service designer where the copilot is useful and where it is merely present.
A 2026 randomised field experiment in Alibaba’s after-sales service operation offers a useful caution. Its assistant proposed issue diagnoses and customer responses. It improved service speed and subjective quality overall, but did not produce a significant improvement in the objective measure of customers trying again for help. Top-performing agents showed little speed improvement and declines in quality, which the researchers linked to increased multitasking. Generative AI in Action, February 2026
That finding changes the implementation question. The goal is not maximum tool use. It is better service economics. If high performers ignore the copilot and continue to lead the queue, that can be the correct design. Their conversations may be more valuable as governed learning material for the system than as a place to require adoption.
What the Inference Institute can help decide
The useful engagement is not a model comparison in isolation. It is a scoped service design and evaluation: select the queue, establish the baseline, define the point where assistance enters, connect approved knowledge, specify what the agent may do with a suggestion and design the comparison that determines whether the workflow improved.
The output should let a customer-service director make a funded decision. It should show which cohort benefits, which issue families qualify, what additional capacity appeared, whether quality held, what operating controls are required and which result would stop the rollout. The organisation can then procure or build against a measured workflow rather than buying licences and searching for value afterwards.
What this does not tell you
Two field studies do not establish a universal return. They concern particular companies, support tasks, tools and periods. The first used a staggered deployment rather than individual random assignment. The second found different effects depending on the performance of the agent and on the quality measure used. Neither supplies a forecast for another organisation.
The argument also does not assume that faster service reduces headcount. Demand may absorb the added capacity. The organisation may choose shorter waits, wider coverage, more coaching or better retention instead. Those are different commercial outcomes and should be named before the pilot begins.
The customer-service director should make the next decision at cohort level. Start with the agents whose experience gap the copilot can plausibly close, in one queue with observable outcomes. If they reach competent performance sooner without creating repeat work, the business case has evidence. If they do not, the organisation has learned where not to spend before a broad rollout turns licence adoption into its only measure of progress.