Start a conversation Contact
← Research

Automate the interview script, not the hiring decision

AI voice interviews can make high-volume screening more consistent, but the business case depends on who continues, who starts and who stays. Automate information collection while people retain the hiring decision.

A head of recruitment is trying to fill a service operation. Recruiters repeat the same first interview all day, each following the same guidance with a slightly different order, emphasis and level of detail.

The cost is not only the minutes spent speaking. Different interviews produce different evidence about similar applicants, good candidates wait for a slot, and hiring managers inherit a signal whose quality depends on who happened to be available that afternoon.

Automate the interview script, not the hiring decision. A voice agent can collect a more consistent first-round record while a human still evaluates the evidence and decides who receives an offer. The commercial test is whether the workflow fills capable people sooner and keeps them, after counting the applicants the format causes to leave.

Consistency can change the hiring economics

Jabarian and Henkel ran a natural field experiment with 70,884 applications for entry-level customer-service roles. Applicants who passed an initial screen were assigned to a human recruiter, an AI voice agent or a choice between the two. Human recruiters made every final hiring decision. Jabarian and Henkel, 2026

The AI-interview group received offers in 9.73 per cent of cases, compared with 8.70 per cent for human interviews. Applicants assigned to the AI agent were 18 per cent more likely to start the job and 18 per cent more likely to remain employed for at least one month. The study found no meaningful difference in handling time, customer satisfaction or employer quality scores for the workers who were hired. Jabarian and Henkel, 2026

The authors connect the result to a specific mechanism. The agent followed the same topic order more closely, covered a more consistent set of topics and used more standardised prompts and follow-ups. That reduced interviewer-driven variation while preserving responses to the individual applicant.

The result is a design clue rather than a licence-count argument. If the first interview is a noisy information-collection step, making it more comparable can improve the signal reaching a recruiter. The organisation still needs to show that the signal predicts the work it is hiring people to do.

Where an AI interview can change the workflow without owning the hiring decision Fig. 01
  1. 01 Invite Tell candidates what the interview is for, who will review it and what happens next.
  2. 02 Collect Use one approved question path with room for relevant follow-ups.
  3. 03 Evaluate Let a trained recruiter review the same evidence and apply the hiring standard.
  4. 04 Decide Compare offers, starts, retention and work quality with the existing process.

The format can spend the value before evaluation begins

Avery and colleagues randomised more than 3,000 applicants into asynchronous audio interviews, asynchronous video interviews, live online interviews or no screening. They measured whether applicants continued and how the resulting interviews were assessed. Avery and colleagues, 2026

Continuation fell by about 53 per cent, or 45 percentage points, after an asynchronous interview. The live online format reduced continuation by about 20 per cent, or 17 percentage points, against the no-screening control. Women were 5.1 percentage points less likely than men to complete the asynchronous interview. The paper attributes the deterrence to beliefs about fairness and the number of people competing for each role. Avery and colleagues, 2026

That finding changes the rollout question. A recorded interview may reduce recruiter time per applicant and still shrink the pool before the organisation can assess anyone. The saved interview is not capacity if it removes people the business would have wanted to hire.

The same study found that an AI assessment tool scored women and underrepresented racial minorities higher than human evaluators, and that its scores were more predictive of later employment outcomes in that sample. Those results concern the evaluation of interview answers, not a universal claim that an automated assessor should make the decision. Avery and colleagues, 2026

Keep the two interventions separate in the pilot. Changing who asks the questions and changing who scores the answers create different mechanisms, risks and evidence requirements. Combining them can make a positive result impossible to explain and a negative result impossible to repair.

Measure the whole funnel, not the interview queue

The buyer for this decision is usually a chief people officer or a head of high-volume recruitment. Their outcome is not interviews completed. It is a filled role that reaches the required level of performance and remains useful long enough to justify the hiring effort.

Start with a baseline for the existing process. Record invitations, starts, completion by stage, interview duration, recruiter time, offer rate, acceptance, time to start and early retention. Segment the record by role, location, experience and any group for which continuation or selection consequences matter.

Then define the smallest operational change the pilot is meant to support. It could be more completed interviews per recruiter, a shorter interval from application to offer, fewer interviews required per start or better retention in the first month. Name the measure before the agent is introduced.

Do not hide the cost of obtaining the signal. Record failed calls, technical handoffs, candidate requests for a person, reviewer time and cases where the transcript is too incomplete to use. The Jabarian and Henkel study reports that 5 per cent of applicants ended their interview because they did not want to speak to an AI, while technical difficulties affected 7 per cent of cases. Jabarian and Henkel, 2026

Read hiring outcomes after the interview. Offer rate can rise because the threshold changed or because the evidence improved. Start rate can rise while retention falls. A useful comparison follows applicants into the work, where quality, attendance, customer results or supervisor assessment can be observed.

Put the human boundary where the evidence becomes consequential

The voice agent should collect and organise information. It should not decide that a person is acceptable because a transcript resembles a successful case. The recruiter needs the question path, the candidate’s answers, the scoring rubric and the points where the system failed or deviated.

That record also makes supplier evaluation concrete. Ask whether the agent can preserve the exact prompt and response sequence, expose unanswered questions, replay a technical failure and hand the case to a person without losing context. Ask how the system is tested when accents, interruptions, silence or a candidate asking for a human change the expected path.

The approval boundary belongs after evidence review. A recruiter can override a recommendation, record why and identify the evidence that would have changed the decision. If the system makes the final choice, the pilot is testing an automated selection policy, not an automated interview, and it needs a different governance and measurement design.

What the Inference Institute can help decide

The useful engagement starts with the hiring workflow, not a voice-model demo. It maps the current screening path, identifies where interviewer variation enters, defines the information the first interview must collect and separates candidate experience from selection quality.

The result is a pilot specification a people leader can fund and review. It names the role, question path, human decision point, control design, cohort measures, post-hire observation and stopping rule. It also makes clear whether the intended gain is recruiter capacity, faster starts, better matching or retention.

What this does not tell you

The AI voice study concerns one recruitment-process outsourcing firm, entry-level customer-service roles in the Philippines and a period when the partner’s workflow was already structured. Its retention measure is one month in a high-turnover market. It does not forecast another role, country or hiring standard.

The asynchronous-interview study concerns three technology jobs and a platform used by the participating employer. Its AI assessment findings come from one commercial tool and a selected set of interview answers. Neither paper establishes that a particular vendor, script or scoring threshold will improve your hiring.

The decision owner should begin with one role where interview volume and recruiter variation are visible. Automate the repeatable information collection, keep a person accountable for the hiring decision and follow the applicants into the work. If the process fills capable people sooner without losing the people you need, the business case has evidence. If it does not, the organisation has found that out before replacing judgement with a queue.

Filed under · Method · Recruitment · AI voice agents · Workforce Inference Institute · 12 Sept 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.