Automate the interview script, not the hiring decision
A voice agent may standardise early interview evidence, but candidate participation and later work outcomes determine its value. Test information collection separately from scoring, with a person accountable for selection.
Keep development and test groups separate before comparing performance. No measured results are shown.
High-volume recruitment needs comparable evidence from early interviews. Repeated questions can vary by interviewer, while candidates wait for available appointments. A voice agent offers a way to standardise collection, but also changes the experience that determines whether applicants continue.
The evaluation must therefore cover both the quality of the record and the people lost before that record is complete. A reduction in recruiter time can be offset by technical failures, inaccessible interaction or suitable applicants choosing not to participate.
Test automated information collection separately from the hiring decision. A voice agent can follow an approved question path while a trained recruiter reviews the resulting evidence. Assess offers, starts, retention and work quality alongside continuation and review effort before widening the workflow.
Consistency can change the hiring economics
Jabarian and Henkel ran a natural field experiment with 70,884 applications for entry-level customer-service roles. Applicants who passed an initial screen were assigned to a human recruiter, an AI voice agent or a choice between the two. Human recruiters made every final hiring decision. Jabarian and Henkel, 2026
The AI-interview group received offers in 9.73 per cent of cases, compared with 8.70 per cent for human interviews. Applicants assigned to the AI agent were 18 per cent more likely to start the job and 18 per cent more likely to remain employed for at least one month. The study found no meaningful difference in handling time, customer satisfaction or employer quality scores for the workers who were hired. Jabarian and Henkel, 2026
The authors connect the result to a specific mechanism. The agent followed the same topic order more closely, covered a more consistent set of topics and used more standardised prompts and follow-ups. That reduced interviewer-driven variation while preserving responses to the individual applicant.
The proposed mechanism is more comparable information reaching the recruiter. Treat that as a local hypothesis to test, rather than assuming that standardisation alone improves hiring. The record still needs to contain evidence relevant to the role and be usable by the person making the decision.
- 01 Invite Tell candidates what the interview is for, who will review it and what happens next.
- 02 Collect Use one approved question path with room for relevant follow-ups.
- 03 Evaluate Let a trained recruiter review the same evidence and apply the hiring standard.
- 04 Decide Compare offers, starts, retention and work quality with the existing process.
The format can spend the value before evaluation begins
Avery and colleagues randomised more than 3,000 applicants into asynchronous audio interviews, asynchronous video interviews, live online interviews or no screening. They measured whether applicants continued and how the resulting interviews were assessed. Avery and colleagues, 2026
Continuation fell by about 53 per cent, or 45 percentage points, after an asynchronous interview. The live online format reduced continuation by about 20 per cent, or 17 percentage points, against the no-screening control. Women were 5.1 percentage points less likely than men to complete the asynchronous interview. The paper attributes the deterrence to beliefs about fairness and the number of people competing for each role. Avery and colleagues, 2026
Interview format belongs in the business case. Measure who continues and who leaves before assessment, including consequential differences across groups. Recruiter time saved through applicant withdrawal is not equivalent to useful capacity released through better collection.
The same study found that an AI assessment tool scored women and underrepresented racial minorities higher than human evaluators, and that its scores were more predictive of later employment outcomes in that sample. Those results concern the evaluation of interview answers, not a universal claim that an automated assessor should make the decision. Avery and colleagues, 2026
Keep the two interventions separate in the pilot. Changing who asks the questions and changing who scores the answers create different mechanisms, risks and evidence requirements. Combining them can make a positive result impossible to explain and a negative result impossible to repair.
Measure the whole funnel, not the interview queue
The buyer for this decision is usually a chief people officer or a head of high-volume recruitment. Their outcome is not interviews completed. It is a filled role that reaches the required level of performance and remains useful long enough to justify the hiring effort.
Start with a baseline for the existing process. Record invitations, starts, completion by stage, interview duration, recruiter time, offer rate, acceptance, time to start and early retention. Segment the record by role, location, experience and any group for which continuation or selection consequences matter.
Then define the smallest operational change the pilot is meant to support. It could be more completed interviews per recruiter, a shorter interval from application to offer, fewer interviews required per start or better retention in the first month. Name the measure before the agent is introduced.
Do not hide the cost of obtaining the signal. Record failed calls, technical handoffs, candidate requests for a person, reviewer time and cases where the transcript is too incomplete to use. The Jabarian and Henkel study reports that 5 per cent of applicants ended their interview because they did not want to speak to an AI, while technical difficulties affected 7 per cent of cases. Jabarian and Henkel, 2026
Read hiring outcomes after the interview. Offer rate can rise because the threshold changed or because the evidence improved. Start rate can rise while retention falls. A useful comparison follows applicants into the work, where quality, attendance, customer results or supervisor assessment can be observed.
Put the human boundary where the evidence becomes consequential
The voice agent should collect and organise information. It should not decide that a person is acceptable because a transcript resembles a successful case. The recruiter needs the question path, the candidate’s answers, the scoring rubric and the points where the system failed or deviated.
That record also makes supplier evaluation concrete. Ask whether the agent can preserve the exact prompt and response sequence, expose unanswered questions, replay a technical failure and hand the case to a person without losing context. Ask how the system is tested when accents, interruptions, silence or a candidate asking for a human change the expected path.
The approval boundary belongs after evidence review. A recruiter can override a recommendation, record why and identify the evidence that would have changed the decision. If the system makes the final choice, the pilot is testing an automated selection policy, not an automated interview, and it needs a different governance and measurement design.
What the Inference Institute can help decide
The Institute can review the screening workflow, identify sources of interviewer variation and specify the information the early interview should collect. Architecture and risk advice should separate candidate experience, information quality and selection outcomes, with implementation retained by the client or chosen partner.
The result is a pilot specification a people leader can fund and review. It names the role, question path, human decision point, control design, cohort measures, post-hire observation and stopping rule. It also makes clear whether the intended gain is recruiter capacity, faster starts, better matching or retention.
What this does not tell you
The AI voice study concerns one recruitment-process outsourcing firm, entry-level customer-service roles in the Philippines and a period when the partner’s workflow was already structured. Its retention measure is one month in a high-turnover market. It does not forecast another role, country or hiring standard.
The asynchronous-interview study concerns three technology jobs and a platform used by the participating employer. Its AI assessment findings come from one commercial tool and a selected set of interview answers. Neither paper establishes that a particular vendor, script or scoring threshold will improve your hiring.
Begin with a defined role and an approved information-collection boundary. Compare candidate continuation, recruiter effort, selection and later work outcomes with the existing process. A decision to expand should account for the applicants lost and the review burden, as well as any improvement in starts or retention.