Seventy thousand people applied for jobs. Some of them picked up the phone and answered questions from a human recruiter. The rest answered the same kind of questions from a machine. Neither group knew they were part of one of the largest hiring experiments ever run inside working firms.
That setup is the core of a new preprint by Brian Jabarian and Luca Henkel, posted to arXiv on July 30, 2026. The economists ran what they call a natural field experiment, meaning the study happened inside a real hiring pipeline rather than a lab, with real jobs and real consequences for the people involved. Applicants were randomly assigned to be interviewed either by human recruiters or by AI voice agents, software that conducts a spoken conversation over the phone.
One detail matters enormously for reading the results, and the authors are careful about it. The AI never decided anything. In both arms of the experiment, human recruiters evaluated the interviews and made the hiring calls. What varied was only who asked the questions and collected the answers.
Applicants interviewed by the AI agents were 12 percent more likely to receive a job offer. On its own, that number could mean the machine was simply easier to impress, or that recruiters reading AI transcripts were softer graders. The authors followed the applicants further to check. The gains carried through to higher job starts, meaning more people actually showed up and began working, and to higher retention, meaning more of them stayed. And the workers hired through AI interviews were no less productive than those hired through human ones.
That combination is what makes the finding hard to dismiss as a scoring artifact. If the AI interviews were just flattering weaker candidates, you would expect the extra hires to wash out somewhere downstream: in offers not accepted, in early departures, in weaker performance. Jabarian and Henkel report that they did not.
What the transcripts showed
The more interesting part of the paper is the attempt to explain why. The authors analyzed the interview transcripts themselves, comparing what the AI agents did with what the human recruiters did.
Their summary is a phrase worth sitting with: controlled variance. The AI interviews were more structured and more consistent from one candidate to the next, which is roughly what you would guess about software. But they were not rigid scripts. The transcripts showed the agents still responding to individual applicants, following up on what a particular person actually said. Structure without stiffness, in other words.
And that consistency was associated with more hiring-relevant information ending up in the transcript. This is the mechanism the authors propose for the whole result. A human recruiter conducting the fifteenth interview of the day is a variable instrument: tired, distracted, warmed up by a good previous conversation, quietly bored. Two equally qualified applicants can walk away from that recruiter having been asked meaningfully different questions. The recruiter reading the transcript later then has less to work with for one of them than the other.
The authors frame their finding in terms of information collection rather than judgment. Automation here is not replacing the decision. It is improving the raw material the decision is made from, by making sure every applicant gets asked the things that actually predict whether they can do the job.
A word on what the paper does not claim. It reports an association between the AI agents' consistency and the amount of hiring-relevant information gathered, not a clean causal chain from one to the other. And this is a single experiment, however large, in a specific hiring context. The abstract does not specify which industry, which countries, or what kinds of roles were involved, so how far these numbers travel to other labor markets remains an open question.
Why it matters
Most public argument about AI in hiring runs on a fear of automated judgment: an algorithm scoring résumés, filtering people out, encoding old biases at scale and at speed. This study describes something different, and the distinction is worth holding onto. The humans kept the judgment. The machine did the interviewing.
If that division of labor holds up in other settings, it suggests a general principle that reaches past hiring. Plenty of organizational decisions rest on information gathered by people who are inconsistent through no fault of their own: intake interviews, medical histories, loan applications, benefits assessments. Anywhere a human collects information under time pressure, the variation in how the collecting is done becomes noise in whatever decision follows. Standardizing that step, while leaving the judgment to a person, is a narrower use of automation than most current proposals.
There are real questions the paper leaves open, including how applicants themselves experienced being interviewed by software, and whether the offer gains were shared evenly across different kinds of candidates. Those matter for anyone deciding whether to actually deploy this. But the finding stands on its own terms: on this evidence, the machine was not a better judge of people. It was a more reliable listener.