Inside a multinational industrial company, 128 employees recently became research subjects. Each worked through three pieces of representative knowledge work: one task built around acquiring knowledge, one around packaging it, and one around creating it. Some were randomly assigned to do the work with generative AI at their side. Others worked without it. The researchers running the experiment wanted an answer to a question that surveys and anecdotes keep circling: when a chatbot enters real office work, does the work get better, or merely faster?

The answer, posted to arXiv this week as a preprint by Sven Bottesch, Chiara Schwenke, Jakob Zimmermann, Maximilian Förster and Mathias Klier, is that it depends on what the work is. Speed improved everywhere: participants using generative AI completed all three types of task more efficiently than those working unaided. Quality was a different story. When the job was to package existing knowledge or to create something new, the AI users produced better work. When the job was to acquire knowledge, quality declined.

The study is what the authors call a randomized lab-in-the-field experiment. It has the controlled structure of a laboratory study, with participants assigned by chance to the AI or no-AI condition, but the subjects are practicing knowledge workers at a real company rather than students or crowdworkers recruited online. The team grounds its design in task-technology fit theory, a long-standing idea in information systems research holding that a tool improves performance only to the extent that its capabilities match the demands of the particular job. Generative AI, on this view, is not one technology with one effect. It is a tool whose value shifts with the task it is pointed at.

The three task types were chosen to represent the span of knowledge work. Acquisition is the part of the job where you find out and absorb something. Packaging is where you shape what is already known into a form others can use. Creation is where you produce something new. The abstract does not spell out why acquisition fared worst, and that pattern will be worth scrutinizing when the full results go through peer review. But the direction of the finding is striking on its own: the one task category centered on getting hold of correct information is the one where the chatbot dragged quality down.

The study's most textured finding concerns not averages but spread. For packaging and creation, generative AI tended to reduce the variance in quality across workers, and the authors report that the benefits went primarily to lower-performing knowledge workers. The tool acted as a leveler, pulling the bottom of the distribution up toward the middle. For knowledge acquisition, it did the opposite: variance in quality grew. On the task where AI hurt on average, it also made outcomes less predictable.

The usual cautions apply. This is a preprint, not yet peer reviewed. The participants all came from a single multinational industrial organization, and the experiment covered three specific tasks meant to stand in for broad categories of knowledge work. The publicly available abstract reports the directions of the effects rather than their sizes, so how much faster, and how much better or worse, remains to be read in the full paper.

Why it matters

Debates about AI and productivity tend to collapse into a single number: some percentage faster, some percentage more output. This experiment suggests a single number is the wrong shape for the answer. In the authors' data, the same tool used by the same kind of workers helped or hurt depending on what the task asked of them. That means an organization rolling out generative AI uniformly across roles could be improving some work while quietly degrading other work, all while everything gets faster and therefore looks more productive.

The leveling result carries its own weight. If AI mainly lifts lower performers on packaging and creation tasks, it compresses skill differences that hiring, pay and promotion currently rest on. And the acquisition result reads as a warning about one of the most common office uses of chatbots today, asking them questions in order to learn something. That is precisely the mode in which this study found quality falling and becoming more erratic.

The findings are preliminary and come from one company. But they push in a useful direction: away from asking whether AI makes knowledge workers more productive, and toward asking which parts of their work it actually fits.