Ask a room of people to invent a fresh metaphor and you get a scattering of images, some clumsy, some startling, few alike. That scattering is the point. Mengchen Dong and Hiromu Yakura, in a preprint posted to arXiv on July 29, 2026, set out to measure what happens to it when generative AI joins the room, and found that the answer depends entirely on which job the AI is given.

The pair ran a preregistered creative writing experiment with two groups: people writing in English as their first language (L1) and people writing in English as a second language (L2). Each writer worked under one of three conditions. Some wrote metaphors alone. Some received AI-generated ideas to start from, a setup the authors call AI ideation. Others wrote their own ideas first and then had an AI refine them, called AI refinement. The measure that mattered was not how good any single metaphor was, but how different the metaphors were from one another across the whole pool. That collective spread is the raw material innovation draws on, and it can shrink even when every individual contribution looks fine.

The L2 writers, the ones working in a language that was not their first, contributed more collective diversity than the L1 writers did. The most varied pools of all came from ideation in a writer's native language. This runs against a common assumption that fluency and creative range travel together, and the authors treat it as evidence that a writer's linguistic and cultural starting point is itself a source of variety, not a handicap to be smoothed over.

Then the AI conditions split apart. When writers began from AI-generated ideas, collective diversity compressed for everyone, L1 and L2 alike, and the L2 advantage became undetectable in the data. When the AI only refined ideas the writers had already produced, both the overall diversity and the L2 advantage survived. Same models, same people, different point of entry, opposite outcome for the group.

Simulated people were not a substitute

The second challenge the authors tested is the tempting shortcut: if diverse participants are hard to recruit, why not simulate them? They built personas from the real backgrounds of their actual participants and used those personas to generate a synthetic version of the entire writer pool. They did not do this cheaply or once. They used three different model families, prompted models in writers' native languages, and raised the sampling temperature, the setting that makes a model's word choices less predictable and more scattered.

Every simulated pool fell below every human pool on collective diversity. Not on average, and not mostly: every one, below every one. Pushing the models harder did produce more variation, but the authors report that the extra variety came through degenerate text, output that scatters because it is breaking down rather than because it is saying something new. Turning up the randomness dial bought noise, not perspective.

One more result complicates any simple lesson. At the level of the individual writer, AI ideation raised writers' own ratings of their work. So the arrangement that shrank the group's collective range made individual writers feel better about what they produced, setting private benefit against the shared good. There was one exception the authors flag: when L2 writers generated ideas in their native language, both the individual and the collective came out ahead.

Why it matters

Most worry about AI and creativity is about quality, whether machine-assisted writing is any good. This study points somewhere less obvious. A tool can improve each person's output while quietly narrowing what the group as a whole is capable of producing, and the loss shows up only when you look across people rather than at any one of them. No individual writer in the ideation condition had an obvious problem. The pool did.

That has practical bearing on how AI features get built. The difference between an assistant that offers you ideas and one that sharpens ideas you brought yourself is small from a product standpoint and large in this data. The authors' framing is that the design of human-AI collaborative workflows determines whether human diversity survives contact with these tools, and their results give that claim a concrete shape: sequence matters, and who goes first matters most.

The simulation result deserves its own caution. Synthetic participants built from real demographic detail are increasingly used to stand in for human samples in research and product testing. On the specific thing measured here, collective creative range, the stand-ins underperformed the people they were modeled on, consistently, across model families and settings.

This is one experiment, on one creative task, with the models available at the time of writing, and it has not yet been through peer review. The authors are careful to say current AI, not AI in principle. What they document is narrow and worth taking seriously: a group's variety is a resource that can be spent without anyone noticing it going.