Amplifying Creativity
Series date follows the editorial schedule. First published ; updated .
A collection can become less varied even while each item gets better reviews. That is a useful warning for anyone using AI to generate options: judging the best proposal and judging the range of proposals are different jobs.
The problem is easy to miss in a writing tool. A suggestion helps one person get past a blank page. The finished piece reads better. If many people receive similar starting points, though, the collection may contain fewer distinct directions. The individual and the editor looking at the whole collection can have different reasons to be pleased.
Better stories, closer together
Anil Doshi and Oliver Hauser tested this in a 2024 Science Advances experiment. Participants wrote an eight-sentence story for a teenage and young-adult audience. Writers were randomly assigned to work without AI ideas, to have access to one GPT-4 starting idea, or to have access to up to five.
The study collected 293 stories. A separate group of 600 evaluators assessed stories without initially knowing which writers had been offered AI. Writers could choose whether to request the ideas. They could not customise the model prompt or have an iterative conversation with it.
Access to AI ideas improved evaluators' ratings of novelty and usefulness on average. In the up-to-five-ideas condition, the reported increases relative to the human-only baseline were 8.1% for novelty and 9.0% for usefulness. “Usefulness” here was a story-specific rating about suitability for the audience and potential for further publication, not proof that a publisher would buy the work.
The researchers also examined how similar stories were to other stories in the same condition, using text embeddings and cosine similarity. AI-assisted stories were closer to the average of other stories in their condition. That collection-level finding can coexist with an evaluator judging an individual story more novel.
This was not a test of output volume: the task called for one short story per writer. Nor was it a comparison showing that people with AI beat a standalone AI on the same writing task. The study supports a narrower claim about access to starting ideas in this constrained setting.
Benefits were larger for writers with lower scores on a preliminary divergent-association task. That score is one creativity-related measure, not a complete assessment of a person's talent. It cannot establish that adaptable generalists should replace specialists in a hiring plan.
The first suggestion can set the direction
The stories were also more similar to the AI ideas the writers received. That is consistent with anchoring on a starting point. It does not prove that every shared model produces a creative monoculture, or that a model cannot propose something original.
The experiment's constraints matter here. A professional writer might reject a suggestion, change the prompt, draw on outside material, or revise repeatedly. Those behaviours were not the treatment tested. The paper is a reason to watch for convergence, not a forecast of all creative work.
There is a parallel design concern outside fiction. If a team asks for twenty product concepts and receives twenty variants of the same assumption, counting concepts overstates the range it explored. But that product-team example is an inference to investigate, not a measured outcome from the story experiment.
Keep the alternatives different where it matters
Imagine a hypothetical team designing an exhibition about a river. One approach follows its ecology, another follows a family's memories, and another asks visitors to decide how scarce water should be allocated. These are different organising ideas, not just three colour palettes.
Try the collection test on an exhibition brief
Write one initial concept before requesting suggestions. Then ask an assistant for alternatives and list the main assumption behind each. Does the visitor observe, investigate, or make a choice? Whose perspective is missing?
Have another person assess the concepts without seeing which were AI-assisted. Rate the strength of each separately from how much new ground it adds to the set. A polished version of an existing concept may score well on one and poorly on the other.
This is a proposed design exercise. The cited study did not test it, and it does not guarantee more original work. Keep it only if the additional range is useful enough to justify the effort.
Changing models or prompts may produce different wording without changing the underlying idea. Do not assume that purchasing several tools buys independence between their suggestions. Look at the outputs and the assumptions they share.
Nor should a team preserve every unusual idea indefinitely. Variety has a cost. An exhibition still needs to make sense to visitors and fit its space and budget. The point is to avoid eliminating a direction before anyone has understood what it offers, then to make a deliberate choice.
When the task requires consistency, convergence may even be useful. A set of instructions should not be made less clear merely to maximise novelty. Decide what kind of variation serves the audience before treating variation as a target.
Credit and evidence belong in the process
AI-generated material can still contain unattributed borrowing, invented facts or a confident description of something that never happened. Check factual claims against sources and follow the relevant publication and rights requirements. A high creativity rating is not clearance to publish.
Keep enough provenance to explain the process honestly: which ideas were supplied, what the author changed, and which sources support the factual material. That also makes it easier to notice when supposedly independent proposals descend from the same starting point.
I would ask a creative team to show its strongest option and one materially different option it seriously considered. The second may lose. Seeing why it lost is more informative than being shown ten polished versions of the winner.
Source and scope
Doshi and Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances, 2024, DOI 10.1126/sciadv.adn5290. See experimental design, creativity ratings, embedding comparisons and limitations. This revision removes unsupported volume claims, universal claims about the origin of breakthroughs, and a private architecture anecdote. The exhibition exercise is hypothetical.