TL;DR: Generative AI raises the quality of your individual output and lowers the diversity of everyone’s output at the same time. In a controlled Science Advances experiment, writers given AI story ideas produced work rated as more creative, with the biggest gains going to the least creative writers, yet those AI-assisted stories were measurably more similar to one another than stories written without help. A separate Creativity and Cognition study found that people brainstorming with ChatGPT produced ideas that were significantly more alike at the group level (effect size d=0.47), even though each individual stayed just as varied as before. And an ICLR study found that writing with a feedback-tuned model reduced the lexical and content diversity across authors. Put together, these results describe a social dilemma: each person is better off using the tool, but collectively a narrower band of ideas gets produced. If you and your competitors all prompt the same model with the same question, you converge on the same answer. This guide explains the mechanism, gives you the Homogenization Risk Matrix to locate where your own work is collapsing toward the average, and lays out the practices of someone who leads like a CEO who protects a differentiated position and learns like a student who keeps generating from first principles.
The uncomfortable finding of the last two years of AI research is not that these tools make people worse. In most measured tasks they make individuals better. The finding is that they make people more alike. When a tool is trained to produce the statistically most likely helpful response, and millions of people query it with similar prompts, the tool becomes a giant averaging machine for human thought. Your output improves. So does everyone else’s. And the gap between your ideas and the next person’s quietly shrinks.
For anyone whose value depends on being different, and that includes almost every professional, founder, and creator, this is the strategic risk of the AI era that almost no one is pricing in.
What the homogenization effect actually is
Homogenization is not the same as decline. A field can get better on average while its members get harder to tell apart. That is exactly what the evidence shows.
The distinction that matters is between the individual level and the collective level. At the individual level, AI is an enhancer: your single essay, pitch, or design is likely to be more polished with model help than without. At the collective level, AI is a compressor: the range of essays, pitches, and designs produced by a population of AI users is narrower than the range that same population would have produced alone. Both things are true simultaneously, and missing either half leads to the wrong decision. People who only see the individual gain over-adopt and lose their edge. People who only see the collective risk refuse the tool and lose the productivity.
Think of it as a distribution. AI pulls the low end up, which is the real and valuable gain. But it also pulls the whole distribution toward a central peak, so the tails, where genuine originality lives, get thinner. A CEO reads a distribution before celebrating an average. A student asks what the average is hiding. This is the flip side of a trend we have argued is otherwise good news, the judgment economy, where the human faculties AI cannot average become the scarce and valuable ones.
The evidence, measured directly
The reason this is worth treating as fact rather than fear is that researchers have now measured it under controlled conditions, with the diversity of output as the explicit variable rather than an afterthought. The pattern replicates across different tasks, models, and diversity metrics.
| Study (year, venue) | Setup | What was measured | Core finding |
|---|---|---|---|
| Doshi and Hauser, 2024, Science Advances | Online experiment; some writers received short-story ideas from a large language model, others did not | Evaluated creativity of each story, plus similarity between stories | AI ideas raised individual creativity ratings, with the largest gains for the least creative writers, while AI-assisted stories were more similar to each other than unaided stories |
| Anderson, Shah and Kreminski, 2024, ACM Creativity and Cognition | 33 participants, within-subjects, divergent-thinking tasks using ChatGPT versus Oblique Strategies | Cosine similarity of sentence embeddings at individual and group level | Group-level output was significantly more homogeneous in the ChatGPT condition (d=0.47, P=0.038); individual-level diversity showed no significant difference (P=0.352) |
| Padmakumar and He, 2024, ICLR | Users wrote argumentative essays with a base model, a feedback-tuned model, or no model | Lexical and content diversity across different authors | Writing with the feedback-tuned model produced a statistically significant reduction in diversity and higher similarity between authors; the base model did not |
Read the middle finding again, because it is the most important one for daily work. The homogenization did not show up when researchers looked at individuals in isolation. Each person, using ChatGPT, was about as varied as they were using a non-AI creativity tool. The convergence appeared only when you compared different people to each other. This is why the effect is nearly invisible from the inside. Working alone with the model, you feel fully creative. You cannot see that a thousand other people, prompted similarly, are being nudged toward the same neighborhood of ideas you just landed in. The loss is real but it is only observable from above, which is precisely the CEO’s vantage point and not the individual contributor’s.
Why the mechanism is structural, not a bug
It is tempting to assume better models will fix this. The mechanism suggests otherwise, because homogenization comes from how these systems work, not from how good they are.
A large language model generates by estimating the most probable continuation given its training data and your prompt. Reinforcement learning from human feedback, the step that makes models feel helpful and safe, further concentrates outputs toward responses that a broad set of human raters approved. The Padmakumar and He result points straight at this: the base model did not reduce diversity, but the feedback-tuned model did. The very training that makes a model pleasant to use is the training that narrows what it says. A more capable model tuned the same way may produce a higher-quality center of gravity, but it is still a center of gravity. The pull toward the middle is the product working as designed.
Add the network effect on top. The more people rely on the same few frontier models, the more a shared substrate underlies human output across an entire industry. Everyone drinking from the same well tastes the same water. This is not a claim that ideas will become literally identical. It is a claim, now backed by measurement, that the variance between people’s ideas shrinks, and variance is exactly what competitive advantage, scientific breakthroughs, and memorable creative work are made of.
The Homogenization Risk Matrix
Not every use of AI carries the same risk. Using a model to format a table or fix grammar does not threaten your originality; using it to generate your core strategic idea does. The mistake is treating all AI use as equally safe or equally dangerous. The following matrix, a CEOtudent editorial framework, sorts common work contexts by how exposed they are and what the corrective move is. It is a judgment tool, not a scoreboard.
| Work context | Homogenization risk | Why | Counter-move |
|---|---|---|---|
| Grammar, formatting, translation, summarizing your own draft | Low | The idea is already yours; AI only touches the surface | Use freely; no defense needed |
| Research and fact-gathering | Low to medium | Sources converge, but you still synthesize | Cross-check across independent sources; add data the model did not surface |
| First-draft ideation for a strategy, positioning, or creative concept | High | This is where the averaging bites hardest and is least visible | Generate your own ideas first, then use AI to critique, not to originate |
| Naming, taglines, campaign concepts, product angles | High | Everyone prompts the same model for the same category; outputs cluster | Treat AI output as the baseline everyone will reach; deliberately depart from it |
| Competitive or market analysis shared across an industry | High | Rivals using the same tool arrive at the same read of the market | Anchor on proprietary data and lived context the model cannot access |
| Personal reflection, judgment, and taste development | Critical | Outsourcing this erodes the faculty that makes you differentiated over years | Keep a no-AI zone for forming your own views before consulting any model |
The organizing principle is simple. The closer AI gets to originating your differentiated idea, the higher the risk. The further it stays toward execution and polish, the safer. Most people have this backwards, guarding their spelling while handing over their thinking.
What to do about it: lead like a CEO, learn like a student
The answer is not to abandon the tool. Refusing AI to preserve originality is like refusing electricity to preserve craftsmanship: you keep your difference and lose everywhere else. The answer is a discipline of use.
Generate before you retrieve. The single most protective habit is sequence. Form your own answer, however rough, before you open the model. Doshi and Hauser found the diversity loss came from AI supplying the ideas. If your idea exists first, the model becomes an editor of something original rather than the author of something average. This is the student’s move: do the hard cognitive work yourself, then check it, because the checking is worthless if there was no independent attempt to check.
Use AI as an adversary, not an oracle. Ask the model to argue against your idea, list its weaknesses, or name the ten most obvious versions of it so you can avoid them. Turning the tool into a critic exploits its strength, breadth of pattern knowledge, without letting it dictate your position. The obvious-versions prompt is especially useful: if the model can list your idea among the predictable ten, so can your competitors’ models. Reasoning up from fundamentals rather than accepting the model’s defaults is the habit we cover in first principles versus best practices, and it is the most reliable route out of the average.
Protect a proprietary input. Homogenization is strongest where everyone feeds the model the same public information. Your defense is private information: your own data, your customers’ exact words, your specific context, a constraint only you face. The model averages the commons; it cannot average what only you hold. A CEO builds moats out of proprietary inputs, and in the AI era proprietary input is the moat.
Keep a no-AI zone for judgment. Taste and judgment are faculties that atrophy when outsourced. Reserve some category of decision, your strategic direction, your point of view on your field, your sense of what is good, that you form without a model, precisely so the muscle stays strong. You can consult AI afterward. Forming the view yourself first is what keeps it yours. This is the same case we make for deliberately building the taste stack: the aesthetic judgment that lets you tell a good departure from a bad one has to be trained, not delegated.
Assume the average and price it in. Treat the model’s first answer as the answer your entire market will also receive. That reframing is liberating: whatever the AI hands you is, by definition, the new baseline, not the finish line. The work that gets noticed, cited, and remembered starts where the model’s output ends.
What this means for teams and organizations
At the individual level these habits protect your edge. At the organizational level the stakes compound, because a company is a population of AI users and the collective effect is exactly what the research measured.
If every analyst prompts the same model, the organization’s collective read of the market narrows toward what every competitor also sees. If every marketer generates concepts the same way, the brand’s voice drifts toward the category mean. The organizational counter-move mirrors the individual one: mandate independent generation before shared tools on the decisions that matter, invest in proprietary data that models cannot access, and reward the departures from the AI baseline rather than the fastest arrival at it. The firms that will stand out in the AI era are not the ones that adopt fastest. They are the ones that adopt while deliberately protecting the variance the tool erodes.
FAQ
Is the homogenization effect proven or just a theory?
It is measured under controlled conditions in multiple peer-reviewed studies, using explicit diversity metrics such as similarity of text embeddings and lexical variety. It is one of the better-established findings about generative AI’s effect on human output, precisely because researchers set out to measure diversity directly rather than infer it.
Does using AI make me less creative as an individual?
Not according to the evidence. Individually, AI assistance tends to raise creativity ratings, especially for people who start with lower creativity, and individual-level diversity was not significantly reduced in the brainstorming study. The loss shows up between people, at the collective level, which is why it is invisible when you work alone.
Will better AI models solve this?
Unlikely on their own. The reduction in diversity was linked to the feedback-tuning that makes models helpful, not to a lack of capability. A more powerful model tuned the same way still pulls outputs toward a probable center. The defense is how you use the tool, not which tool you use.
What is the single most effective countermeasure?
Sequence. Generate your own idea before consulting the model, then use AI to critique and refine rather than to originate. If your original thought exists first, the model sharpens your difference instead of replacing it with the average.
Does this apply to code and analysis, or only creative writing?
The controlled studies focused on writing and ideation, so the direct evidence is strongest there. The mechanism, probable-completion generation plus shared models plus similar prompts, applies anywhere many people query the same system with similar inputs, which includes code patterns, market analysis, and strategy.
Kaynakça
- Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances.
- Anderson, B. R., Shah, J. H. and Kreminski, M. (2024). Homogenization Effects of Large Language Models on Human Creative Ideation. Proceedings of the 16th ACM Conference on Creativity and Cognition.
- Padmakumar, V. and He, H. (2024). Does Writing with Language Models Reduce Content Diversity? International Conference on Learning Representations (ICLR).
- OECD (2024). OECD Employment Outlook: analysis of artificial intelligence and the labour market.
- Stanford Institute for Human-Centered AI (2025). Artificial Intelligence Index Report, chapters on generative AI capabilities and adoption.
- Brynjolfsson, E., Li, D. and Raymond, L. (2023). Generative AI at Work. National Bureau of Economic Research working paper.
This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.
This post is also available in:













