Large language models have a sameness problem that goes beyond individual quirks. Ask ChatGPT, Claude or Gemini to pick a random number between 1 and 10 and each will almost certainly return 7. Ask them to name a car and the answer is likely Toyota or Honda. The predictability is not a bug but a feature of how these models are built and trained.
The Root of AI Groupthink
In November 2025, researchers from multiple institutions published a paper at NeurIPS that won best paper award for documenting this exact phenomenon. They asked 25 different LLMs 50 times each to write a metaphor about time. Most of the 1,250 responses were some variation of Time is a river or Time is a weaver. The team speculated that homogeneity arises because all major models are trained on similar web data using similar reinforcement learning techniques. They all learn the same patterns, the same tropes and the same safe answers.
Examples of this repetition are easy to reproduce. Consider these prompts:
One Startup's Unconventional Fix
Springboards, an Australian startup cofounded by Pip Bingemann and Kieran Browne, built Flint specifically to break out of this rut. Unlike conventional approaches that raise the temperature parameter across all tokens, Flint injects randomness at strategic decision points where models typically settle on the most probable token. This technique produces outputs that are still coherent but far more varied. “Most language models are fighting hallucinations,” Bingemann said. “We welcome them.” When asked for a random number, Flint returned 3.7916. That is the kind of response that exposes the bias in standard LLM training: they default to round, familiar digits.
Bingemann gave another example. He prompted ChatGPT, Claude and Flint to create a tagline for New Balance running shoes. Both ChatGPT and Claude returned “Run your way.” Flint returned “Built to last, run to win.” It is not a poetic line, but at least it is different from the others. The homogeneity extends beyond simple prompts. Browne noted that the design of chat interfaces makes users feel they are having a private conversation. That feeling is misleading. Most people are getting the same advice, the same metaphors and the same recommendations as everyone else.
Why This Matters
The practical consequences of AI homogeneity are significant for businesses and creatives who rely on LLMs for fresh ideas. If every travel agent, marketer and writer uses the same AI tools to generate content, the resulting recommendations and narratives will all sound alike. This erodes competitive advantage and reduces the value of AI as a brainstorming partner. Startups like Springboards offer one path forward, but their approach also carries risks. Injecting randomness can produce nonsensical or harmful outputs if not carefully controlled. The industry will need to decide whether the goal is always-safe answers or genuinely useful variety. For now, the safest output remains a predictable one. That may soon change as more companies recognize the cost of groupthink in their AI systems.



