The next time a job application lands in an AI screening system, the algorithm may judge you more harshly than a human recruiter. Researchers at Princeton University and the University of Chicago ran a simulated hiring experiment and found that LLMs including ChatGPT, Claude and Gemini developed ethnic stereotypes faster and more aggressively than human participants. The models, however, scored roughly 65% higher on a segregation scale. That finding pushes back against the assumption that AI removes human bias from hiring.

What You Need to Know

When LLMs process résumés, they can quickly form generalizations from limited data and lock in stereotypes. The same reasoning abilities that make models like OpenAI's o3 strong at logic puzzles also make them prone to overgeneralizing about social groups. As chatbots gain memory features, early biases risk becoming permanent unless system objectives are redesigned.

How the Study Measured Bias

In the experiment, each model was told to act as a hiring consultant for a fictional city. Candidates came from four fictitious ethnic groups: Tufa, Aima, Reku and Weki. Over 40 rounds, the AI hired for jobs such as doctor, lawyer, janitor and child-care aide. Unbeknownst to the models, every candidate had an equal chance of success. But the LLMs quickly segregated Aima candidates into lower-status roles after observing a single failure, while human participants showed less bias.

  • LLMs scored 1.83 on segregation scale: Human participants averaged 0.84, while OpenAI's reasoning model o3 reached near the maximum of 2.0.
  • Higher-reasoning models showed stronger bias: OpenAI's o3 and DeepSeek's R1 stereotyped more than simpler models, because they are optimized to generalize from few examples.
  • Memory features risk locking in bias: As chatbots remember user history, they may over-index on past patterns and reinforce early stereotypes.

Why This Matters

If AI hiring tools adopt these biases at scale, job seekers from underrepresented groups could face systematic exclusion. The study shows that simply telling a model to be fair has little effect. Instead, the design of AI objectives matters more. When LLMs are rewarded for exploration or diverse hiring outcomes, bias decreases. That suggests employers and AI developers must rethink how they define success for hiring algorithms. Regulators, too, should pay attention: the European Union's AI Act and similar frameworks may need to address the specific risk of AI forming its own biases from experience, not just from training data.

Ryan Liu, a PhD student at Princeton University and coauthor of the study, says the core issue is that LLMs are eager to create generalizations from limited data. “When LLMs rush to generalize in social settings, that’s when things tend to go wrong,” Liu said. The study was presented at ICML in Seoul.

What This Means for Automated Hiring

Companies that deploy AI for resume screening should test not only for inherited biases from training data but also for biases that emerge during use. The experiment suggests that even a well-intentioned model can develop harmful stereotypes after a few interactions. OpenAI and other developers did not respond to requests for comment. As the industry races toward agentic AI that remembers user details, the risk of locked-in bias grows. The fix, researchers argue, lies in aligning AI goals with fair outcomes rather than merely instructing models to be unbiased.