The next time a job application lands in an AI screening system, the algorithm may judge you more harshly than a human recruiter. Researchers at Princeton University and the University of Chicago ran a simulated hiring experiment and found that LLMs including ChatGPT, Claude and Gemini developed ethnic stereotypes faster and more aggressively than human participants. The models, however, scored roughly 65% higher on a segregation scale. That finding pushes back against the assumption that AI removes human bias from hiring.
How the Study Measured Bias
In the experiment, each model was told to act as a hiring consultant for a fictional city. Candidates came from four fictitious ethnic groups: Tufa, Aima, Reku and Weki. Over 40 rounds, the AI hired for jobs such as doctor, lawyer, janitor and child-care aide. Unbeknownst to the models, every candidate had an equal chance of success. But the LLMs quickly segregated Aima candidates into lower-status roles after observing a single failure, while human participants showed less bias.
Why This Matters
If AI hiring tools adopt these biases at scale, job seekers from underrepresented groups could face systematic exclusion. The study shows that simply telling a model to be fair has little effect. Instead, the design of AI objectives matters more. When LLMs are rewarded for exploration or diverse hiring outcomes, bias decreases. That suggests employers and AI developers must rethink how they define success for hiring algorithms. Regulators, too, should pay attention: the European Union's AI Act and similar frameworks may need to address the specific risk of AI forming its own biases from experience, not just from training data.
Ryan Liu, a PhD student at Princeton University and coauthor of the study, says the core issue is that LLMs are eager to create generalizations from limited data. “When LLMs rush to generalize in social settings, that’s when things tend to go wrong,” Liu said. The study was presented at ICML in Seoul.
What This Means for Automated Hiring
Companies that deploy AI for resume screening should test not only for inherited biases from training data but also for biases that emerge during use. The experiment suggests that even a well-intentioned model can develop harmful stereotypes after a few interactions. OpenAI and other developers did not respond to requests for comment. As the industry races toward agentic AI that remembers user details, the risk of locked-in bias grows. The fix, researchers argue, lies in aligning AI goals with fair outcomes rather than merely instructing models to be unbiased.



