AI hiring tools are more likely to stereotype job applicants than humans are, new research finds
A Princeton and University of Chicago study put ChatGPT, Claude, and Gemini through a simulated hiring game. The models quickly sorted fictional ethnic groups into job categories and did it far more aggressively than human test subjects did.

Key points
- Princeton and University of Chicago researchers published findings in July 2025 showing that large language models (LLMs) stereotype job candidates more severely than humans do in controlled hiring simulations.
- On a segregation scale where 2.0 means complete group-to-job confinement, human participants scored 0.84. OpenAI's reasoning model o3 scored 1.83.
- Telling the models to be fair had little effect, but offering a bonus reward for diverse hiring significantly reduced biased behaviour.
- Models became less biased when given relevant personal details about candidates, such as age and education, but reverted to group stereotyping when given irrelevant details like hair colour.
- As reported by MIT Technology Review, companies now deploying AI to screen résumés face a serious and largely unresolved fairness problem.
Before a single human recruiter reads your résumé, an AI system may have already decided you are not the right fit. That is no longer a distant worry. It is a present reality at many large employers.
New academic research shows just how badly that can go wrong.
Researchers at Princeton University and the University of Chicago ran three major large language models (LLMs, the AI technology behind chatbots like ChatGPT, Claude, and Gemini) through a simulated hiring game. Each model played the role of a hiring consultant for a fictional city. Over 40 rounds, the model chose candidates from four made-up ethnic groups to fill 20 jobs ranging from doctors and lawyers to child-care aides and janitors.
Here is the catch the models did not know: every candidate had an equal chance of succeeding, regardless of group.
The models did not behave that way. When an early hire from one group failed, the model quickly stopped considering that group for similar roles and started funnelling them into lower-status jobs instead. One bad result became a rule applied to an entire population.
Humans do this too. But the models did it roughly 65 percent more aggressively than human participants in the original psychology study the experiment was adapted from.
Should job seekers be worried?
Yes, and here is the honest picture. The study was a simulation, not a live recruitment system. Real-world AI hiring tools do not get instant feedback on whether a new hire worked out, which slows bias formation. But the research team warns that even delayed feedback, the kind companies receive months after hiring, could still cause AI systems to over-generalise from small samples.
Ryan Liu, a PhD student at Princeton and a co-author of the study, says LLMs are explicitly trained to build generalisations from limited data. That is what makes them useful for coding problems and logic puzzles. In social settings, that same instinct turns into stereotyping.
Newer, more capable reasoning models made things worse, not better. OpenAI's o3 and DeepSeek's R1 (a competing AI reasoning model built in China) both showed stronger biases than older, simpler models.
Telling a model to "be fair" changed almost nothing. But rewarding it for diverse hiring outcomes worked. The study also found that giving the AI genuinely relevant personal information about a candidate, such as their work history or education level, reduced stereotyping. Feeding it irrelevant details, like physical appearance, pushed it straight back to sorting people by group.
The practical takeaway for anyone applying to jobs right now: ask employers directly whether AI screens your application and, if so, what bias audits they run on that system. It is a reasonable question, and a company serious about fair hiring should have a straight answer ready.



