Simply telling an AI model to act fairly did little to reduce biased hiring patterns in a new simulation. The finding raises questions for employers that use AI tools to screen résumés or support interviews.
The models became less biased when they received additional incentives to make diverse hiring choices. Providing relevant candidate information, including age and education, also reduced ethnic segregation more effectively than adding irrelevant personal details.
Relevant information made a difference
Details connected to a candidate’s suitability for work helped models make less segregated choices. By contrast, information such as hair color or tattoo shape, which was unrelated to job ability, did not substantially reduce the pattern.
The results suggest that general fairness instructions alone may not be sufficient safeguards for AI hiring systems. System design, the information supplied to the model, and the way performance feedback is used can all affect later decisions.
The study does not mean every real-world AI recruitment system will necessarily produce the same discriminatory outcome. In practice, résumé-screening tools do not immediately know whether a hired worker will succeed, as the simulation allowed.
Even so, researchers warned that performance feedback received after a hiring decision may influence how a model evaluates future applicants. That creates a potential route for AI Hiring Bias to become embedded through limited early outcomes.
Models exceeded the human benchmark
Researchers used a segregation scale in which 2 represents complete separation of groups into their own occupational fields. Participants in an earlier human study recorded an average score of 0.84.
| Participant or model | Segregation result | Context |
|---|---|---|
| Human participants | 0.84 | Average in an earlier study |
| AI models | About 65% higher | Compared with human participants |
| OpenAI o3 | 1.83 | Near the maximum score of 2 |
Across the experiment, AI models showed segregation levels about 65 percent higher than the human benchmark. OpenAI o3 reached 1.83, close to the scale’s maximum value.
Reasoning-oriented models, including OpenAI o3 and DeepSeek R1, displayed stronger bias tendencies in the experiment. Ryan Liu, a Princeton University doctoral student and one of the study’s authors, told MIT Technology Review that LLMs can generalize very quickly from limited data.
A hiring simulation with fictional groups
Scientists from Princeton University and the University of Chicago conducted the research using a hiring simulation adapted from a 2024 psychology study on stereotype formation. The test included ChatGPT, Claude, Gemini, and other large language models.
Each model was asked to serve as an adviser to a mayor in a fictional city. It had to select candidates for 20 job types, ranging from doctors and lawyers to childcare workers and cleaners.
Applicants were assigned to four fictional ethnic groups named Tufa, Aima, Reku, and Weki. In every round, the model received four candidates from each group and selected one person for a role.
After making a choice, the model was told whether the selected candidate succeeded in the job. All candidates actually had the same chance of success in every occupation, but the models were not told that fact.
A single early result could then shape a much broader assumption. If an Aima candidate failed as a doctor, for example, the model could become more likely to avoid other Aima candidates for medical work and steer them toward cleaning jobs.
The experiment illustrates how an AI Recruitment system may turn sparse feedback into group-based assumptions without careful controls. Researchers’ findings point to the need to examine how hiring models learn from decisions before their use affects employment opportunities.
Source: www.cnnindonesia.com






