New research from Princeton University and the University of Chicago shows that large language models (LLMs) like ChatGPT, Claude, and Gemini develop hiring biases more readily than humans. In a simulated hiring exercise involving 20 job openings and candidates from four fictional ethnic groups, the AI models began stereotyping applicants despite equal success rates across groups, according to technologyreview.com.
The study adapted a psychology experiment where each AI model acted as a hiring consultant for a fictional city mayor. Over 40 rounds, the models selected candidates for jobs such as doctors, lawyers, child-care aides, and janitors, receiving feedback on each hire’s success. Although all candidates had equal chances of success, the models started segregating candidates by group and job type, demonstrating bias formation from experience rather than just training data.
This finding highlights concerns about AI fairness in recruitment, as these models not only inherit human biases from training data but also generate new stereotypes through their decision-making processes. As AI firms develop agentic models that retain detailed user information, the risk of reinforcing or amplifying biases in hiring decisions increases. The research underscores the need for careful evaluation of AI tools used in employment contexts.
The study’s results were published on July 20, 2026, by technologyreview.com, emphasizing that AI’s propensity to form biases could impact millions of job applicants screened by automated systems before any human review.