Recent research from Princeton University and the University of Chicago suggests that large language models (LLMs) are prone to developing original biases when tasked with hiring, potentially stereotyping applicants more aggressively than humans. The study, presented at the ICML conference in Seoul, utilized a simulated hiring game where models like ChatGPT, Claude, and Gemini were asked to fill various job roles using candidates from four fictional ethnic groups.
In the experiment, all candidates were equally qualified, yet the models quickly began segregating groups into specific job categories based on limited outcomes. For instance, if a model observed a single failure for a candidate from a specific group in a high-skill role, it frequently stopped hiring members of that group for similar positions. On a segregation scale where human participants scored 0.84, models scored significantly higher; OpenAI’s o3 model reached 1.83, nearing the maximum possible score.
Ryan Liu, a PhD student at Princeton and coauthor of the study, noted that LLMs are optimized to generalize from limited data, a trait that aids in solving logic puzzles but leads to problematic stereotyping in social contexts. The study found that newer, more advanced models with higher reasoning capabilities often exhibited stronger biases. Furthermore, attempts to instruct models to be fair were largely ineffective, though offering bonuses for diverse hiring successfully reduced biased behavior.
The research also highlighted that providing models with relevant personal information about candidates, such as education and age, helped mitigate ethnic segregation. However, when models were provided with irrelevant data, such as physical characteristics, they reverted to stereotyping. Experts suggest that as AI systems gain improved memory and personalization, the risk of these models forming and reinforcing novel, unlearned biases in real-world applications like lending or recruitment remains a significant concern.
Source: MIT Technology Review
No comments yet. Be the first to share your thoughts.