Tests show GPT-3.5 and GPT-4 systematically produce biases that disadvantage protected groups, based on their names, when screening and ranking job candidates
Recruiters are eager to use generative AI, but a Bloomberg experiment found bias against job candidates based on their names alone
Context & Ripple Effects
This finding extends a longer record of automated employment tools creating unequal outcomes: earlier coverage found that algorithmic application screening could disproportionately exclude poorer applicants (algorithmic hiring screens that penalized poorer applicants).
It also puts a hiring-specific test alongside evidence that generative systems can amplify demographic stereotypes in job-related imagery (stereotypes in AI-generated job imagery), despite OpenAI having used external experts to probe GPT-4 for prejudice and other risks before release (pre-release bias testing for GPT-4).
First-order effects
- Recruiters using GPT-3.5 or GPT-4 to screen or rank applicants face an immediate risk that name signals, rather than job-relevant qualifications, influence candidate treatment.
- The results give employers a concrete reason to pause or constrain automated ranking workflows and test outputs before they shape hiring decisions.
Second-order effects
- AI hiring-tool buyers are likely to demand auditable evaluations and human review controls from vendors, shifting deployment risk from a model demo to the employer's operating process.
- Model providers and recruitment-software vendors face pressure to show whether mitigations hold in high-stakes workflows, not merely that a general-purpose model can generate useful text.
Third-order effects
- If organizations continue adopting general-purpose models in employment decisions, hiring may become a leading test case for operational AI governance: documented evaluation, escalation paths, and accountability for outcomes.
- The broader shift is from treating model bias as a pre-release safety issue to treating it as a deployment-control problem shared by model makers and the organizations that use them.
The trend: Generative AI is moving into consequential workplace decisions faster than governance practices can establish whether automated outputs are equitable and accountable.