A comparison between US data and 5,100 Stable Diffusion-generated images related to job title and crime finds the tool amplifies gender and racial stereotypes
of some some logical and objective thinking machine. — In reality generative AI, is just fancy autocomplete based on its training data. So it'll always reflect the biases of its underlying data and be harder to fix than traditional systems. … Tweets: Leonardo Nicoletti / @leonardonclt : 🚨Generative AI has a serious problem with bias🚨 Over months of reporting, @dinabass and I looked at thousands of images from @StableDiffusion and found that text-to-image AI takes gender and racial stereotypes to extremes worse than in the real world. 🧵 1/13 [video] Aamer Zaheer / @azaamerzaheer : Bias and stolen data are the core problems with “AI” not X-risk. https://twitter.com/... Bilal Zuberi / @bznotes : Our biases...amplified by Generative AI. And people seem to be rushing into brute forcing GAI into every field, without paying sufficient attention to inherent risks. One piece of the solution resides with companies like @fiddlerlabs, but we have a LOT more work to do. https://twitter.com/... Davide / @criticofpolecon : wait a god damn minute, are you telling me AI isn't objective but a product of human information? https://twitter.com/... Bardia Khosravi / @khosravi_bardia : This is a serious issue with AI generated content. Other than this Bloomberg post, here is an interesting paper focusing on #StableDiffusion inherent biases: https://arxiv.org/... #Bias #GenerativeAI https://twitter.com/... Casey Becker / @casey_becker : Text-to-image GANs visualise societal bias. Can they be used to study it? It's not quite my field, but it seems like a useful application for an unfortunate phenomenon. https://twitter.com/... Rachel Metz / @rachelmetz : We talk a lot about how AI is biased but don't often get to really see it in a meaningful way. This in-depth piece from @dinabass and @Leonardonclt is a fantastic way in to the ways in which generative AI (in this case @StabilityAI Stable Diffusion) is biased). Check it out! https://twitter.com/... Dina Bass / @dinabass : An analysis of more than 5,000 images created with Stable Diffusion found that it takes racial and gender disparities to extremes — worse than those found in the real world. Here's what @Leonardonclt and I found: https://www.bloomberg.com/... Forums: r/artificial : Humans Are Biased. Generative AI Is Even Worse. Stable Diffusion's text-to-image model amplifies stereotypes about race and gender — here's why that matters
Context & Ripple Effects
Bloomberg's months-long investigation compared 5,100 Stable Diffusion images against US data on job titles and crime and found the model doesn't just mirror stereotypes around gender and race — it exaggerates them past real-world levels, confirming the reporter's framing that these systems are 'fancy autocomplete' that inherits and intensifies their training data's biases.
The finding lands on Stability AI at a moment when its data practices were already under fire: artists whose work was swept into the training set without consent, payment, or notice had already framed the dataset as extracted rather than curated, and Stanford researchers later found the underlying LAION-5B corpus riddled with CSAM, forcing it offline. The bias study adds a third axis — demographic distortion — to the same root cause: nobody audited what went in.
First-order effects
- Stability AI now faces documented evidence its flagship tool skews depictions of jobs and crime by gender and race more severely than reality — a direct liability for any business embedding Stable Diffusion output in hiring, media, or advertising workflows.
Second-order effects
- Enterprises deploying image generators will push vendors for dataset provenance and bias audits before adoption, turning curation quality from an invisible cost into a competitive differentiator among model providers.
Third-order effects
- As bias, consent, and illegal-content findings accumulate around web-scraped corpora like LAION-5B, the industry drifts toward documented, permissioned training data — with regulators likely to treat unaudited scraping as the systemic risk the studies keep proving it to be.
The trend: Generative image models are being pushed from scale-at-any-cost web scraping toward audited, consented datasets as each new study converts training-data shortcuts into commercial and regulatory exposure.