Over 500 AI researchers are working on a project to build an open-source large language model that will be used to conduct research independent of any company
Hundreds of scientists around the world are working together to understand one of the most powerful emerging technologies before it's too late.Source:BigScience WorkshopandBigScience WorkshopTweets:@jon_severs,@lucianabenotti,@techreview,@_karenhao,@phalcon7,@ict4peace,@techreview,@techreview,@mor10,@glichfield, and@quantamagazineSource:BigScience Workshop:Frequently Asked Questions (FAQ)BigScience Workshop:The Summer of Language Models 21 (BigScience)Tweets:Jon Severs /@jon_severs:This is a ver
MIT Technology ReviewKaren Hao
Context & Ripple Effects
This story opens the arc that ends with BigScience releasing BLOOM: after OpenAI's GPT-3 made state-of-the-art language models effectively corporate property, over 500 researchers organized as the BigScience Workshop to build an equivalent they could study without any company's permission. The project is also a direct answer to the concerns laid out in Timnit Gebru's draft paper on the risks of large language models — if the risks are real, researchers need access to audit them.
First-order effects
Independent academics get a path to study large-model behavior without signing agreements with OpenAI or other commercial gatekeepers, since the model will be open-source by design.
Hugging Face emerges as the organizing hub for a distributed research effort spanning hundreds of contributors across institutions.
Corporate labs face a new accountability dynamic: once an open-access model of comparable scale exists, claims about proprietary systems can be checked against something researchers can inspect directly.
Third-order effects
If the pattern holds — the workshop grew from 500 to over 900 researchers by early 2022 per VentureBeat's follow-up — frontier-scale research splits into two tracks, corporate and commons-based, with open models becoming shared infrastructure rather than one-off releases.
The 2026 reporting on researchers at OpenAI, Anthropic, and Google treating LLMs as subjects of scientific study shows where this leads: model understanding itself becomes a field that depends on access norms BigScience helped establish.
The trend: Frontier language-model research is splitting into proprietary lab tracks and commons-based open-access efforts, with BigScience the template for the latter.
This is a very good article, on very scary stuff - it also details a company called HUGGINGFACE. And Huggingface is sort of the hero of the piece. https://www.technologyreview.com/ ...
“Soon enough, all of our digital interactions—when we email, search, or post on social media—will be filtered through large language models”. https://www.technologyreview.com/ ...
Google recently announced an AI system that can chat to users about any subject, but didn't discuss the ethical debate surrounding such cutting-edge systems. Studies have already shown how racist, sexist, and abusive ideas are embedded in these models. https://www.technologyrevie…
Ever since Google fired @timnitGebru & @mmitchell_ai, it's continued to deploy the very technology it punished them for scrutinizing. Now hundreds of scientists are racing to investigate the technology's risks before it's too late to avoid its harms. https://www.technologyreview.…
Very concerning use of large language models (LLM). „Unfortunately, very little research is being done to understand how the flaws of this technology could affect people in real-world applications, or to figure out how to design better LLMs that mitigate these challenges." https:…
Google fired its ethical AI co-leads after they raised concerns about the racist, sexist and abusive ideas embedded in one of its most prized AI technologies. The company recently unveiled ambitious new plans to deploy this technology across its products. https://www.technologyre…
Hundreds of scientists around the world are working together to understand one of the most powerful emerging technologies before it's too late. https://www.technologyreview.com/ ...
“LLMs (Large Language Models) are increasingly being integrated into the linguistic infrastructure of the internet atop shaky scientific foundations.” “The race to understand the thrilling, dangerous world of language AI” https://www.technologyreview.com/ ...
“Soon enough, all of our digital interactions—when we email, search, or post on social media—will be filtered through LLMs.” An important piece from @_KarenHao about an increasingly core piece of tech infrastructure https://twitter.com/...
Tech companies use programs that read and write without understanding. But researchers are studying the disturbing limits of these glorified autocomplete functions, @_KarenHao writes for @techreview. https://www.technologyreview.com/ ...