/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Over 500 AI researchers are working on a project to build an open-source large language model that will be used to conduct research independent of any company

Hundreds of scientists around the world are working together to understand one of the most powerful emerging technologies before it's too late.Source:BigScience WorkshopandBigScience WorkshopTweets:@jon_severs,@lucianabenotti,@techreview,@_karenhao,@phalcon7,@ict4peace,@techreview,@techreview,@mor10,@glichfield, and@quantamagazineSource:BigScience Workshop:Frequently Asked Questions (FAQ)BigScience Workshop:The Summer of Language Models 21 (BigScience)Tweets:Jon Severs /@jon_severs:This is a ver

MIT Technology Review Karen Hao

Context & Ripple Effects

This story opens the arc that ends with BigScience releasing BLOOM: after OpenAI's GPT-3 made state-of-the-art language models effectively corporate property, over 500 researchers organized as the BigScience Workshop to build an equivalent they could study without any company's permission. The project is also a direct answer to the concerns laid out in Timnit Gebru's draft paper on the risks of large language models — if the risks are real, researchers need access to audit them.

First-order effects

  • Independent academics get a path to study large-model behavior without signing agreements with OpenAI or other commercial gatekeepers, since the model will be open-source by design.
  • Hugging Face emerges as the organizing hub for a distributed research effort spanning hundreds of contributors across institutions.

Second-order effects

  • Startups building general-purpose language tools for non-English markets — the effort covered in Wired's look at Chinese, South Korean, Israeli, and German teams following GPT-3 — gain an open reference model they can build on rather than competing against closed APIs alone.
  • Corporate labs face a new accountability dynamic: once an open-access model of comparable scale exists, claims about proprietary systems can be checked against something researchers can inspect directly.

Third-order effects

  • If the pattern holds — the workshop grew from 500 to over 900 researchers by early 2022 per VentureBeat's follow-up — frontier-scale research splits into two tracks, corporate and commons-based, with open models becoming shared infrastructure rather than one-off releases.
  • The 2026 reporting on researchers at OpenAI, Anthropic, and Google treating LLMs as subjects of scientific study shows where this leads: model understanding itself becomes a field that depends on access norms BigScience helped establish.

The trend: Frontier language-model research is splitting into proprietary lab tracks and commons-based open-access efforts, with BigScience the template for the latter.

Discussion

  • @jon_severs Jon Severs on x
    This is a very good article, on very scary stuff - it also details a company called HUGGINGFACE. And Huggingface is sort of the hero of the piece. https://www.technologyreview.com/ ...
  • @lucianabenotti Luciana Benotti on x
    “Soon enough, all of our digital interactions—when we email, search, or post on social media—will be filtered through large language models”. https://www.technologyreview.com/ ...
  • @techreview @techreview on x
    Google recently announced an AI system that can chat to users about any subject, but didn't discuss the ethical debate surrounding such cutting-edge systems. Studies have already shown how racist, sexist, and abusive ideas are embedded in these models. https://www.technologyrevie…
  • @_karenhao Karen Hao on x
    Ever since Google fired @timnitGebru & @mmitchell_ai, it's continued to deploy the very technology it punished them for scrutinizing. Now hundreds of scientists are racing to investigate the technology's risks before it's too late to avoid its harms. https://www.technologyreview.…
  • @phalcon7 Phyllis D.K. Hildreth on x
    This. The danger of a single scientific story.... Large language models (LLMs) https://twitter.com/... https://twitter.com/...
  • @ict4peace @ict4peace on x
    Very concerning use of large language models (LLM). „Unfortunately, very little research is being done to understand how the flaws of this technology could affect people in real-world applications, or to figure out how to design better LLMs that mitigate these challenges." https:…
  • @techreview @techreview on x
    Google fired its ethical AI co-leads after they raised concerns about the racist, sexist and abusive ideas embedded in one of its most prized AI technologies. The company recently unveiled ambitious new plans to deploy this technology across its products. https://www.technologyre…
  • @techreview @techreview on x
    Hundreds of scientists around the world are working together to understand one of the most powerful emerging technologies before it's too late. https://www.technologyreview.com/ ...
  • @mor10 Morten Rand-Hendriksen on x
    “LLMs (Large Language Models) are increasingly being integrated into the linguistic infrastructure of the internet atop shaky scientific foundations.” “The race to understand the thrilling, dangerous world of language AI” https://www.technologyreview.com/ ...
  • @glichfield Gideon Lichfield on x
    “Soon enough, all of our digital interactions—when we email, search, or post on social media—will be filtered through LLMs.” An important piece from @_KarenHao about an increasingly core piece of tech infrastructure https://twitter.com/...
  • @quantamagazine Quanta on x
    Tech companies use programs that read and write without understanding. But researchers are studying the disturbing limits of these glorified autocomplete functions, @_KarenHao writes for @techreview. https://www.technologyreview.com/ ...