/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Stack Overflow CEO says the company plans to charge large AI developers for access to the 50M questions and answers on its service as soon as mid-2023

The programmer Q&A site joins Reddit in demanding compensation when its data is used to train algorithms and ChatGPT-style bots Mastodon: @paul@social.paulkedrosky.com . Tweets: @nwaisb , @jason , @evijitghosh , @abacaj , @oliverjumpertz , @sarthakgh , @vivek7ue , @peard33 , @technologypoet , and @ruchowdh . Thanks: @tsimonite Mastodon: Paul Kedrosky / @paul@social.paulkedrosky.com : More and more companies belatedly realizing they own valuable AI training sets, and that they have had them looted.  Stack Overflow joins the list of companies newly charging for AI training access. … Tweets: Noah Waisberg / @nwaisb : LLMs raise sooooo many interesting IP issues. One big one: were previous un-paid/un-permissioned training runs done on sites including these breach of copyright or breach of contract (the contract being the website terms and conditions)? https://twitter.com/... @jason : The list of data holders who will require AI products to license datasets grows: Reddit, Twitter, & now Stackoverflow https://www.wired.com/... Avijit Ghosh / @evijitghosh : This money is coming back to users who create the actual content. Right? RIGHT? https://www.wired.com/... Anton / @abacaj : Looks like the data moats are now making moves to defend that https://twitter.com/... Oliver Jumpertz / @oliverjumpertz : I said it first, and now it has happened. You just cannot expect to scrape data from all over the internet to build LLMs without compensating the platforms and people who created that content in the first place. https://www.wired.com/... Sar Haribhakti / @sarthakgh : StackOverflow and Reddit want to charge for their UGC content Quora wants to own the consumer interface by working with multiple model providers https://www.wired.com/... Vivek Raghunathan / @vivek7ue : Another shoe drops ... https://www.wired.com/... Wonder how Reddit and Stack will draw the line between who pays and who doesn't? What about open-source LLM efforts like @AiEleuther GPT-J/GPT-Neo or @StanfordCRFM Alpaca or @FacebookAI Llama / OPT or @databricks Dolly? Paresh Dave / @peard33 : Stack Overflow has 50 million questions and answers related to computer programming on its service. They are all available free for others to use under a Creative Commons license. But Stack's CEO says big tech companies are violating the license https://www.wired.com/... Vanessa Harris / @technologypoet : Subscription fees will save us from super advanced AI. https://twitter.com/... @ruchowdh : Seeing a trend now between Reddit and Stack Overflow - it seems the only party NOT cashing in on tech billions are users - you know, the actual content generators. https://www.wired.com/... Thanks: @tsimonite

Wired Paresh Dave

Discussion

  • @jason @jason on x
    The list of data holders who will require AI products to license datasets grows: Reddit, Twitter, & now Stackoverflow https://www.wired.com/...
  • @evijitghosh Avijit Ghosh on x
    This money is coming back to users who create the actual content. Right? RIGHT? https://www.wired.com/...
  • @abacaj Anton on x
    Looks like the data moats are now making moves to defend that https://twitter.com/...
  • @oliverjumpertz Oliver Jumpertz on x
    I said it first, and now it has happened. You just cannot expect to scrape data from all over the internet to build LLMs without compensating the platforms and people who created that content in the first place. https://www.wired.com/...
  • @ruchowdh @ruchowdh on x
    Seeing a trend now between Reddit and Stack Overflow - it seems the only party NOT cashing in on tech billions are users - you know, the actual content generators. https://www.wired.com/...
  • @sarthakgh Sar Haribhakti on x
    StackOverflow and Reddit want to charge for their UGC content Quora wants to own the consumer interface by working with multiple model providers https://www.wired.com/...
  • @vivek7ue Vivek Raghunathan on x
    Another shoe drops ... https://www.wired.com/... Wonder how Reddit and Stack will draw the line between who pays and who doesn't? What about open-source LLM efforts like @AiEleuther GPT-J/GPT-Neo or @StanfordCRFM Alpaca or @FacebookAI Llama / OPT or @databricks Dolly?
  • @peard33 Paresh Dave on x
    Stack Overflow has 50 million questions and answers related to computer programming on its service. They are all available free for others to use under a Creative Commons license. But Stack's CEO says big tech companies are violating the license https://www.wired.com/...
  • @technologypoet Vanessa Harris on x
    Subscription fees will save us from super advanced AI. https://twitter.com/...