Analysis: 13+ datasets used by tech companies without permission to train AI models contain 15.8M+ YouTube videos from 2M+ channels, including 1M how-to videos
www.theatlantic.com/technology/ a... Forums: r/technology : AI Is Coming for YouTube Creators | At least 15 million videos have been snatched by tech companies. See also Mediagazer
AI Watchdog is indeed an excellent resource and is, for example, how I know that my books have been ingested into the LibGen dataset and used without my consent to train commercial AI systems [embedded post]
I'm proud to share a new project that @alexreisner.bsky.social and I launched at @theatlantic.com today! It's called AI Watchdog, and it's our new home for all of the investigations into training data sets, such as LibGen, Books3, and OpenSubtitles. www.theatlantic.com/category/…
the ingenuous scholarly and amateur instructional material on youtube, and its large proportion of overall content there, is one of the 21st century's great collective achievements and its enclosure it for the profit of the tiny grifter few is very terrible [embedded post]
This is a really important resource and reporting cataloguing the scale at which AI is scraping the internet and allowing people to search these datasets — www.theatlantic.com/technology/ a...