Facebook announces a project to train its AI to understand what is happening in videos, using data from videos publicly available on Facebook
starting with Reels' recommendations: https://ai.facebook.com/... https://twitter.com/... Yann LeCun / @ylecun : Another Self-Supervised Learning project at Facebook AI. This time from video and joint video+audio+text. https://twitter.com/... Lance Ulanoff / @lanceulanoff : I can hear the AI now, “So many cat videos.” https://www.theverge.com/... @mattnavarra : “[its] part of our broader efforts toward building machines that learn like humans do” gulp! https://www.theverge.com/... Adrian Hon / @adrianhon : Pleased to “worldscraping” take off as an AR term for capturing (and monetising) information about the world with smart glasses! I was surprised it hadn't been used before my initial blog post tbh https://www.theverge.com/... https://twitter.com/... Victor Zambrano / @argonaut : Really? 75% of the time we don't even know what's happening in a tiktok. But Facebook's evil AI will know. Right. https://twitter.com/... James Vincent / @jjvincent : Facebook is training its AI on users' videos, teaching them to understand content like humans do (or so it hopes). It's a vague announcement, but one that could be incredibly significant https://www.theverge.com/... James Vincent / @jjvincent : Immediate uses cases include improving content recommendation and moderation. But Facebook also hints at integrating AI video analysis with its upcoming smartglasses https://www.theverge.com/... https://twitter.com/...
Context & Ripple Effects
Facebook had already applied computer vision to automated video tagging and used Rosetta to read text in images and video frames for search and harmful-content detection. The new project extends that work from recognizing isolated people or text to learning jointly from video, audio, and text.
Its initial use in Reels recommendations makes the effort a distribution product as well as a research program. Later coverage of a video recommendation system spanning Facebook’s services shows how video understanding became central to the company’s ranking stack.
First-order effects
- Facebook can use publicly available Facebook videos as training material for a model intended to improve Reels recommendations, tying model development directly to an existing short-video surface.
- Reels viewers and creators are the immediate audience affected: the stated first deployment is a recommendation system informed by richer understanding of video content.
Second-order effects
- Facebook’s earlier tagging, text extraction, and toxic-content efforts gain a more general multimodal foundation, potentially reducing the need to treat visual, audio, and text signals as separate inputs across its products.
- For creators, publicly shared video becomes both publishable media and an input to Facebook’s recommendation models, increasing the value of content signals that the platform can interpret.
Third-order effects
- The project points toward social platforms competing on proprietary multimodal training data and the ability to turn it into ranking quality, rather than on recommendation logic alone.
- As video-understanding models are applied across discovery and safety functions, the boundary between content distribution and content analysis becomes structurally thinner.
The trend: Social platforms are turning user-posted video into multimodal inference input for recommendation, search, and content-safety systems.