Facebook details DINO self-supervised AI for object discovery and segmentation in images/videos, and PAWS, a new ML approach to classify poorly labelled images
VentureBeatKyle Wiggers
Context & Ripple Effects
DINO and PAWS are the latest step in a deliberate arc: two months ago Facebook showed that its SEER model trained on a billion public Instagram images beat fully supervised baselines, then announced a project to train AI on what happens in videos from its own platform. DINO extends self-supervision to discovering and segmenting objects in stills and video without labels, while PAWS attacks the opposite bottleneck — datasets whose existing labels are sparse or unreliable.
The through-line runs back further than this spring: since deploying Rosetta to read text inside images and video frames in 2018, Facebook has been converting its media corpus into structured understanding. What changed with DINO and PAWS is that the labeling work itself is being automated, closing the loop between the data Facebook uniquely holds and the models trained on it.
First-order effects
Teams working on image and video understanding get two concrete tools: DINO removes the need to hand-label objects for discovery and segmentation, and PAWS makes classification viable on poorly labelled datasets where supervised training stalls.
Second-order effects
Rivals with large unlabeled media archives — YouTube-scale video platforms especially — face pressure to match the same self-supervised approach, while commercial data-annotation suppliers lose pricing power as label-free methods cover more vision tasks.
Third-order effects
If the pattern holds, computer vision consolidates around platforms that own both the raw media and the compute to train on it, widening the gap between data-rich incumbents and everyone else — with questions about consent for training on user-posted material unresolved.
The trend: Frontier computer vision is shifting from human-labeled datasets toward self-supervised training on platform-owned media, making proprietary data access the decisive research advantage.
In addition to sharing DINO, a self-supervised model that can discover and segment objects in an image or video with no supervision, we're also sharing PAWS, a new method for 10x more efficient training. Get the code: https://ow.ly/... #computervision https://twitter.com/...
Raise your eyebrow if you understand they are definitely using the same algorithms to profile people based on skin color and facial features. (*raises eyebrow*) https://twitter.com/...
And there's more! It's not just that these AI models perform better in so many ways — the time and computing power needed to train them is an order of magnitude less. Our team wrote all about it today: https://ai.facebook.com/...
and yet you still have a worse problem of it than any other platform, and the company internal comms have said it is intentionally ignored as they like the bump in engagement it offers. we remember mark using the “amazing tech will fix this all” defense before congress *eyeroll* …
Object segmentation is considered one of the hardest challenges in computer vision because it requires that AI truly understands what is in an image. @FacebookAI is publishing new work today to show advances in this field using self-supervised learning: https://ai.facebook.com/..…
New AI/computer vision system “managed to connect categories based on visual properties, a bit like humans do. For example, we see that animal species are clearly separated, with a coherent structure that resembles the biological taxonomy” https://ai.facebook.com/... https://twit…
Where is our AI fueled future headed? Here's a hint: The biggest and most expensive AI efforts in the world are all trying to build something that can aggressively censor the conversations of billions of people in real-time. https://twitter.com/...
Thanks for open-sourcing the code! it's pretty amazing what cheap, un-labeled objectives can do thesedays https://github.com/... https://twitter.com/...