Meta release SAM 3, a model for detection, segmentation, and tracking of objects in images and video, and SAM 3D, which can reconstruct objects and humans in 3D
Meta Platforms Inc. today is expanding its suite of open-source Segment Anything computer vision models with the release of SAM 3 and SAM 3D …
Context & Ripple Effects
Meta’s Segment Anything line began with a model and large mask dataset for object identification, then expanded to video-capable segmentation with SAM 2’s image-and-video support. SAM 3 extends that trajectory from delineating objects to detecting, segmenting and tracking them across visual media.
SAM 3D adds a spatial layer to Meta’s broader computer-vision work: Meta had also released V-JEPA 2 for understanding 3D environments and object movement, while earlier research addressed generation of 3D assets. The new releases bring reconstruction into the Segment Anything family.
First-order effects
- Researchers and developers gain open-source tools for object detection, segmentation and tracking in images and video, plus 3D reconstruction of objects and people.
- Meta broadens Segment Anything from a primarily segmentation-oriented model family into a more complete visual-perception stack spanning 2D video and 3D scenes.
Second-order effects
- Teams building video-analysis, robotics or spatial-computing workflows can evaluate a shared set of Meta models for object-level inputs rather than assembling separate perception components.
- Competing computer-vision vendors face greater pressure to differentiate on task accuracy, deployment tooling, data controls or specialized applications as Meta makes broader capabilities available as open source.
Third-order effects
- If follow-on releases continue to combine tracking, scene understanding and reconstruction, open model families could become a common lower layer for applications that treat video as structured, machine-readable data.
- The strategic contest may shift from supplying a single vision model to controlling the surrounding data, workflow and deployment layers; the extent of that shift will depend on adoption and real-world performance.
The trend: Meta is turning the Segment Anything line into a broader open visual-perception platform, connecting object-level understanding across images, video and 3D.