Meta releases Segment Anything Model 2 with support for object segmentation in videos and images; the code and weights are available under an Apache 2.0 license
Meta had a palpable hit last year with Segment Anything, a machine learning model that could quickly and reliably identify and outline just about anything in an image.
TechCrunchDevin Coldewey
Context & Ripple Effects
Meta’s original Segment Anything release and 1-billion-mask dataset established the project as a reusable computer-vision building block for object identification. SAM 2 extends that arc from still images to video, while retaining a permissive distribution model.
The release also sits alongside Meta’s work on video representation learning, including V-JEPA’s approach to learning from masked video. It makes segmentation a more accessible component for developers building video-analysis workflows.
First-order effects
Developers and researchers can use SAM 2’s code and weights under Apache 2.0, reducing licensing friction for image- and video-segmentation experiments and deployments.
Meta broadens Segment Anything from image outlining to object segmentation in video, giving users a single model family for both media types.
Second-order effects
Video-tooling providers and computer-vision teams can incorporate a widely available segmentation layer rather than building that capability from scratch, increasing pressure to differentiate on workflow integration, data, or application-specific performance.
The availability of weights makes deployment choices more flexible: teams can evaluate or run the model in their own environments instead of depending solely on a hosted inference endpoint.
Third-order effects
If open-weight vision models continue to add temporal capabilities, foundational segmentation is likely to become commodity infrastructure, shifting competitive value toward data pipelines, product interfaces, and specialized downstream models.
This is one data point in an open release strategy for core vision tooling that can expand the ecosystem of complementary applications around Meta’s research assets, though adoption will depend on real-world performance and integration cost.
The trend: Open-weight computer-vision models are moving from static-image primitives toward reusable video-understanding components, widening the market for complementary tools and applications.
We're also announcing the open source release of Meta's Segment Anything Model 2 (SAM 2), our state-of-the-art segmentation model which introduces high-quality, fully-promptable object tracking for videos for the first time. The model operates in real-time and also performs bett…
@BenjaminDEKR I can see SAM 2 being a game changer in fintech and healthcare industries. Imagine walking into a hospital room with AR glasses, and SAM 2 recognizes the medical equipment, providing real-time instructions to medical staff.
Memory Attention: adding object permanence with $50k in compute @AIatMeta continues to lead Actually Open AI. SAM2 generalizes SAM1 from image segmentation to video, releasing task, model, and dataset as Apache 2! Notable aspects from reading the paper: - shockingly efficient: [i…
SAM 2 looks very interesting, and more importantly, very performant in terms of speed. Of note, the “video” implementation takes in a series of frames, which means it's possible to integrate it with a webcam stream, which opens *possibilities* https://ai.meta.com/sam2/
SAM was an important gut check for the earth observation industry - fine tuning it yielded close to SOTA results with much, much less work than the SOTA approaches. I wonder if fine tuning SAM 2 will eclipse alternatives. Cool moment for the industry.