Meta SAM 3: Text-Prompt Segmentation, Tracking, and SAM 3D

Meta’s SAM 3 finds and tracks every image or video object matching a phrase, while SAM 3D reconstructs objects and human bodies from one image.

A phrase instead of a click

Meta released SAM 3 on November 19, 2025, alongside a paper describing the underlying task as Promptable Concept Segmentation. Where SAM 1 and SAM 2 took a click, a box, or a mask as a prompt and returned one matching object, SAM 3 accepts a short noun phrase, such as "red baseball cap," or an example crop of an object, and returns masks and tracked identities for every instance in an image or video that matches it. The model pairs an image-level detector with a memory-based video tracker sharing one backbone, and adds a presence head that separates the question of whether the concept appears at all from where it appears, which the paper credits with the detection accuracy gain. [1][2]

To train the presence head and the detector on phrases rather than clicks, Meta built a data engine that produced a dataset of 4 million unique concept labels, including hard negative examples, across images and video, and released a matching evaluation set, the SA-Co benchmark, along with the model weights. [2]

A doubling, by Meta's own measure

Meta's paper reports that SAM 3 roughly doubles accuracy on the SA-Co benchmark compared to the systems it compares against, in both the image and video settings, and describes further gains on zero-shot LVIS detection and on object-counting tasks. Independent coverage published two days later described the same figure and illustrated the underlying change with a safari clip: SAM 2 needed a separate click for each elephant in frame, while SAM 3 finds and tracks every elephant from the single word. [2][3]

Atlas interpretation: The doubling figure is Meta's own comparison against systems Meta selected, on a benchmark Meta built for exactly this task, so it demonstrates that concept segmentation is a real improvement over asking a click-based model to approximate the same job rather than a score an outside lab reproduced against SAM 3 on neutral ground. [2]

A same-day release for reconstructing objects and bodies

Meta shipped SAM 3D as a separate pair of models the same day: one for reconstructing the shape and layout of objects and scenes from a single photograph, and one for estimating human body shape and pose from a single image. Meta released checkpoints and inference code for both and introduced SAM 3D Artist Objects, an evaluation set built with artists, saying the models substantially outperform prior methods on it. [1]

Both SAM 3 and SAM 3D reached Meta's own products on release day: SAM 3 powers new effects in the Edits video app and the Vibes feature on Meta AI, and SAM 3D underlies a View in Room feature that lets Facebook Marketplace shoppers place a listed piece of furniture into a photo of their own space. [1]

Sources

  1. New Segment Anything Models Make it Easier to Detect Objects and Create 3D Reconstructions

    Meta · Nov 19, 2025

  2. SAM 3: Segment Anything with Concepts

    arXiv · Nov 20, 2025

  3. Meta AI's New Segment Anything Model: Exploring SAM 3

    Ultralytics · Nov 21, 2025