A drop-in replacement, not a new model
SAM 3.1 is Meta's update to SAM 3, the version of Segment Anything that finds and tracks every instance matching an open-vocabulary phrase rather than one object per click. SAM 3.1 keeps that capability and is positioned as a drop-in replacement, aimed specifically at making video tracking faster rather than adding a new detection capability. [1]
Before this update, SAM 3 processed each tracked object through the model separately, one forward pass per object, even when a video held several instances of the same class at once. That per-object repetition is what the multiplexing change removes. [1]
Sixteen objects, one pass, shared context
Object multiplexing lets SAM 3.1 process up to sixteen tracked objects together in a single forward pass instead of running the model once per object, which Meta describes as eliminating redundant computation and memory bottlenecks. The paired mechanism, described as global reasoning, gives the model inter-object communication: shared context across the tracked objects in a scene, rather than each one being scored in isolation. [1]
Atlas interpretation: The two changes solve different problems. Multiplexing is a throughput fix, batching work the model was already doing. Global reasoning is an accuracy fix aimed at the case that breaks a per-object tracker: a scene with several similar-looking instances of the same class, where knowing about the other fifteen objects helps the model tell them apart and avoid swapping identities between them. [1]
Meta states that inference latency scales with the number of tracked objects, and that the model holds close to real-time performance up to roughly five concurrent objects in video, short of the sixteen-object ceiling the multiplexing pass supports. [1]
16 to 32 frames per second, on one GPU
Meta's headline throughput number is a doubling: a mid-density video, meaning a scene with a moderate number of simultaneously tracked objects rather than one or a crowd, goes from 16 to 32 frames per second on a single H100 GPU. [1]
Atlas interpretation: That figure is Meta's own measurement, on Meta's own definition of mid-density, run on the specific hardware Meta chose to quote. No independent benchmark of SAM 3.1's throughput was found at the time of this check, so the 32 frames per second figure describes what Meta reports, not a number a third party has replicated. [1]
Weights on Hugging Face
Meta published SAM 3.1 checkpoints on Hugging Face alongside the announcement, continuing the pattern set by SAM 3, SAM 3D and the earlier SAM 2, each released with open weights rather than held back as an API-only model. [1]