SAM 2.1 Hiera-Large
Release in the Segment Anything (SAM) family · version 2.1 (sam2.1_hiera_large)
Maintained by Meta (FAIR)1
SAM 2.1 Hiera-Large is the largest checkpoint (224.4M parameters, per the README) in Meta's SAM 2.1 suite, an improved set of SAM 2 checkpoints released in September 2024. SAM 2 is a transformer with streaming memory that segments objects in images and tracks them through video from point, box, or mask prompts.12
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: SAM 2 repository (external site: github.com)
- License: LICENSE (Apache 2.0) (external site: github.com)
- Paper: SAM 2: Segment Anything in Images and Videos (arXiv 2408.00714) (external site: arxiv.org)
- Release notes: SAM 2 release notes (external site: github.com)
Availability and license
Overall availability
The checkpoint downloads directly from Meta's servers or from Hugging Face without gating.110
Availability is separate from permission: read the license before using or redistributing.
The README states that the SAM 2 model checkpoints, demo code, and training code are licensed under Apache 2.0. The fonts bundled with the web demo are under the SIL Open Font License 1.1, and an optional connected-components post-processing module adapted from cc_torch carries its own BSD 3-Clause license.14
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Direct download links in the README and an ungated Hugging Face repository.110 |
| Inference codeIs code for running the model published? | Public | The repository provides image and video predictors and notebooks; the Hugging Face card also documents use with Transformers.19 |
| Training codeIs the code used to train the model published? | Public | Training and fine-tuning code was released with SAM 2.1, including the trainer, loss functions, dataset loaders, and launch scripts for single- and multi-node jobs.25 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The paper lists the training mix as SA-1B, the SA-V dataset, an internal video dataset, and open-source video datasets. SA-V is downloadable under CC BY 4.0, but the internal data is not released.86 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The paper's appendix tabulates pre-training and full-training hyperparameters, and the paper states its results use SAM 2.1, but it gives few details of what changed from the July 2024 checkpoints.8 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The README reports SA-V test, MOSE val, and LVOS v2 results for each SAM 2.1 checkpoint and a benchmarking script for speed. The repository publishes a VOS inference script documented with the SAM 2.1 configs and checkpoints (with a flag for LVOS-style datasets) and an SA-V evaluator for the released val and test sets; MOSE and LVOS scores come from those datasets' own evaluation tools or servers.176 |
What it is useful for
Promptable segmentation of objects in images, automatic mask generation, and segmenting and tracking objects across video frames.1
Run and use notes
- The README requires Python 3.10 or later with PyTorch 2.5.1 and torchvision 0.20.1 or later and notes that installation compiles an optional CUDA extension; it recommends WSL with Ubuntu on Windows.1
Organization context
Other releases in the Segment Anything (SAM) family
- SAM 3.1Model-disclosure tier (USASI rubric v0.1): Restricted weights
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.