DINOv3
Family overview — summarizes releases; licenses and availability belong to each release.
Maintained by Meta (FAIR)12
DINOv3 is Meta's third generation of self-supervised vision backbones, announced in August 2025. The suite includes a 6.7B-parameter ViT-7B/16, distilled Vision Transformers from 21M to 840M parameters, ConvNeXt models, and two backbones trained on satellite imagery, plus task heads for classification, depth, detection, and segmentation. The web-image models were trained on LVD-1689M, a curated set of about 1.7 billion images.142
- Repository: DINOv3 repository (external site: github.com)
- Model hub: DINOv3 collection on Hugging Face (external site: huggingface.co)
- Paper: DINOv3 paper (arXiv 2508.10104) (external site: arxiv.org)
- License: DINOv3 License (external site: github.com)
Releases assessed
| Release | Weights | Tier (USASI rubric v0.1) | License | Released |
|---|---|---|---|---|
| DINOv3 ViT-7B/16 (LVD-1689M) | Weights: Partial | Model-disclosure tier (USASI rubric v0.1): Restricted weights | DINOv3 License | Aug 14, 2025 |
What it is useful for
Frozen image features for downstream tasks such as classification, retrieval, depth estimation, semantic segmentation, and video tracking; the model card recommends fine-tuning only as a last resort.2
Organization context
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.