V-JEPA 2 ViT-g/16 (384 px)
Release in the V-JEPA 2 family · version 2 ViT-g/16, 384 resolution (vjepa2-vitg-fpc64-384)
Maintained by Meta (FAIR)31
The 1-billion-parameter ViT-g/16 encoder from Meta's V-JEPA 2, in the version that takes 64-frame clips at 384-pixel resolution. It was pretrained with self-supervision on VideoMix22M, a mix of about 22 million video and image samples. Meta's action-conditioned V-JEPA 2-AC world model was post-trained from a ViT-g/16 V-JEPA 2 encoder.137
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: vjepa2 repository (external site: github.com)
- License: LICENSE (MIT, repository) (external site: github.com)
- Paper: V-JEPA 2 paper (arXiv 2506.09985) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable without gating from Hugging Face and as a direct checkpoint link in the repository README.23
Availability is separate from permission: read the license before using or redistributing.
MIT License (vjepa2 repository) (external site: github.com)43
Apache License 2.0 (Hugging Face model card) (external site: huggingface.co)12
The repository README says most of V-JEPA 2 is licensed under MIT, with three data-augmentation and worker-initialization files under Apache 2.0, but does not name a separate license for the checkpoints linked from the README. The Hugging Face card for this checkpoint declares Apache-2.0. Meta's announcement says the code and checkpoints are available for commercial and research applications.318
Model-disclosure tier
Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Ungated Hugging Face repository and direct download link.23 |
| Inference codeIs code for running the model published? | Public | The repository provides PyTorch Hub loaders and a demo; the model card documents use with Hugging Face Transformers.31 |
| Training codeIs the code used to train the model published? | Public | The repository publishes the V-JEPA 2 pretraining loop and launch commands, with pretraining and cooldown configurations for ViT-g/16 (including a 384-pixel cooldown).35 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The paper lists VideoMix22M's components (Something-Something v2, Kinetics, HowTo100M, YT-Temporal-1B, and ImageNet) and says all sources are publicly available. The YT-Temporal-1B portion was filtered by retrieval-based curation; this review did not find the curated subset list published.7 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | Pretraining and cooldown configuration files are published, and the paper documents pretraining hyperparameters in its appendix.57 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The repository publishes attentive-probe evaluation code, evaluation configs for the ViT-g 384-pixel model (for example SSv2, Diving48, EPIC-KITCHENS-100, Kinetics-400, ImageNet-1k), and trained probe checkpoints.36 |
What it is useful for
The model card describes use as a video and image encoder for video classification, retrieval, and as the video encoder of vision-language models.1
Run and use notes
Organization context
Other releases in the V-JEPA 2 family
- V-JEPA 2.1 ViT-G/16 (384 px)Model-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.