FastVLM
Family overview — summarizes releases; licenses and availability belong to each release.
FastVLM is a family of vision-language models from Apple, described in a CVPR 2025 paper, built around FastViTHD, a hybrid convolutional-transformer vision encoder that outputs fewer visual tokens to reduce encoding time for high-resolution images. Apple released 0.5B, 1.5B, and 7B variants that pair the encoder with Qwen2 language models, with PyTorch checkpoints and versions exported for Apple silicon.145
- Repository: ml-fastvlm repository (external site: github.com)
- Model hub: FastVLM 7B on Hugging Face (external site: huggingface.co)
- Paper: FastVLM: Efficient Vision Encoding for Vision Language Models (arXiv 2412.13303) (external site: arxiv.org)
- Website: Apple Machine Learning Research article (external site: machinelearning.apple.com)
Releases assessed
| Release | Weights | Tier (USASI rubric v0.1) | License | Released |
|---|---|---|---|---|
| FastVLM 7B | Weights: Public | Model-disclosure tier (USASI rubric v0.1): Open-weight | Apple Machine Learning Research Model License Agreement, Apple software license (ml-fastvlm code) | Unknown |
What it is useful for
Organization context
U.S. eligibility
Eligible · basis: U.S. headquarters
The model license states the models are developed and released by Apple Inc., and the paper lists Apple as the authors' affiliation. Apple Inc. has its principal executive offices in Cupertino, California, per its Form 10-K for fiscal year 2025. The language model components are third-party Qwen2 models from the Qwen Team at Alibaba Group, recorded on the release record; they are not U.S.-developed.3476
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.