CLIP
Family overview — summarizes releases; licenses and availability belong to each release.
CLIP (Contrastive Language-Image Pre-Training) is an OpenAI family of paired image and text encoders trained with a contrastive objective on 400 million image-text pairs, which lets the model classify images zero-shot from natural-language labels. OpenAI released the models in stages from January 2021 to April 2022, from ResNet-50 and ViT-B/32 up to ViT-L/14 and ViT-L/14 at 336-pixel resolution.124
- Repository: openai/CLIP repository (external site: github.com)
- Documentation: Model card (repository) (external site: github.com)
- Model hub: CLIP ViT-L/14 on Hugging Face (external site: huggingface.co)
- Paper: Learning Transferable Visual Models From Natural Language Supervision (arXiv 2103.00020) (external site: arxiv.org)
Releases assessed
| Release | Weights | Tier (USASI rubric v0.1) | License | Released |
|---|---|---|---|---|
| CLIP ViT-L/14 | Weights: Public | Model-disclosure tier (USASI rubric v0.1): Open-weight | MIT | Jan 2022 |
What it is useful for
The model card describes CLIP as a research output for studying zero-shot image classification, robustness, and bias. It states that any deployed use, commercial or not, is currently out of scope, that surveillance and facial recognition uses are always out of scope, and that use should be limited to English.1
Organization context
U.S. eligibility
Eligible · basis: U.S. headquarters
The model card states that CLIP was developed by researchers at OpenAI, and the models are published in OpenAI's openai/CLIP repository. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.126
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.