CLIP ViT-L/14
Release in the CLIP family · version ViT-L/14
CLIP ViT-L/14 pairs a ViT-L/14 Vision Transformer image encoder with a masked self-attention Transformer text encoder, trained contrastively to match images with their captions. OpenAI released it in January 2022; a variant further trained at 336-pixel resolution (ViT-L/14@336px) followed in April 2022.139
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: openai/CLIP repository (external site: github.com)
- License: LICENSE (MIT, code) (external site: github.com)
- Paper: CLIP paper (arXiv 2103.00020) (external site: arxiv.org)
Availability and license
Overall availability
Downloadable without gating from Hugging Face, and fetched by the openai/CLIP package's clip.load("ViT-L/14"). No license for the weights is stated (see license notes).26
Availability is separate from permission: read the license before using or redistributing.
The repository's MIT license covers "the Software" (the code). This review found no license statement for the model weights in the repository README, the model card, or the Hugging Face card, whose metadata carries no license field. Separately from licensing, the model card places all deployed uses out of scope.5432
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
No license for the weights is recorded in this catalog.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published on Hugging Face without gating and downloadable through clip.load().26 |
| Inference codeIs code for running the model published? | Public | The openai/CLIP package loads the model and encodes images and text; the Hugging Face card documents use with Transformers.41 |
| Training codeIs the code used to train the model published? | Unknown | The repository provides loading, inference, and evaluation examples; this review found no published training code.4 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The paper describes a dataset of 400 million image-text pairs gathered from public internet sources using 500,000 search queries. The model card says OpenAI will not release the dataset; only a list identifying a YFCC100M subset used in an ablation is published.938 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The paper describes training for 32 epochs with Adam, decoupled weight decay, a cosine schedule, and a 32,768 minibatch, and says the largest Vision Transformer took 12 days on 256 V100 GPUs. No training configuration files are published.9 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The repository publishes the class names and prompt templates used for the paper's zero-shot results, an ImageNet prompt-engineering notebook, and a linear-probe example.74 |
What it is useful for
Run and use notes
Organization context
Other releases in the CLIP family
No other releases in this family have been assessed.
U.S. eligibility
Eligible · basis: U.S. headquarters
The model card states that CLIP was developed by researchers at OpenAI, and the checkpoint is published by OpenAI in its openai/CLIP repository and Hugging Face account. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.3610
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.