Molmo2-O 7B
Release in the Molmo family · version Molmo2-O-7B
Maintained by Ai2 (Allen Institute for AI)1
Molmo2-O-7B is the Molmo 2 vision-language model built on Ai2's Olmo 3 7B Instruct language model with a SigLIP 2 vision encoder, released in December 2025. It accepts images, multiple images, and video with text, and can answer questions, caption, and point to or track objects. Ai2 describes it as the variant whose full model flow is open, since its language model is also Ai2's.12
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: Training and inference code (allenai/molmo2) (external site: github.com)
- Paper: Molmo 2 paper (arXiv 2601.10611) (external site: arxiv.org)
- Release notes: Molmo 2 announcement (external site: allenai.org)
Availability and license
Overall availability
Downloadable from Hugging Face without gating. The card says the model is intended for research and educational use under Ai2's Responsible Use Guidelines and was trained on third-party datasets that are limited to academic and non-commercial research use.1
Availability is separate from permission: read the license before using or redistributing.
The weights are Apache 2.0, but the card states the model is intended for research and educational use in line with Ai2's Responsible Use Guidelines, and that some of its third-party training data is limited to academic and non-commercial research use, so users should review those sources. Ai2's Molmo2 datasets are ODC-BY; the Molmo2-Cap card adds that it contains text captions generated with GPT-4.1 and GPT-5, subject to OpenAI's terms.176
Model-disclosure tier
Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Final weights are on Hugging Face; the molmo2 repository also links pretraining and SFT checkpoints for Molmo2-O-7B.14 |
| Inference codeIs code for running the model published? | Public | The card documents inference with Transformers 4.57.1 (with trust_remote_code), and the repository documents vLLM inference.14 |
| Training codeIs the code used to train the model published? | Public | The card said training code would follow later; the allenai/molmo2 repository now publishes pretraining, SFT, long-context SFT, and evaluation scripts, with Molmo2-O-7B trained using the olmo3_7b_instruct setting.14 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | Ai2's nine Molmo2 datasets are public on Hugging Face (ODC-BY), and the repository provides download scripts for the academic datasets; some of those must be downloaded manually under their own licensing agreements. The Olmo 3 language model's training data is documented on its own record.1246 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The Molmo 2 paper and the repository describe the pretraining, SFT, and long-context SFT stages.34 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The card reports an average over 15 academic benchmarks and points to the technical report; the repository includes evaluation scripts.14 |
What it is useful for
Video and image question answering, captioning, counting, and grounding (pointing and tracking), as shown in the model card's examples.1
Run and use notes
- The model card's quick start installs transformers 4.57.1 with torch, accelerate, decord2, and molmo_utils, and loads the model with trust_remote_code=True.1
Organization context
Provenance and derivatives
Built by Ai2 from its own Olmo 3 7B Instruct language model and a SigLIP vision encoder published by Google (the card names SigLIP 2 and links google/siglip-so400m-patch14-384), then trained on Ai2's Molmo2 datasets and third-party academic datasets. The paper states the new Molmo2 data was collected without using closed vision-language models; the Molmo2-Cap card states that some of its text captions were generated with OpenAI's GPT-4.1 and GPT-5.136
- Derived from: Olmo 3 7B Instruct — Language model backbone (Ai2).
- Derived from: SigLIP (google/siglip-so400m-patch14-384) (external site: huggingface.co) — Vision encoder published by Google.
- Derived from: Molmo2 datasets (external site: huggingface.co) — Ai2's image, multi-image, and video training data (ODC-BY).
Other releases in the Molmo family
No other releases in this family have been assessed.
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.