NVIDIA Cosmos3-Nano
Release in the NVIDIA Cosmos family · version Cosmos3-Nano (16B)
Cosmos3-Nano is the 16B-parameter model in NVIDIA's Cosmos 3 family, released on May 31, 2026. It shares the family's mixture-of-transformers design, generating text autoregressively and images, video, audio, and robot actions by diffusion, from combinations of text, image, video, audio, and action-trajectory inputs.19
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- License: OpenMDW License 1.1 (external site: openmdw.ai)
- Repository: Cosmos-Framework repository (external site: github.com)
- Paper: Cosmos 3 technical report (external site: research.nvidia.com)
Availability and license
Overall availability
Weights are downloadable from Hugging Face without an access gate under OpenMDW-1.1; the card states the model is ready for commercial and non-commercial use and lists global deployment geography.21
Availability is separate from permission: read the license before using or redistributing.
OpenMDW License Agreement, version 1.1 (OpenMDW-1.1) (external site: openmdw.ai)123
OpenMDW License Agreement, version 1.1 (Cosmos-Framework code) (external site: raw.githubusercontent.com)5
OpenMDW-1.1 grants permission to deal in the model, its parameters, and related data, documentation, and software without restriction under copyright, patent, database, and trade secret rights. Distributors must keep a copy of the agreement and applicable notices, and rights end for anyone who sues claiming the materials infringe a patent or copyright, unless responding to a suit brought first. It places no restrictions on outputs. It is not on this catalog's OSI-approved list.34
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Weights are published on Hugging Face without an access gate.21 |
| Inference codeIs code for running the model published? | Public | The card documents inference with vLLM-Omni, vLLM (through a Cosmos-Framework plugin package), Hugging Face Diffusers, and SGLang Diffusion.16 |
| Training codeIs the code used to train the model published? | Partial | Cosmos-Framework publishes a distributed trainer whose documented recipes cover supervised fine-tuning (including a Nano vision SFT recipe). The report describes the internal training infrastructure used for both towers but does not identify it as Cosmos-Framework, and pre-training configurations were not found.679 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The card summarizes the corpus (about 1.3B data points from 393 dataset entries, drawn from NVIDIA-owned data and public datasets, plus synthetic data) and links a public summary of training content. The report lists five synthetic datasets released on Hugging Face; most training data is not released.19 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The technical report describes the training stages, data mixtures, optimizer settings, learning-rate schedules, and the token and GPU counts used for Nano.9 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card and report publish benchmark results, and the report releases one human evaluation benchmark (Cosmos-HUE); code to re-run all reported evaluations was not verified.19 |
What it is useful for
Run and use notes
- The card lists PyTorch, vLLM-Omni, Diffusers, and SGLang as runtimes; NVIDIA Ampere, Hopper, and Blackwell GPUs; and Linux as the tested operating system. It states that only BF16 precision is tested and gives a recommended vLLM-Omni serving command for an H200 GPU.1
Organization context
Provenance and derivatives
Trained by NVIDIA. The technical report states that Cosmos3-Nano adapts the Qwen3-VL 8B architecture and is initialized from pre-trained Qwen3-VL-8B weights (developed by the Qwen team at Alibaba Cloud), and that visual generation uses the video VAE encoder from Wan2.2-TI2V-5B (published on Hugging Face by Wan-AI). The card lists synthetic training data generated with HiDream-I1, Qwen-Image-2512, and Qwen3-VL.911211
- Derived from: Qwen3-VL-8B (external site: github.com) — Initialization weights for the language and vision backbone.
- Derived from: Wan2.2-TI2V-5B video VAE (external site: huggingface.co) — Video VAE encoder used for visual generation; published by Wan-AI under Apache 2.0.
Other releases in the NVIDIA Cosmos family
- NVIDIA Cosmos3-SuperModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
The model card names NVIDIA as the model developer. NVIDIA's principal executive offices are in Santa Clara, California, per its Form 10-Q for the quarter ended July 26, 2026. The model is initialized from third-party Qwen3-VL weights, recorded under provenance; that base is not U.S.-developed.113912
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.