NVIDIA Cosmos3-Super
Release in the NVIDIA Cosmos family · version Cosmos3-Super (64B)
Cosmos3-Super is the 64B-parameter model in NVIDIA's Cosmos 3 family, released on May 31, 2026. It uses a mixture-of-transformers design with an autoregressive tower that generates text and a diffusion tower that generates images, video, audio, and robot actions, and it takes combinations of text, image, video, audio, and action-trajectory inputs.19
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- License: OpenMDW License 1.1 (external site: openmdw.ai)
- Repository: Cosmos-Framework repository (external site: github.com)
- Paper: Cosmos 3 technical report (external site: research.nvidia.com)
Availability and license
Overall availability
Weights are downloadable from Hugging Face without an access gate under OpenMDW-1.1; the card states the model is ready for commercial and non-commercial use and lists global deployment geography.21
Availability is separate from permission: read the license before using or redistributing.
OpenMDW License Agreement, version 1.1 (OpenMDW-1.1) (external site: openmdw.ai)123
OpenMDW License Agreement, version 1.1 (Cosmos-Framework code) (external site: raw.githubusercontent.com)5
OpenMDW-1.1 grants permission to deal in the "Model Materials" (model architecture and parameters plus related data, documentation, and software) without restriction under copyright, patent, database, and trade secret rights. Distributors must keep a copy of the agreement and applicable notices, and rights end for anyone who sues claiming the materials infringe a patent or copyright, unless responding to a suit brought first. It places no restrictions on outputs. The technical report calls it the Linux Foundation's OpenMDW-1.1 License. It is not on this catalog's OSI-approved list.349
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Weights are published on Hugging Face without an access gate.21 |
| Inference codeIs code for running the model published? | Public | The card documents inference with vLLM-Omni, vLLM, Hugging Face Diffusers, and SGLang Diffusion, and uses the Cosmos-Framework package for prompt upsampling.16 |
| Training codeIs the code used to train the model published? | Partial | Cosmos-Framework publishes a distributed trainer whose documented recipes cover supervised fine-tuning (including LoRA for Super) of released checkpoints. The report describes the internal training infrastructure used for both towers but does not identify it as Cosmos-Framework, and pre-training configurations were not found.679 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The card summarizes the corpus (about 1.3B data points from 393 dataset entries, drawn from NVIDIA-owned data and public datasets, plus synthetic data) and links a public summary of training content. The report lists five synthetic datasets released on Hugging Face; most training data is not released.19 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The technical report describes the reasoner and generator training stages (pre-training, mid-training, and post-training), data mixtures, optimizer settings, learning-rate schedules, token counts, and GPU counts.9 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The card and report publish benchmark results, and the report releases one human evaluation benchmark (Cosmos-HUE); code to re-run all reported evaluations was not verified.19 |
What it is useful for
Run and use notes
- The card lists PyTorch, vLLM-Omni, Diffusers, and SGLang as runtimes; NVIDIA Ampere, Hopper, and Blackwell GPUs; and Linux as the only tested operating system. It states that only BF16 precision is tested and gives a recommended vLLM-Omni serving setup for eight H200, H100, or A100 GPUs.1
Organization context
Provenance and derivatives
Trained by NVIDIA. The technical report states that Cosmos3-Super adapts the Qwen3-VL 32B architecture and is initialized from pre-trained Qwen3-VL-32B weights (developed by the Qwen team at Alibaba Cloud), and that visual generation uses the video VAE encoder from Wan2.2-TI2V-5B (published on Hugging Face by Wan-AI). The card lists synthetic training data generated with HiDream-I1, Qwen-Image-2512, and Qwen3-VL.911211
- Derived from: Qwen3-VL-32B (external site: github.com) — Initialization weights for the language and vision backbone.
- Derived from: Wan2.2-TI2V-5B video VAE (external site: huggingface.co) — Video VAE encoder used for visual generation; published by Wan-AI under Apache 2.0.
Other releases in the NVIDIA Cosmos family
- NVIDIA Cosmos3-NanoModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
The model card names NVIDIA as the model developer. NVIDIA's principal executive offices are in Santa Clara, California, per its Form 10-Q for the quarter ended July 26, 2026. The model is initialized from third-party Qwen3-VL weights, recorded under provenance; that base is not U.S.-developed.113912
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.