gpt-oss-20b
Release in the gpt-oss family · version gpt-oss-20b
gpt-oss-20b is the smaller gpt-oss model: a 24-layer mixture-of-experts transformer with 20.9B total and about 3.6B active parameters per token, using 32 experts with the top 4 selected per token. It is text-only, supports context up to 131,072 tokens, and ships with its MoE weights quantized to MXFP4.1
- Model hub: Hugging Face model card (external site: huggingface.co)
- Repository: gpt-oss repository (external site: github.com)
- Paper: Model card (arXiv) (external site: arxiv.org)
- License: LICENSE (Apache 2.0) (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.2
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | MXFP4-quantized MoE weights plus an original-format checkpoint are in the Hugging Face repository.2 |
| Inference codeIs code for running the model published? | Public | The GitHub repository has reference PyTorch, Triton, and Metal implementations, tools, and a Responses API server.5 |
| Training codeIs the code used to train the model published? | Unknown | The official repository covers inference, tools, and evaluation; no training code was found. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card describes a text-only dataset of trillions of tokens focused on STEM, coding, and general knowledge, filtered for hazardous biosecurity content, with a June 2024 knowledge cutoff. The data itself is not released.1 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The model card documents architecture, tokenizer, training compute (almost 10x fewer H100-hours than the 2.1 million used for gpt-oss-120b), and post-training with chain-of-thought reinforcement learning at a high level, not in reproducible detail.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The model card reports benchmark and safety results; the repository includes evaluation code adapted from simple-evals for running GPQA and HealthBench.17 |
What it is useful for
Presented for lower-latency, local, or specialized use cases, including tool use and function calling; the model card says it can be fine-tuned on consumer hardware.2
Run and use notes
- The model card documents running the model with Transformers, vLLM, Ollama, and LM Studio, and OpenAI's repository provides reference PyTorch, Triton, and Metal implementations.25
- OpenAI states that with its MoE weights in MXFP4 precision (4.25 bits per parameter), gpt-oss-20b can run on systems with as little as 16GB of memory; the model card lists the checkpoint size as 12.8 GiB.12
- The model must be prompted with OpenAI's harmony response format; the model card says it will not work correctly otherwise.2
Organization context
Other releases in the gpt-oss family
- gpt-oss-120bModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.