Whisper large-v3-turbo
Release in the Whisper family · version large-v3-turbo
Whisper large-v3-turbo is a multilingual speech recognition model that OpenAI derived from Whisper large-v3 by cutting the decoder from 32 layers to 4 and fine-tuning for two more epochs. It has about 0.8 billion parameters (809M in the README and Hugging Face card; the repository model card lists 798M) and is the default model in the openai-whisper package.4615
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: openai/whisper repository (external site: github.com)
- Release notes: large-v3-turbo release discussion (external site: github.com)
- License: LICENSE (MIT) (external site: github.com)
Availability and license
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published on Hugging Face without gating and downloadable through the openai-whisper package.28 |
| Inference codeIs code for running the model published? | Public | The openai/whisper repository provides the inference code and command-line tool; the Hugging Face card documents inference with Transformers.61 |
| Training codeIs the code used to train the model published? | Unknown | The repository covers inference and evaluation-data preparation; this review found no published training code.6 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | OpenAI says turbo was fine-tuned on the same amount of multilingual transcription data used to train large-v3, excluding translation data. The large-v3 training mixture (1 million hours of weakly labeled audio and 4 million hours pseudo-labeled with large-v2) is described but not released.43 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The release discussion describes the procedure at a high level (decoder reduced to 4 layers, two further epochs of fine-tuning); no training configuration is published.4 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The release discussion reports comparisons with other Whisper models. The repository's data README documents how the paper's evaluation datasets were prepared, but no turbo-specific evaluation scripts were located.49 |
What it is useful for
The README presents turbo as a faster version of large-v3 for transcription with a small loss of accuracy. It is not trained for translation: it returns the source language even when the translate task is requested, and the README directs users to the other multilingual models, such as medium or large, for translation into English.64
Run and use notes
Organization context
Provenance and derivatives
Derived by OpenAI from its own Whisper large-v3: the decoder was pruned from 32 to 4 layers and the model was fine-tuned for two more epochs on large-v3's multilingual transcription data.41
- Derived from: Whisper large-v3 — Base model; same developer.
Other releases in the Whisper family
- Whisper large-v3Model-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.