Gemma 4 12B Unified
Release in the Gemma family · version 4 (12B Unified)
Maintained by Google DeepMind14
Gemma 4 12B Unified is an 11.95-billion-parameter Gemma 4 model with an encoder-free design: image patches and audio waveforms are projected directly into the language model instead of passing through separate encoders. It accepts text, image, audio, and video (as frames) input and generates text, with a 256K-token context window. Google's release page dates its release to June 3, 2026, after the other Gemma 4 sizes.1498
- Model hub: Hugging Face (instruction-tuned) (external site: huggingface.co)
- Model hub: Hugging Face (pre-trained) (external site: huggingface.co)
- Documentation: Gemma 4 model card (external site: ai.google.dev)
- License: Gemma 4 license (external site: ai.google.dev)
- Paper: Gemma 4 Technical Report (external site: arxiv.org)
Availability and license
Overall availability
Weights are downloadable from Hugging Face; the repositories were not gated at the time of review. Use is governed by the Apache License 2.0.1238
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Pre-trained and instruction-tuned checkpoints are published in safetensors format on Hugging Face.123 |
| Inference codeIs code for running the model published? | Public | The model card documents inference with Hugging Face Transformers; Google DeepMind's Apache-2.0 gemma JAX library also supports Gemma 4.110 |
| Training codeIs the code used to train the model published? | Unknown | The gemma JAX library includes fine-tuning code. This catalog did not find published code used to pre-train Gemma 4.10 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card and technical report describe the data types (web documents, code, mathematics, images, and audio for this size) and a January 2025 cutoff; the data itself is not released.17 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The technical report describes the architecture, compute setup (TPUv4 and TPUv6e, JAX, Pathways), and data filtering, and says pre-training and post-training follow the Gemma 3 approach; it is not a complete recipe.7 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | Benchmark results for the instruction-tuned models are reported in the model card and technical report.17 |
What it is useful for
The model card lists text generation, chatbots, summarization, image data extraction, NLP and vision-language research, and language-learning tools among intended uses, and documents a configurable thinking mode and native function calling. It also lists audio processing, such as speech recognition and speech translation, for this size.1
Run and use notes
- The model card shows loading the model with Hugging Face Transformers (AutoModelForMultimodalLM) and recommends sampling with temperature 1.0, top_p 0.95, and top_k 64.1
Organization context
Other releases in the Gemma family
- Gemma 4 26B A4BModel-disclosure tier (USASI rubric v0.1): Open-weight
- Gemma 4 31BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: Documented U.S. control
The model card names Google DeepMind as the author. Google DeepMind is a research unit of Google, announced by Google's CEO in 2023, and Google's parent Alphabet Inc. has its principal executive offices in Mountain View, California, per its fiscal 2025 Form 10-K.11112
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.