Grok-1
Release in the Grok family · version Grok-1
Maintained by xAI (SpaceXAI)12
Grok-1 is a 314-billion-parameter mixture-of-experts language model (8 experts, 2 used per token) that xAI trained from scratch. The published checkpoint is the raw base model from pretraining that ended in October 2023 and is not fine-tuned for dialogue or other applications.12
- Repository: Grok-1 repository (external site: github.com)
- Model hub: Hugging Face model repository (external site: huggingface.co)
- Release notes: Open Release of Grok-1 (external site: x.ai)
- License: LICENSE (Apache 2.0) (external site: github.com)
Availability and license
Overall availability
Downloadable from Hugging Face without gating or by torrent.42
Availability is separate from permission: read the license before using or redistributing.
The repository README states that the Apache 2.0 license applies only to the source files in the repository and the Grok-1 weights.2
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Base-model checkpoint published on Hugging Face (described there as int8) and via a torrent link.42 |
| Inference codeIs code for running the model published? | Public | JAX example code for loading and sampling; xAI says its MoE implementation is not efficient.2 |
| Training codeIs the code used to train the model published? | Unknown | Not assessed. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Unknown | The release notes say only that the base model was trained on a large amount of text data. |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Unknown | Not assessed. |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Unknown | Not assessed. |
Run and use notes
- The GitHub repository documents running the checkpoint with its JAX example code (run.py) and says a machine with enough GPU memory is required; the Hugging Face card says a multi-GPU machine is needed. The model supports activation sharding and 8-bit quantization.24
- Documented specifications: 64 layers, 48 query heads and 8 key/value heads, a 131,072-token SentencePiece tokenizer, rotary embeddings, and a maximum context of 8,192 tokens.2
Organization context
Provenance and derivatives
xAI states that Grok-1 was trained from scratch using a custom training stack built on JAX and Rust.1
Other releases in the Grok family
- Grok 2Model-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
Developed by xAI, whose AI operations are headquartered in Palo Alto, California, according to the June 2026 prospectus of its parent, Space Exploration Technologies Corp., a Texas corporation.5
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.