SmolLM3 3B
Release in the SmolLM family · version SmolLM3-3B
Maintained by Hugging Face (Smol Models Research, HuggingFaceTB)110
SmolLM3-3B is the instruct model of Hugging Face's SmolLM3 release (July 2025), a 3B-parameter decoder-only transformer pretrained on about 11 trillion tokens of web, code, math, and reasoning data in stages, then mid-trained on reasoning data and aligned with supervised fine-tuning and Anchored Preference Optimization. It can answer with or without an extended reasoning trace, supports tool calling, and natively covers six European languages.12
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Release notes: SmolLM3 blog post (external site: huggingface.co)
- Repository: Pretraining configs (huggingface/smollm) (external site: github.com)
- Repository: Post-training recipe (alignment-handbook) (external site: github.com)
- Model hub: Intermediate checkpoints (external site: huggingface.co)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.1
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (huggingface/smollm) (external site: raw.githubusercontent.com)5
Apache License 2.0 (alignment-handbook) (external site: raw.githubusercontent.com)7
The post-training dataset SmolTalk2 states that its newly created subsets are Apache 2.0 and that the existing public datasets it includes keep their original licenses.9
Model-disclosure tier
Every item in the model checklist is public, including the training data itself, and the weights and code are under OSI-approved licenses.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Final weights are on Hugging Face without gating; intermediate checkpoints, including mid-training and SFT checkpoints, are published separately.1 |
| Inference codeIs code for running the model published? | Public | The card documents inference with Transformers v4.53.0 or later and with vLLM and SGLang, and lists quantized options for llama.cpp, ONNX, MLX, MLC, and ExecuTorch.1 |
| Training codeIs the code used to train the model published? | Public | Nanotron pretraining configs for the three stages and long-context extension are in the huggingface/smollm repository; the mid-training, SFT, and APO configs are in the alignment-handbook SmolLM3 recipe.146 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Public | The pretraining datasets are listed in a public Hugging Face collection, and the mid-training, SFT, and preference data are published as SmolTalk2. Several listed datasets are published by third parties (for example DCLM and The Stack v2) under their own terms.189 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The blog post gives the web, code, and math proportions for each pretraining stage, the long-context extension, and the post-training stages; the configs include exact data weights.2 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The card reports base and instruct evaluation results run with Hugging Face's lighteval and states that evaluation configs and code are in the huggingface/smollm repository.1 |
What it is useful for
Run and use notes
- The card states the shipped config supports 65,536-token context and describes enabling YaRN rope scaling (factor 2.0) for inputs up to about 128k tokens; it recommends temperature 0.6 and top_p 0.95.1
Organization context
Provenance and derivatives
Post-trained by Hugging Face from its own SmolLM3-3B-Base, which was pretrained from scratch on public datasets. The mid-training data (NVIDIA's Llama-Nemotron post-training dataset and OpenThoughts3) contains reasoning traces, and several SmolTalk2 SFT subsets were generated with Qwen3-32B or DeepSeek-V3, so the post-training data includes outputs of third-party models.1289
- Derived from: SmolLM3-3B-Base (external site: huggingface.co) — Hugging Face base model; no separate catalog record.
- Derived from: FineWeb-Edu (subset of FineWeb) — One of the web pretraining datasets.
- Derived from: SmolTalk2 (external site: huggingface.co) — Mid-training, SFT, and preference data.
Other releases in the SmolLM family
No other releases in this family have been assessed.
U.S. eligibility
Eligible · basis: U.S. headquarters
Developed and published by Hugging Face's HuggingFaceTB organization ("Hugging Face Smol Models Research"). Hugging Face's terms of service identify Hugging Face, Inc., a Delaware corporation, as the provider of its services, and its privacy policy states that the company is located in the United States. See the hugging-face organization record for the dual-country assessment.1101112
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.