Inkling-Small
Release in the Inkling family · version Inkling-Small
Maintained by Thinking Machines Lab13
Inkling-Small is a 42-layer decoder-only mixture-of-experts transformer with 276B total and 12B active parameters, routing each token to 6 of 256 experts plus 2 shared experts. Like Inkling, it accepts text, image, and audio input, outputs text, and supports a context window of up to 1M tokens.12
- Model hub: Hugging Face (BF16) (external site: huggingface.co)
- Model hub: Hugging Face (NVFP4) (external site: huggingface.co)
- Documentation: Inkling-Small Model Card (external site: thinkingmachines.ai)
- Release notes: Introducing Inkling-Small (external site: thinkingmachines.ai)
- License: Model Acceptable Use Policy (external site: thinkingmachines.ai)
Availability and license
Overall availability
BF16 and NVFP4 checkpoints are downloadable from Hugging Face without gating; the developer's Model Acceptable Use Policy also applies.34
Availability is separate from permission: read the license before using or redistributing.
The Hugging Face repository declares Apache 2.0 in its metadata and links to the Apache license text; it has no separate LICENSE file. The developer's Model Acceptable Use Policy states that anyone who accesses, downloads, or uses its model materials agrees to be bound by it, and it lists prohibited uses.34
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published as a BF16 checkpoint and a separate NVFP4-quantized checkpoint.31 |
| Inference codeIs code for running the model published? | Public | The model card documents deployment with open-source SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face Transformers.13 |
| Training codeIs the code used to train the model published? | Unknown | Not assessed. |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card and training-data documentation describe source categories (public, third-party, and synthetic data), deduplication, and filtering; the release post says the pretraining data mix differs from Inkling's. The data is not released.152 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The release post describes, at a high level, changes to the pre-training data mix and recipe, post-training that used on-policy distillation from Inkling in part, and further agentic-coding reinforcement learning.2 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | Benchmark results are published in the model card and as evaluation result files in the Hugging Face repository.3 |
What it is useful for
The model card describes use by developers building agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems.1
Run and use notes
- The model card states that the BF16 checkpoint needs at least 600 GB of aggregated GPU memory (for example 4x NVIDIA B300 or 8x H200), and that the NVFP4-quantized checkpoint needs at least 180 GB (W4A4 on 1x B300, which requires SM100+ architecture, or W4A16 on 2x H200).1
Organization context
Provenance and derivatives
Thinking Machines Lab says Inkling-Small was pretrained with a changed data mix and recipe, and that an earlier checkpoint was post-trained in part using on-policy distillation with Inkling as the teacher, followed by further agentic-coding reinforcement learning.2
- Derived from: Inkling — Teacher model for on-policy distillation during post-training; not a weight initialization.
Other releases in the Inkling family
- InklingModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.