Phi-4-mini-flash-reasoning
Release in the Phi family · version 4-mini-flash-reasoning
Phi-4-mini-flash-reasoning is a 3.8-billion-parameter, English, text-only Microsoft model fine-tuned for mathematical reasoning. It uses a hybrid "SambaY" decoder-hybrid-decoder architecture that mixes state space model layers with attention and Differential Attention, and supports a 64K-token context.1
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- License: License (MIT) (external site: huggingface.co)
- Repository: ArchScale training codebase (external site: github.com)
- Paper: Paper (arXiv 2507.06607) (external site: arxiv.org)
Availability and license
Overall availability
Weights can be downloaded from Hugging Face without a gated access request. Microsoft also offers the model on Azure AI Foundry. Microsoft announced on July 9, 2025 that the model was available that day on Azure AI Foundry, the NVIDIA API Catalog, and Hugging Face; the model card gives June 2025 as its release date.412
Availability is separate from permission: read the license before using or redistributing.
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Published on Hugging Face without gating.4 |
| Inference codeIs code for running the model published? | Public | The model card documents Transformers inference with pinned package versions and links vLLM pull requests that add support.1 |
| Training codeIs the code used to train the model published? | Partial | Microsoft's ArchScale repository (MIT) says it released code for large-scale pre-training of Phi-4-mini-flash. The code for the reasoning fine-tuning stage is not identified.51 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card lists 5T pre-training tokens and 150B reasoning-training tokens, with a February 2025 cutoff for public data. It says the reasoning data is synthetic math content generated by DeepSeek-R1 plus curated public math questions. The data is not released.1 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | The model card gives architecture, hardware, duration, and token counts for each stage, and includes the abstract of the accompanying architecture paper, which this review did not read in full. The card does not publish a complete fine-tuning recipe.1 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The model card reports results and the sampling protocol (Pass@1 averaged over 64 samples for AIME24/25 and 8 for Math500 and GPQA Diamond). ArchScale includes LightEval-based reasoning evaluation for AIME, MATH-500, and GPQA.15 |
What it is useful for
The model card lists multi-step mathematical problem solving where memory, compute, or latency is constrained, such as formal proof generation and advanced word problems. It says the model is designed and tested for math reasoning only.1
Run and use notes
- The model card lists pinned packages for Transformers inference (flash_attn 2.7.4.post1, torch 2.6.0, mamba-ssm 2.2.4, causal-conv1d 1.5.0.post8, transformers 4.46.1). It says the model uses flash attention by default and has been tested on NVIDIA A100 and H100 GPUs.1
Organization context
Provenance and derivatives
Microsoft trained the model. According to the model card, its reasoning fine-tuning data is synthetic math content generated by DeepSeek-R1 in order to distill that model's reasoning, along with curated public math questions and part of the SFT data used for the base Phi-4-mini-flash model. No DeepSeek weights are part of this release.1
- Derived from: DeepSeek-R1 — Teacher model whose outputs were used as synthetic fine-tuning data; not a base-weight dependency.
- Derived from: Phi-4-mini-flash (base) — Microsoft base model whose SFT data is partly reused, per the model card.
Other releases in the Phi family
- Phi-4-reasoning-vision-15BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S. headquarters
Published by Microsoft on its Hugging Face account under a Microsoft Corporation copyright license. Microsoft Corporation lists Redmond, Washington as its address on its Form 10-K cover page. The reasoning fine-tuning data was generated by another developer's model (see provenance), but the weights were trained by Microsoft.136
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.