USASI
Model release

Phi-4-reasoning-vision-15B

Release in the Phi family · version 4-reasoning-vision-15B

Maintained by Microsoft1

Phi-4-reasoning-vision-15B is a 15-billion-parameter multimodal reasoning model from Microsoft. It combines the Phi-4-Reasoning language model with a SigLIP-2 vision encoder, takes text and images as input, produces text, and has a 16,384-token context length. It can either reason step by step or answer directly, depending on the task.1

Last reviewedEntry updated Documented release Mar 4, 2026

Availability and license

Overall availability

Public

Weights can be downloaded from Hugging Face without a gated access request. Microsoft also offers the model as a hosted deployment on Microsoft Foundry.21

Availability is separate from permission: read the license before using or redistributing.

The Hugging Face model card lists the MIT License, but that repository has no separate LICENSE file. The GitHub repository has a standard MIT LICENSE file (Copyright 2026 Microsoft). The model card says nothing in it restricts or modifies the license.124

Model-disclosure tier

Computed from the checklist below using USASI rubric v0.1. An editorial category, not a certification.
Model-disclosure tier (USASI rubric v0.1): Open-weight

The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.

How tiers are computed

Public materials checklist

Items for a model under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for Phi-4-reasoning-vision-15B
ItemStatusNotes and evidence
WeightsCan the general public download the model parameters for this release?PublicPublished on Hugging Face without gating.21
Inference codeIs code for running the model published?PublicThe GitHub repository has Transformers and vLLM directories, and the model card documents the required torch, transformers, and vllm versions.41
Training codeIs the code used to train the model published?UnknownThe Microsoft Research blog says fine-tuning code was released with the model, but the GitHub repository contains only Transformers and vLLM inference code, and this review did not locate the fine-tuning code elsewhere.54
Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access.PartialThe model card and blog describe about 200 billion tokens of multimodal data. Most of it comes from filtered open-source vision-language datasets, with added internal Microsoft domain data and targeted acquisitions. The blog says some responses were regenerated with GPT-4o and o4-mini. An EU-format data card is published; the training data itself is not.153
Training recipeAre the training configuration and procedure documented in enough detail to follow?PartialThe model card describes supervised fine-tuning on a mix of reasoning and non-reasoning data, the mid-fusion architecture, and compute (240 NVIDIA B200 GPUs for 4 days). A technical report is linked but was not reviewed here.1
Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only.PartialThe model card reports results and names the open-source evaluation frameworks used (Eureka ML Insights and VLMEvalKit). Microsoft says benchmark logs were released; this review did not locate them.15

What it is useful for

The model card lists scientific and mathematical reasoning over visual inputs such as diagrams, charts, and documents, and computer-use agent tasks such as locating interface elements on screens. It also covers captioning, visual question answering, and OCR, and is trained primarily on English.1

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The model card lists torch 2.7.1 or later and transformers 4.57.1 or later (vllm 0.15.2 or later for vLLM). It says the model was tested on NVIDIA A6000, A100, H100, and B200 GPUs under Ubuntu 22.04.5 and recommends serving it with vLLM in bf16 precision.1

Organization context

Provenance and derivatives

Built on Microsoft's Phi-4-Reasoning language model and a SigLIP-2 vision encoder, according to the model card. The Microsoft Research blog says some training responses were regenerated with GPT-4o and o4-mini.15

Other releases in the Phi family

Phi family overview

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S. headquarters

The model card names Microsoft Corporation as developer, with Microsoft Ireland Operations Limited as authorized representative for the EU. The EU-format data card lists the Irish entity in its developer field, but the release is a Microsoft product published on Microsoft's Hugging Face and GitHub accounts. Microsoft Corporation lists Redmond, Washington as its address on its Form 10-K cover page.136

Assessed Sep 29, 2026

Sources

  1. 1.
    microsoft/Phi-4-reasoning-vision-15B model card (external site: huggingface.co)

    Microsoft (Hugging Face) · Model card · published Mar 4, 2026 · accessed Sep 29, 2026

  2. 2.
  3. 3.
    Phi-4-Reasoning-Vision-15B data card (DATACARD.md) (external site: huggingface.co)

    Microsoft (Hugging Face) · Model card · published Aug 31, 2026 · accessed Sep 29, 2026

  4. 4.
  5. 5.
    Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model (external site: microsoft.com)

    Microsoft Research · Announcement · published Mar 4, 2026 · accessed Sep 29, 2026

  6. 6.
    Microsoft Corporation Form 10-K for the fiscal year ended June 30, 2026 (external site: sec.gov)

    Microsoft Corporation (U.S. SEC filing) · Filing · published Jul 29, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project