SGLang
Version 0.5.20
Maintained by LMSYS (hosting organization)27, RadixArk (company described as a maintainer in a post on the LMSYS blog)910
SGLang is an open-source serving framework for large language models and multimodal models, designed for low-latency, high-throughput inference on anything from a single GPU to large distributed clusters. Its runtime features include RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, continuous batching, several forms of parallelism, structured outputs, and quantization.2
- Repository: GitHub repository (external site: github.com)
- Documentation: Installation guide (external site: docs.sglang.io)
- Release notes: Release v0.5.20 (external site: github.com)
- License: License (Apache 2.0) (external site: raw.githubusercontent.com)
Availability and license
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| Source codeIs the source code publicly readable? | Public | Published on GitHub under the Apache License 2.0.123 |
| DocumentationIs user documentation published? | Public | The documentation includes installation, quick-start, backend and frontend tutorials, a per-model deployment cookbook, and a contribution guide.24 |
| InstallationAre installation instructions or packages publicly available? | Public | Documented methods include pip or uv, installing from source, Docker images, Kubernetes, and cloud deployment options.4 |
| Supported platformsAre supported operating systems or hardware documented? | Public | The install guide mainly covers NVIDIA GPUs and has separate pages for AMD GPUs, Apple Metal, Intel Xeon CPUs, Intel XPU, Google TPU, NVIDIA Jetson and DGX Spark, and Ascend NPUs.42 |
| Release statusAre versioned releases published? | Public | Versioned releases are published on GitHub and PyPI; v0.5.20 was released on 2026-09-18. Nightly wheels are also available.564 |
What it is useful for
Serving language, embedding, reward, and diffusion models behind an OpenAI-compatible API, and acting as the rollout backend in reinforcement-learning post-training frameworks.2
Run and use notes
- The installation guide requires Python 3.10 or higher. It states that SGLang requires CUDA 13, that the CUDA 12 wheels and images are retired, and that 0.5.19 was the last release with CUDA 12 builds.4
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The SGLang README states that the project is currently hosted under the non-profit open-source organization LMSYS. LMSYS describes itself as a 501(c)(3) nonprofit incorporated in September 2024. The IRS extract for California lists LMSYS Corp, in Sacramento, as a 501(c)(3) corporation with a September 2024 ruling date. A July 2026 post on the LMSYS blog, credited to RadixArk and Google, also describes the company RadixArk as a maintainer; this assessment rests on LMSYS hosting, and RadixArk's location was not verified. SGLang joined the PyTorch Ecosystem in 2025, but the pages read do not describe it as a PyTorch Foundation-hosted project.278911
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.