HELM (Holistic Evaluation of Language Models)
Version 0.5.16
Maintained by Stanford Center for Research on Foundation Models (CRFM)1310
HELM is an open-source Python framework from Stanford's Center for Research on Foundation Models for evaluating foundation models, including language and multimodal models. It provides benchmarks in a standardized format, one interface to models from several providers, metrics beyond accuracy such as efficiency, bias, and toxicity, and a web UI and leaderboards for inspecting results. HELM entered maintenance mode on June 1, 2026.14
- Repository: GitHub repository (external site: github.com)
- Documentation: Documentation (external site: crfm-helm.readthedocs.io)
- Documentation: Maintenance mode policy (external site: crfm-helm.readthedocs.io)
- Paper: HELM paper (arXiv) (external site: arxiv.org)
- License: License (Apache 2.0) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Source code on GitHub under the Apache License 2.0; installable from PyPI as crfm-helm. Raw leaderboard results are in a public Google Cloud Storage bucket that needs no authentication.1237
Availability is separate from permission: read the license before using or redistributing.
Benchmark data is not bundled with the Apache-licensed code. HELM scenarios download data from its original sources, such as URLs or Hugging Face datasets, and cache it locally.8
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Published on GitHub under the Apache License 2.0.12 |
| Tasks / dataAre the tasks or test data available? | Partial | Scenario code fetches each benchmark's data from its original source. Raw results for the HELM leaderboards (Capabilities, Safety, MedHELM, VHELM, and others) can be downloaded from a public storage bucket. The reproduction guide divides MedHELM benchmarks into public, gated (credentials or approval required), and private ones.876 |
| MethodologyIs the method for scoring described? | Public | The HELM paper (TMLR 2023) describes the taxonomy of scenarios and metrics, the seven metrics measured, and the standardized evaluation conditions.9 |
| ReproducibilityAre instructions for reproducing results published? | Public | The documentation covers installation and a step-by-step procedure for reproducing leaderboards from the run-entry and schema files in the repository.56 |
| LimitationsAre known limitations documented? | Partial | The HELM paper names areas the benchmark does not yet cover. The maintenance-mode policy sets out what is no longer supported and warns that scenarios and models that depend on external APIs may break without active support.94 |
What it is useful for
Run and use notes
- The installation guide requires Python 3.10 or later and recommends installing into a virtualenv or conda environment. Text-to-image (HEIM) and vision-language (VHELM) evaluations need extra dependencies.5
- Under the maintenance-mode policy, HELM is maintained by volunteers on a best-effort basis, and significant issues may be addressed when maintainer bandwidth is available. No new features or leaderboard evaluations will be added. Users for whom the lack of active support is a problem are pointed to alternative frameworks: Evalchemy, Inspect AI Evals, Lighteval, EleutherAI's LM Evaluation Harness, and Unitxt.4
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The README states that HELM was created by the Center for Research on Foundation Models (CRFM) at Stanford, and the repository sits in the stanford-crfm GitHub organization. The PyPI package lists Stanford CRFM as its author, and CRFM describes itself as an initiative at Stanford HAI. The IRS extract lists Stanford's trustees in Stanford, California, as a 501(c)(3) organization.131011
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.