USASI
Evaluation toolEvaluation tool

HELM (Holistic Evaluation of Language Models)

Version 0.5.16

Maintained by Stanford Center for Research on Foundation Models (CRFM)1310

HELM is an open-source Python framework from Stanford's Center for Research on Foundation Models for evaluating foundation models, including language and multimodal models. It provides benchmarks in a standardized format, one interface to models from several providers, metrics beyond accuracy such as efficiency, bias, and toxicity, and a web UI and leaderboards for inspecting results. HELM entered maintenance mode on June 1, 2026.14

Last reviewedEntry updated Documented release Apr 30, 2026

Availability and license

Overall availability

Public

Source code on GitHub under the Apache License 2.0; installable from PyPI as crfm-helm. Raw leaderboard results are in a public Google Cloud Storage bucket that needs no authentication.1237

Availability is separate from permission: read the license before using or redistributing.

Benchmark data is not bundled with the Apache-licensed code. HELM scenarios download data from its original sources, such as URLs or Hugging Face datasets, and cache it locally.8

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for HELM (Holistic Evaluation of Language Models)
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicPublished on GitHub under the Apache License 2.0.12
Tasks / dataAre the tasks or test data available?PartialScenario code fetches each benchmark's data from its original source. Raw results for the HELM leaderboards (Capabilities, Safety, MedHELM, VHELM, and others) can be downloaded from a public storage bucket. The reproduction guide divides MedHELM benchmarks into public, gated (credentials or approval required), and private ones.876
MethodologyIs the method for scoring described?PublicThe HELM paper (TMLR 2023) describes the taxonomy of scenarios and metrics, the seven metrics measured, and the standardized evaluation conditions.9
ReproducibilityAre instructions for reproducing results published?PublicThe documentation covers installation and a step-by-step procedure for reproducing leaderboards from the run-entry and schema files in the repository.56
LimitationsAre known limitations documented?PartialThe HELM paper names areas the benchmark does not yet cover. The maintenance-mode policy sets out what is no longer supported and warns that scenarios and models that depend on external APIs may break without active support.94

What it is useful for

Running standardized benchmark evaluations of foundation models, reproducing HELM leaderboard results from the published run configurations, and inspecting individual prompts and model responses.16

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The installation guide requires Python 3.10 or later and recommends installing into a virtualenv or conda environment. Text-to-image (HEIM) and vision-language (VHELM) evaluations need extra dependencies.5
  • Under the maintenance-mode policy, HELM is maintained by volunteers on a best-effort basis, and significant issues may be addressed when maintainer bandwidth is available. No new features or leaderboard evaluations will be added. Users for whom the lack of active support is a problem are pointed to alternative frameworks: Evalchemy, Inspect AI Evals, Lighteval, EleutherAI's LM Evaluation Harness, and Unitxt.4

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The README states that HELM was created by the Center for Research on Foundation Models (CRFM) at Stanford, and the repository sits in the stanford-crfm GitHub organization. The PyPI package lists Stanford CRFM as its author, and CRFM describes itself as an initiative at Stanford HAI. The IRS extract lists Stanford's trustees in Stanford, California, as a 501(c)(3) organization.131011

Assessed Sep 29, 2026

Sources

  1. 1.
    HELM README (external site: raw.githubusercontent.com)

    Stanford CRFM · Repository · accessed Sep 29, 2026

  2. 2.
    HELM LICENSE (external site: raw.githubusercontent.com)

    Stanford CRFM · License · accessed Sep 29, 2026

  3. 3.
    crfm-helm on PyPI (external site: pypi.org)

    Python Package Index · Documentation · published Apr 30, 2026 · accessed Sep 29, 2026

  4. 4.
  5. 5.
    Installation - HELM documentation (external site: crfm-helm.readthedocs.io)

    Stanford CRFM · Documentation · accessed Sep 29, 2026

  6. 6.
  7. 7.
  8. 8.
  9. 9.
    Holistic Evaluation of Language Models (arXiv:2211.09110) (external site: arxiv.org)

    arXiv · Paper · published Nov 16, 2022 · accessed Sep 29, 2026

  10. 10.
  11. 11.

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project