USASI
Evaluation toolEvaluation tool

OpenAI Evals

Project record

Maintained by OpenAI13

OpenAI Evals is an open-source framework from OpenAI for evaluating language models and systems built on them, together with a registry of existing evals. Evals are defined in YAML and run from the oaieval and oaievalset command-line tools against "completion functions", which by default call models through the OpenAI API. The repository is not archived, but activity since mid-2024 has been limited to maintenance, and its README now points users to OpenAI's hosted Evals in the OpenAI dashboard.2467

Last reviewedEntry updated Documented release Mar 2023

Availability and license

Overall availability

Public

Source code on GitHub under the MIT License and installable from PyPI as evals. The eval registry data is stored with Git LFS and must be fetched separately. Running evals against OpenAI models requires an OpenAI API key and incurs API costs.237

Availability is separate from permission: read the license before using or redistributing.

The license file states that MIT applies to everything except the datasets it lists separately, which keep their own licenses (for example ODC-By, CC0, CC BY, CC BY-SA, and, for some components of the steganography and theory-of-mind evals, CC BY-NC 4.0). The README states that contributors agree to license their eval logic and data under MIT and that OpenAI may use contributed data in future service improvements.32

Public materials checklist

Items for a evaluation tool under USASI rubric v0.1. Unknown means unassessed or insufficient evidence.
Public materials checklist for OpenAI Evals
ItemStatusNotes and evidence
CodeIs the evaluation code published?PublicPublished on GitHub under the MIT License.13
Tasks / dataAre the tasks or test data available?PublicEval definitions are YAML files under evals/registry/evals; the data files under evals/registry/data are stored with Git LFS. Some datasets carry their own licenses.243
MethodologyIs the method for scoring described?Publicdocs/eval-templates.md specifies the matching rules of the basic templates and the parameters of the model-graded classification template.5
ReproducibilityAre instructions for reproducing results published?Publicdocs/run-evals.md documents the oaieval and oaievalset commands, threading and timeout settings, and local JSONL logging.4
LimitationsAre known limitations documented?PartialThe documentation notes operational limits (a single eval cannot be resumed mid-run; runs sometimes hang after the final report), and the README says evals with custom code are not being accepted. A general discussion of eval validity was not found in the pages read.42

What it is useful for

Writing evals from templates (exact, includes, fuzzy, and JSON match, or model-graded classification) without custom code, running them against OpenAI API models or other completion functions, and logging results locally.542

Run and use notes

Documented facts only. No hardware or performance claims are made without a cited source and stated assumptions.
  • The README gives Python 3.9 as the minimum version. The latest release on PyPI is 3.0.1.post1, published May 1, 2024.27
  • Recent commits are maintenance: a December 2024 README change linking to the hosted Evals product, removal of an eval suite with defunct dependencies in November 2025, and pinning of CI and pre-commit references in April 2026.6

Organization context

U.S. eligibility

Project eligibility rests on documented governing or maintaining entities, not on contributors.

Eligible · basis: U.S.-governed project

The project is published in OpenAI's GitHub organization, and its MIT license names OpenAI as copyright holder. OpenAI Group PBC lists its address as 1455 3rd Street, San Francisco, California, in a February 2026 agreement filed with the SEC.138

Assessed Sep 29, 2026

Sources

  1. 1.
    openai/evals (external site: github.com)

    OpenAI · Repository · accessed Sep 29, 2026

  2. 2.
    openai/evals README.md (external site: raw.githubusercontent.com)

    OpenAI · Documentation · accessed Sep 29, 2026

  3. 3.
  4. 4.
  5. 5.
  6. 6.
    Commits · openai/evals (external site: github.com)

    OpenAI · Repository · accessed Sep 29, 2026

  7. 7.
    evals on PyPI (JSON metadata) (external site: pypi.org)

    Python Package Index · Documentation · published May 1, 2024 · accessed Sep 29, 2026

  8. 8.
    Exhibit 10.1: Equity commitment letter agreement between OpenAI Group PBC and Amazon (external site: sec.gov)

    U.S. Securities and Exchange Commission (Amazon.com, Inc. filing) · Filing · published Feb 27, 2026 · accessed Sep 29, 2026

This listing is not an endorsement, a safety assessment, or a federal approval.

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project