OpenAI Evals
Project record
OpenAI Evals is an open-source framework from OpenAI for evaluating language models and systems built on them, together with a registry of existing evals. Evals are defined in YAML and run from the oaieval and oaievalset command-line tools against "completion functions", which by default call models through the OpenAI API. The repository is not archived, but activity since mid-2024 has been limited to maintenance, and its README now points users to OpenAI's hosted Evals in the OpenAI dashboard.2467
- Repository: GitHub repository (external site: github.com)
- Documentation: How to run evals (docs/run-evals.md) (external site: github.com)
- Documentation: Eval templates (docs/eval-templates.md) (external site: github.com)
- License: License (MIT, with per-dataset licenses) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Source code on GitHub under the MIT License and installable from PyPI as evals. The eval registry data is stored with Git LFS and must be fetched separately. Running evals against OpenAI models requires an OpenAI API key and incurs API costs.237
Availability is separate from permission: read the license before using or redistributing.
The license file states that MIT applies to everything except the datasets it lists separately, which keep their own licenses (for example ODC-By, CC0, CC BY, CC BY-SA, and, for some components of the steganography and theory-of-mind evals, CC BY-NC 4.0). The README states that contributors agree to license their eval logic and data under MIT and that OpenAI may use contributed data in future service improvements.32
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Published on GitHub under the MIT License.13 |
| Tasks / dataAre the tasks or test data available? | Public | Eval definitions are YAML files under evals/registry/evals; the data files under evals/registry/data are stored with Git LFS. Some datasets carry their own licenses.243 |
| MethodologyIs the method for scoring described? | Public | docs/eval-templates.md specifies the matching rules of the basic templates and the parameters of the model-graded classification template.5 |
| ReproducibilityAre instructions for reproducing results published? | Public | docs/run-evals.md documents the oaieval and oaievalset commands, threading and timeout settings, and local JSONL logging.4 |
| LimitationsAre known limitations documented? | Partial | The documentation notes operational limits (a single eval cannot be resumed mid-run; runs sometimes hang after the final report), and the README says evals with custom code are not being accepted. A general discussion of eval validity was not found in the pages read.42 |
What it is useful for
Run and use notes
- The README gives Python 3.9 as the minimum version. The latest release on PyPI is 3.0.1.post1, published May 1, 2024.27
- Recent commits are maintenance: a December 2024 README change linking to the hosted Evals product, removal of an eval suite with defunct dependencies in November 2025, and pinning of CI and pre-commit references in April 2026.6
Organization context
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.