Humanity's Last Exam (HLE)
Version HLE (2,500 questions, finalized April 2025); HLE-Rolling; HLE-Diamond (September 2026)
Maintained by Center for AI Safety1239, Scale AI19
A multimodal benchmark of 2,500 expert-written, closed-ended academic questions across more than a hundred subjects, including mathematics, the humanities, and the natural sciences, in multiple-choice and short-answer formats suited to automated grading. It was organized by teams at the Center for AI Safety and Scale AI with questions from nearly 1,000 subject-matter contributors, and was published in Nature in January 2026.12
- Website: Website (external site: lastexam.ai)
- Repository: GitHub repository (external site: github.com)
- Dataset hub: Dataset (Hugging Face, cais/hle) (external site: huggingface.co)
- Paper: Paper (arXiv 2501.14249) (external site: arxiv.org)
- Paper: Nature paper (external site: nature.com)
- Release notes: HLE-Diamond release notes (external site: lastexam.ai)
Availability and license
Overall availability
The public question set is on Hugging Face behind an automatic access gate that asks users to share contact information; the evaluation code is public on GitHub. A private held-out question set is not released.4521
Availability is separate from permission: read the license before using or redistributing.
The dataset and repository include a canary string intended to help model developers keep the benchmark out of training data. The separate cais/hle-diamond dataset's metadata also lists the MIT license and uses the same automatic access gate.246
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | The hle_eval scripts generate model predictions through the openai-python interface and grade them with a judge model; MIT license.23 |
| Tasks / dataAre the tasks or test data available? | Partial | The 2,500 public questions are downloadable after accepting the Hugging Face access prompt; a private held-out set is kept back to assess overfitting.451 |
| MethodologyIs the method for scoring described? | Public | The paper (arXiv, with a Nature version) describes the benchmark, whose questions each have a known, unambiguous, verifiable answer that cannot be quickly found by internet search. The website explains that calibration error is measured by asking models for an answer and a 0-100% confidence, with answers graded by a judge model.81 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README gives commands for running predictions and the judge, notes temperature 0 as the default, and advises at least 8,192 completion tokens for reasoning models.2 |
| LimitationsAre known limitations documented? | Public | The website states that HLE tests structured academic problems rather than open-ended research or creative problem solving, and that high accuracy alone would not indicate autonomous research ability. Changes to the question set (removals, re-additions, updates, and additions) are listed in the public HLE-Rolling change log, and HLE-Diamond is described as a subset refined over a year of cleaning.1109 |
What it is useful for
Run and use notes
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The benchmark's organizing team is drawn from the Center for AI Safety and Scale AI; its GitHub repository is in the centerforaisafety organization under an MIT license held by centerforaisafety, and its dataset is published under CAIS's Hugging Face organization. The Center for AI Safety is a San Francisco 501(c)(3) nonprofit listed by the IRS. Scale AI is covered by its own catalog record.12371112
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.