LM Evaluation Harness
Version 0.4.13
Maintained by EleutherAI125
An open-source framework from EleutherAI for evaluating language models on many benchmark tasks through one interface. It supports local models (for example through Hugging Face Transformers or vLLM) and hosted model APIs, and its README lists more than 60 standard academic benchmarks with hundreds of subtasks and variants.1
- Repository: GitHub repository (external site: github.com)
- Release notes: Release v0.4.13 (external site: github.com)
- Documentation: New task guide (external site: github.com)
- License: License (MIT) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Source code on GitHub under the MIT License; installable from PyPI as lm-eval, with model backends installed as optional extras.152
Availability is separate from permission: read the license before using or redistributing.
The new-task guide states that task data is downloaded and managed through the Hugging Face datasets API, so benchmark data comes from separately published datasets rather than from the harness repository.3
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Published on GitHub under the MIT License.12 |
| Tasks / dataAre the tasks or test data available? | Public | Task configurations are in the repository; task data is loaded from datasets on the Hugging Face Hub (or local files) as specified in each task's configuration.31 |
| MethodologyIs the method for scoring described? | Public | The new-task guide describes generative and multiple-choice (log-likelihood) task types and how metrics and aggregations are declared; the README sets out the order of preference used to choose prompting and evaluation procedures.31 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README documents command-line usage, publicly available prompts, result caching, and logging; versioned releases are published on GitHub and PyPI.145 |
| LimitationsAre known limitations documented? | Partial | The README documents operational limitations (no native multi-node evaluation, early-stage support for the PyTorch MPS backend, and request types some backends do not support). A general discussion of benchmark validity limits was not found in the pages read.1 |
What it is useful for
Run and use notes
- Since December 2025 the base package no longer bundles transformers or torch; backends are installed as extras such as lm_eval[hf], lm_eval[vllm], or lm_eval[api], per the README.1
Organization context
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.