τ-bench (tau2-bench)
Version 1.0.1
τ-bench is Sierra's open-source simulation framework for evaluating customer-service AI agents. In each domain an agent must follow a written policy and use tools while a simulated user takes part in the conversation, in turn-based text mode or full-duplex voice mode. The current repository, which carries the τ³-bench update, covers airline, retail, telecom, and banking-knowledge domains plus a mock domain.28
- Repository: GitHub repository (sierra-research/tau2-bench) (external site: github.com)
- Release notes: Release 1.0.1 (external site: github.com)
- Documentation: Task schema and evaluation guide (external site: github.com)
- Paper: τ²-bench paper (arXiv 2506.07982) (external site: arxiv.org)
- License: License (MIT) (external site: raw.githubusercontent.com)
Availability and license
Overall availability
Source code and domain task files on GitHub under the MIT License. Running an evaluation requires API access to the language models used for the agent and the simulated user.132
Availability is separate from permission: read the license before using or redistributing.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| CodeIs the evaluation code published? | Public | Published on GitHub under the MIT License.13 |
| Tasks / dataAre the tasks or test data available? | Public | The domains README documents per-domain data files stored in the repository under data/tau2/domains: policy, task definitions, voice tasks, task splits, and databases.4 |
| MethodologyIs the method for scoring described? | Public | The evaluation guide explains how rewards combine database-state, communication, and assertion checks, and the README cites the τ-bench, τ²-bench, τ-Voice, and τ-Knowledge papers.527 |
| ReproducibilityAre instructions for reproducing results published? | Public | The README documents installation with uv, API-key setup through LiteLLM, and example run commands; versioned releases and a changelog are published, and a pre-v1.0.1 tag is kept for reproducing earlier grading.26 |
| LimitationsAre known limitations documented? | Partial | The README states that, after grading fixes to the banking_knowledge domain, results from versions before 1.0.1 are not comparable with later results (other domains are unaffected), and the evaluation guide notes that the reference action sequence is only one valid solution path. No general discussion of the benchmark's validity limits was found in the pages read.25 |
What it is useful for
Run and use notes
- The README requires Python 3.12 or 3.13 (>=3.12, <3.14) and installs with uv; voice features need extra system packages such as portaudio and ffmpeg.2
Organization context
U.S. eligibility
Eligible · basis: U.S.-governed project
The project is maintained in the sierra-research GitHub organization, its MIT license names Sierra Research as copyright holder, and Sierra's own blog says Sierra introduced τ-bench, presents τ²-bench as building on it, and links to this repository. Sierra Technologies, Inc. is a Delaware corporation headquartered in San Francisco, California, per its modern slavery statement.1389
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.