ModelRefs / Tau-Bench — AI Glossary
Tau-Bench — AI Glossary
A benchmark evaluating AI agents on realistic customer-service and software tool-use tasks with ground-truth verifiable outcomes.
Overview
Tau-Bench (Yao et al. 2024) simulates real-world tool-use scenarios (airline booking, retail customer service) where agents interact with a database and user simulator. Graded by task success rate and policy compliance. Evaluates reliability, multi-step planning, and graceful error recovery in high-stakes service workflows.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Tau-Bench — AI Glossary.
Frequently asked questions
What is Tau-Bench?
A benchmark evaluating AI agents on realistic customer-service and software tool-use tasks with ground-truth verifiable outcomes.
What concepts are related to Tau-Bench?
Closely related concepts include gaia, osworld, tool agent.