ModelRefs / Tau-Bench — AI Glossary

Tau-Bench — AI Glossary

A benchmark evaluating AI agents on realistic customer-service and software tool-use tasks with ground-truth verifiable outcomes.

Overview

Tau-Bench (Yao et al. 2024) simulates real-world tool-use scenarios (airline booking, retail customer service) where agents interact with a database and user simulator. Graded by task success rate and policy compliance. Evaluates reliability, multi-step planning, and graceful error recovery in high-stakes service workflows.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Tau-Bench — AI Glossary.

Frequently asked questions

What is Tau-Bench?

A benchmark evaluating AI agents on realistic customer-service and software tool-use tasks with ground-truth verifiable outcomes.

What concepts are related to Tau-Bench?

Closely related concepts include gaia, osworld, tool agent.