ModelRefs / Agents Benchmarks — Top AI Models
Agents Benchmarks — Top AI Models
Tool-use, function-calling, web/computer-use, and multi-step task execution. Measure end-to-end agentic task completion under real environments.
Overview
Tool-use, function-calling, web/computer-use, and multi-step task execution.
What this category is for: Measure end-to-end agentic task completion under real environments.
Benchmarks in this category
- τ-bench — Tool-augmented agent benchmark on retail & airline customer-service tasks.
- BFCL v3 — Berkeley Function-Calling Leaderboard, parallel + multi-turn tool use.
- WebArena — Realistic web agent tasks across 5 self-hosted sites.
- OSWorld — Computer-use agents across real OS apps (Ubuntu/Windows/macOS).
- GAIA — General AI assistant benchmark — multi-tool real-world questions.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Agents Benchmarks — Top AI Models.