ModelRefs / Terminal-Bench Leaderboard — AI Model Scores

Terminal-Bench Leaderboard — AI Model Scores

Agentic command-line task suite (Stanford / Anthropic). Current leaders, methodology, and citation sources for Terminal-Bench.

Overview

Agentic command-line task suite (Stanford / Anthropic).

How it is measured: Task success rate across 100 sandboxed Linux tasks.

How this benchmark is scored

Categorycoding
Maximum score100 % success
DirectionHigher is better

Primary source: https://www.tbench.ai/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Terminal-Bench Leaderboard — AI Model Scores.