ModelRefs / WebArena Leaderboard — AI Model Scores
WebArena Leaderboard — AI Model Scores
Realistic web agent tasks across 5 self-hosted sites. Current leaders, methodology, and citation sources for WebArena.
Overview
Realistic web agent tasks across 5 self-hosted sites.
How it is measured: End-to-end task success rate.
How this benchmark is scored
| Category | agents |
|---|---|
| Maximum score | 100 % success |
| Direction | Higher is better |
Primary source: https://webarena.dev/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to WebArena Leaderboard — AI Model Scores.