ModelRefs / OSWorld Leaderboard — AI Model Scores
OSWorld Leaderboard — AI Model Scores
Computer-use agents across real OS apps (Ubuntu/Windows/macOS). Current leaders, methodology, and citation sources for OSWorld.
Overview
Computer-use agents across real OS apps (Ubuntu/Windows/macOS).
How it is measured: 369 tasks; success rate with screenshot loop.
How this benchmark is scored
| Category | agents |
|---|---|
| Maximum score | 100 % success |
| Direction | Higher is better |
Primary source: https://os-world.github.io/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to OSWorld Leaderboard — AI Model Scores.