ModelRefs / OSWorld — AI Glossary
OSWorld — AI Glossary
A benchmark evaluating computer-using agents on real desktop tasks across Ubuntu, macOS, and Windows. GPT-4V at launch: ~13% success; humans: 72%.
Overview
OSWorld (Xie et al. 2024) presents 369 real computer tasks (LibreOffice editing, Chrome navigation, file management, VsCode use) with functional verification. Agents observe screenshots and generate mouse/keyboard actions. GPT-4V at launch: ~13% success; humans: 72%. Defines the frontier for GUI/CUA evaluation.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to OSWorld — AI Glossary.
Frequently asked questions
What is OSWorld?
A benchmark evaluating computer-using agents on real desktop tasks across Ubuntu, macOS, and Windows.
What concepts are related to OSWorld?
Closely related concepts include computer using agent, gaia, web agent.