← Glossary Term
OSWorld
A benchmark that measures how well an AI model can operate a real computer desktop, clicking through applications to finish everyday tasks.
Most benchmarks ask a model questions. OSWorld puts it in front of an actual operating system and asks it to get something done: open a file, edit a spreadsheet, change a setting, work across several applications. Scores are the percentage of tasks completed correctly, and they are much lower than on text benchmarks because driving a desktop involves a long chain of steps where any one mistake ends the run.
It has become the standard reference point for what people call computer use. When a lab claims its model can operate your machine on your behalf, OSWorld is usually the number it quotes.
Mentioned in