What is computer use, and why is it harder for AI than answering questions?
When you ask an AI a question, it reads text and produces text — one clean input, one clean output. Computer use is fundamentally different: the AI must look at a screen, decide what to click or type, act, observe the result, and repeat — an ongoing loop of perception and action across software it was never explicitly trained to understand.
Text generation has a clear finish line. Computer use doesn't. The model must plan across a live, changing interface: a button that moves, a page that loads slowly, a pop-up that interrupts. Every step is a new decision point, and small per-step error rates snowball into large task failure rates overall.
OpenAI's GPT-6 Astra leads on OSWorld 2.0, a benchmark simulating open-ended desktop tasks like navigating browsers and filling spreadsheets without human guidance. Reaching that level required not just a smarter model, but a fundamentally different kind of training.