Agent demos are seductive. You prompt an agent to research a competitor, draft a brief, update a spreadsheet, and it does something plausible enough that everyone in the room nods. Then you ship it, and your users slowly stop trusting it, because it keeps getting 80% of the way through tasks and failing at the end in ways that are genuinely hard to diagnose. OSWorld 2.0, released in late June 2026 by the XLANG Research lab, is the benchmark that finally makes that failure mode measurable. The numbers are sobering: Claude Opus 4.8, the current leader, finishes just 20.6% of tasks end-to-end. GPT-5.5 plateaus at 13% regardless of whether you give it 150, 300, or 500 steps. What Makes OSWorld 2.0 Different The original OSWorld measured desktop computer-use on relatively short tasks. 2.0 extends it to 108 long-horizon workflows across seven professional domains: research, creative production, engineering, personal services, business and finance, administration and compliance, and healt...