A model called OX Alpha showed up on OpenRouter on August 20. No lab name. No announcement. No documentation beyond a spec sheet: 1,048,576-token context, text and image and video input, function calling, free to use until August 27. Within 24 hours, people were routing production traffic through it. The benchmark claim driving that adoption: 80% on DeepSWE Pass@1. For context, the numbers being passed around put Claude Fable at 65% and GPT-5.6 Sol at 52%. Those numbers came from a ten-task community test, not an audited leaderboard, but that didn't slow anyone down. I'm not here to litigate whether OX Alpha is actually better at coding. I'm more interested in what happened to people's due diligence. What "Stealth" Actually Means on OpenRouter OpenRouter runs a program where labs can preview models anonymously. The provider shows up as "Stealth," and the model name is whatever the lab chooses. The idea is straightforward: labs get real-world ev...
Agent demos are seductive. You prompt an agent to research a competitor, draft a brief, update a spreadsheet, and it does something plausible enough that everyone in the room nods. Then you ship it, and your users slowly stop trusting it, because it keeps getting 80% of the way through tasks and failing at the end in ways that are genuinely hard to diagnose. OSWorld 2.0, released in late June 2026 by the XLANG Research lab, is the benchmark that finally makes that failure mode measurable. The numbers are sobering: Claude Opus 4.8, the current leader, finishes just 20.6% of tasks end-to-end. GPT-5.5 plateaus at 13% regardless of whether you give it 150, 300, or 500 steps. What Makes OSWorld 2.0 Different The original OSWorld measured desktop computer-use on relatively short tasks. 2.0 extends it to 108 long-horizon workflows across seven professional domains: research, creative production, engineering, personal services, business and finance, administration and compliance, and healt...