For most of this year, if you asked me how to compare coding agents, I'd have pointed you at SWE-bench. The safer answer now is: don't. SWE-bench Verified died in February. OpenAI audited it and found frontier models could reproduce the original human-written patches verbatim, which meant scores reflected training contamination, not capability. They pulled their own numbers from Verified and recommended SWE-bench Pro instead. Then in July, OpenAI audited Pro and found roughly 30% of the 731 tasks are broken. Not "hard." Not "noisy." Broken. Their automated pipeline flagged 200 tasks (27.4%). Human reviewers tagged 249 (34.1%). OpenAI retracted their Pro recommendation and called on the broader evaluation community to start over. So the fallback for the fallback is gone. And models are still shipping press releases citing SWE-bench scores. How We Got Here SWE-bench Verified made sense when it launched. Real GitHub issues, real test suites, tasks that...