Skip to main content

Posts

Showing posts with the label DevTools

OSWorld 2.0 and the Finish-Line Problem in Long-Horizon Agent Evals

Agent demos are seductive. You prompt an agent to research a competitor, draft a brief, update a spreadsheet, and it does something plausible enough that everyone in the room nods. Then you ship it, and your users slowly stop trusting it, because it keeps getting 80% of the way through tasks and failing at the end in ways that are genuinely hard to diagnose. OSWorld 2.0, released in late June 2026 by the XLANG Research lab, is the benchmark that finally makes that failure mode measurable. The numbers are sobering: Claude Opus 4.8, the current leader, finishes just 20.6% of tasks end-to-end. GPT-5.5 plateaus at 13% regardless of whether you give it 150, 300, or 500 steps. What Makes OSWorld 2.0 Different The original OSWorld measured desktop computer-use on relatively short tasks. 2.0 extends it to 108 long-horizon workflows across seven professional domains: research, creative production, engineering, personal services, business and finance, administration and compliance, and healt...

DeepSeek V4-Flash Beats Its Own Pro on Agent Benchmarks

DeepSeek released the official public beta of V4-Flash-0731 on July 31, and the benchmark numbers are worth a second look. Same 284B MoE architecture, same $0.14 per million input tokens, and a post-training rerun that pushed agent benchmark scores past DeepSeek's own V4-Pro-Preview on every metric the company published. What Actually Changed Nothing about the architecture or scale changed. DeepSeek ran a new round of post-training, the phase that shapes how a pre-trained model uses its knowledge, not what knowledge it holds. The underlying structure (284B total parameters, 13B active per token, with CSA and HCA sparse attention layers) is identical to the Flash-Preview build. DeepSeek just worked the behavioral layer on top again. That alone shouldn't be remarkable. Except for how much it moved the needle. The Benchmark Story Here are the numbers DeepSeek published for Flash-0731 vs Flash-Preview vs V4-Pro-Preview: DeepSWE: 7.3 (Flash-Preview) to 54.4 (0731). That...

Supabase Evals: What Task-Specific Benchmarks Teach You About AI Coding Agents

Supabase open-sourced their evals framework last week (supabase/evals, Apache-2.0), and I think it's the most useful thing published about AI coding agent evaluation in months. Not because of which model topped the leaderboard. Because of how they designed the measurement itself. What They're Testing and Why It's Hard Supabase built their evals around a three-axis grid: products (database, auth, storage, edge-functions, realtime, cron, queues, vectors, data-api), topics (RLS, security, migrations, SQL, SDK, observability, self-hosting, declarative-schema), and stages (build, deploy, investigate, resolve). That last axis is where it gets real. "Build" is the easy part. Any capable coding agent can scaffold a schema. "Deploy" and "Investigate" are where agents start to diverge. "Resolve" is where you find out if an agent can fix a broken RLS policy without silently breaking three others it didn't know existed. They're runni...

Commits Are Up 180%. Releases Are Up 30%. Your Testing Pipeline Is the Bottleneck.

A 2026 NBER study quantified something a lot of engineering teams have already felt: AI coding agents increased commit rate by 180%, but software releases only grew by 30%. That gap is the story. You did not solve your velocity problem by adopting Claude Code or Cursor. You moved it downstream. On July 29, BrowserStack launched Test Companion , an agentic test automation tool built directly into the IDE. It's worth understanding why it exists and what it tells you about where AI tooling is headed. What the 180/30 Gap Actually Means When a coding agent can spin up a full feature in an afternoon, the constraint shifts. It's no longer "how fast can we write the code." It's "how fast can we trust that code enough to ship it." I've seen this play out on teams using Claude Code seriously. Output goes up fast. But PR queues get longer, QA cycles stretch out, and the release cadence barely moves. The agents didn't fix deployment velocity. They expo...

n8n's Native MCP Support: Your Existing Workflows Just Became Agent Tools

n8n shipped a significant update on July 29 that I've been waiting for since they announced MCP integration earlier this year. The headline: n8n now works as both an MCP client and an MCP server. That's not just a configuration option. It changes how you think about the workflows you've already built. What the July 29 Release Actually Shipped Four things worth paying attention to (full details in n8n's release notes ): Native MCP server mode. Any n8n workflow can now be exposed as a callable tool to external AI clients. Claude Desktop, a custom agent, your own LLM-powered interface. If the client speaks MCP, it can call your workflow. You toggle this at the workflow level, or bulk-enable it across an entire project folder from the new folder actions menu. Native MCP client mode. n8n's AI Agent nodes can now discover and call external MCP-compliant tools directly, without you writing a custom API wrapper or an HTTP request node. The agent resolves the tool lis...

Google's Gemini Managed Agents Just Got Cron Triggers, Budget Caps, and Hooks

Google pushed an update to Gemini API Managed Agents on July 28 that I've been waiting for. It's not a new model. It's the production scaffolding that makes running agents as background workers actually viable: budget caps, cron triggers, environment hooks, and a proper API to manage the sandboxes. None of this is conceptually new, but it fills the gaps that made me nervous about committing to the platform for anything beyond a demo. What They Shipped Six things landed in this release. Gemini 3.6 Flash is now the default model for managed agents, with no code changes required. But the parts that matter operationally are the new production primitives. Cron triggers. You bind an agent, a prompt, an environment, and a cron expression into a single persistent resource. The trigger fires on schedule without you doing anything. Each run reuses the same sandbox, so files written in one execution are visible to the next. If a run fails five times in a row, the trigger pauses i...

Claude Opus 5 Lands, and It Outperforms Fable 5 Where It Counts

Anthropic shipped Claude Opus 5 on July 24. Four days later I'm still sitting with the benchmark numbers, because they tell a story I didn't expect: on the metrics that actually matter for builder work, Opus 5 outperforms Fable 5 at exactly half the token price. That's not "close enough." That's a genuine inversion. What the Numbers Actually Say Fable 5 has been Anthropic's flagship since June. It costs $10 per million input tokens, $50 per million output tokens. Opus 5 costs $5/$25 per million tokens, which is identical to Opus 4.8. Same price point, better performance on the key evals. On Frontier-Bench v0.1, the agentic coding benchmark I weight most heavily right now, Opus 5 scores 43.3%. Fable 5 scores 33.7%. On SWE-bench Pro it posts 79.2%. On ARC-AGI-3, which tests novel problem-solving rather than memorized patterns, Opus 5 scores 30.2%. The next-best model scored 7.8%. Fable 5 wasn't even tested on it. Artificial Analysis pegged the weig...

Claude Sonnet 5: Cheaper Per Token, Not Always Cheaper Per Task

Anthropic shipped Claude Sonnet 5 on June 30, positioning it as the go-to model for agentic work at a price that doesn't require a flagship budget. At $3 per million input tokens and $15 per million output tokens (after introductory pricing ends August 31), it's priced at roughly 60% of what Opus 4.8 costs. That sounds like an easy call. It's more complicated. Where Sonnet 5 Actually Earns the "Agentic" Label The benchmark numbers that matter for builders aren't the ones that get the most press. On SWE-bench Pro, Sonnet 5 scores 63.2% compared to Opus 4.8's 69.2%. That 6-point gap is real. For a coding agent doing open-ended software engineering, it matters. But Terminal-Bench 2.1 tells a different story. Sonnet 5 scores 80.4%. Opus 4.8 scores 74.6%. That's the first time a mid-tier Sonnet has beaten its flagship sibling on a major coding benchmark, and the margin isn't narrow. For agents that work in the terminal, run shell pipelines, orchest...

What Actually Changed in Claude Sonnet 5 (and Why I'm Switching My Agents to It)

Claude Sonnet 5 shipped on June 30. Anthropic's framing was "a cheaper way to run agents," which is technically accurate but undersells what's actually interesting here. The benchmark I keep coming back to is Terminal-Bench 2.1: Sonnet 5 scores 80.4%, while Opus 4.8 scores 74.6%. The mid-tier model now beats the top tier on agentic terminal tasks. On GDPval-AA v2, a knowledge work eval, Sonnet 5 scores 1,618 Elo versus Opus 4.8's 1,615, basically a dead heat. Opus 4.8 still leads on SWE-bench Pro (69.2% vs Sonnet 5's 63.2%) and OSWorld-Verified (83.4% vs 81.2%). For hard multi-step coding and complex computer-use flows, Opus is still the call. But for most of what shows up in an agent loop, Sonnet 5 at $2 per million input tokens and $10 per million output tokens (introductory pricing, through August 31) is a genuinely different calculus than Sonnet 4.6 was. That said, migrating isn't a find-replace on the model ID. Three changes in Sonnet 5 will break ...

MCP 2026-07-28: The Biggest Revision Since Launch Goes Stateless. Here's What Breaks.

The biggest revision to the Model Context Protocol spec since launch drops July 28, and it has real breaking changes for anyone running MCP servers in production today. Not the kind where you bump a version number and everything still works. The kind where your infrastructure assumptions need to change. I've been through the release candidate. Here's what actually matters. The Stateless Core Is the Real Story The headline change: MCP is now stateless at the protocol layer. The initialize / initialized handshake is gone. The Mcp-Session-Id header is gone. Client metadata, previously exchanged once at connection setup, now travels in _meta on every request. The practical consequence: any server instance can handle any request. No more sticky sessions, no more shared session store, no more deep packet inspection at the gateway to route a client to the right backend. The MCP team puts it plainly : a remote server that needed all of that "can now run behind a plain roun...

Stop Optimizing the Wrong Variable

I've been watching xAI quietly build up to this. On July 8, they launched Grok 4.5 into developer access, then opened it to everyone on grok.com and the X app the next day. They called it their "flagship model for coding, agentic tool calling, and knowledge work." That framing is deliberate. This isn't a general-purpose model with coding bolted on as an afterthought. Most of the coverage led with the per-token price: $2 per million input tokens, $6 per million output. Competitive, but not shocking. The number that actually got my attention was 15,954. That's the average output token count Grok 4.5 uses to solve a task on SWE-Bench Pro. Claude Opus 4.8 uses roughly 67,020 tokens for comparable tasks on the same benchmark. A 4.2x difference. When you're running agent loops at any meaningful scale, that gap is where your bill comes from. Built for Agents, Tested Outside the Lab xAI didn't just run evals and ship. Grok 4.5 went through internal validation...

Claude Sonnet 5: Near-Opus Performance at 40% Lower Cost Changes Your Agent Routing

Anthropic dropped Claude Sonnet 5 on June 30, and the positioning is deliberate: this is the most agentic Sonnet model they've shipped. Close to Opus 4.8 in performance, at 40% lower cost at standard pricing, and 60% cheaper during the introductory window that runs through August 31, 2026. For anyone running multi-step agents at scale, this changes the cost math in a concrete way. What's actually in the release Sonnet 5 ( claude-sonnet-5 ) sits between Haiku 4.5 and Opus 4.8 in the lineup, but Anthropic is pitching it as the default workhorse for most agentic work. It carries a 1M token context window, 128k max output, and adaptive thinking. Fast latency. And it defaults to effort: high on the Claude API and Claude Code, meaning the model engages its full reasoning budget by default. That's the same default as Opus 4.8. Here's the full pricing picture now: Haiku 4.5 : $1/$5 per million tokens. Fast, near-frontier for simple tasks. Sonnet 5 intro (through Aug 31)...

Meta's First Paid Model Is Live. It Wins on Agents, Trails on Code, and Costs a Third of Opus.

Meta released Muse Spark 1.1 on July 9, 2026. It's their first closed-weight model behind a paid API, and the benchmarks tell an interesting story: it beats Opus 4.8 and GPT-5.5 on tool-use and agent tasks, trails them both on pure coding. That split matters for how you route it. Meta Just Went Closed (Sort Of) For years, Meta's play was Llama. Open weights, self-hostable, no API lock-in. That playbook built a huge developer ecosystem and kept Anthropic and OpenAI honest on pricing. Now Meta has a paid closed model at $1.25 per million input tokens and $4.25 per million output tokens. To put that in context: Opus 4.8 runs around $15/$75. GPT-5.5 Sol is $5/$30. Muse Spark 1.1 at $4.25 output is competing on price with Flash-tier models while claiming frontier-tier performance on agent benchmarks. That's a meaningful price point if the benchmarks hold. They're not abandoning Llama. Muse Spark is a separate product line from Meta Superintelligence Labs . Open weights ...

JADEPUFFER: An LLM Agent Just Ran a Complete Ransomware Attack

A few days ago, Sysdig's threat research team published their analysis of what they assess as the first fully autonomous LLM-driven ransomware attack. No human operator directed individual steps. An AI agent broke in, harvested credentials, moved laterally, encrypted a production database, destroyed the originals, and left a ransom note. The whole chain ran without a human at the keyboard. The name Sysdig gave it: JADEPUFFER. I want to walk through the attack chain, because the details matter more than the headline. The Entry Point Was a Tool Builders Use CVE-2025-3248 is an unauthenticated remote code execution flaw in Langflow , the visual builder for LLM workflows and agent applications. No login required. If Langflow is internet-facing and unpatched, anyone can POST arbitrary Python to an endpoint and it runs. CISA added this to its Known Exploited Vulnerabilities catalog in May 2025. Fixed in Langflow 1.3.0. Apparently enough people were still running older versions that...

At $0.87/M Output Tokens, DeepSeek V4-Pro Just Repriced Your Agent Architecture

DeepSeek made the 75% discount on V4-Pro permanent in late June. Not a promo extension, not a trial period. They called it an "efficiency gain being passed through." That framing matters. It means the new price floor is structural, not a marketing play designed to flip later. The numbers: $0.435/M input, $0.87/M output, and cache hits at $0.003625/M. For context: GPT-5.5 sits at $5/M input and $30/M output. Claude Fable 5 is $10/M and $50/M. DeepSeek V4-Pro is roughly 34x cheaper per output token than GPT-5.5. At that delta, you're not comparing pricing tiers anymore. You're looking at different economic regimes. What Actually Changed V4-Pro was already a serious model before the cut. It's a 1.6 trillion parameter MoE with 49B active params, a 1M token context window, and MIT-licensed. It scores 80.6% on SWE-bench Verified , the highest open-weights entry, tied with Gemini 3.1 Pro. The price cut didn't change the model. It changed what's economically v...

Apple's Foundation Models Framework Is Now a Model Router. Here's What Changes for Builders.

At WWDC26, Apple made a move that most coverage missed. They didn't just update the Foundation Models framework with new models. They restructured it into something closer to a model abstraction layer, one where your Swift code stays the same whether you're calling an on-device model, Apple's Private Cloud Compute, or a third-party provider like Claude or Gemini. That changes the architecture of iOS AI apps significantly. What Actually Changed The Foundation Models framework has existed since Apple Intelligence launched. But until now, it was essentially one thing: an on-device Apple model you called from Swift, with the privacy and latency benefits that come from never leaving the device. WWDC26 turned that into three distinct tiers accessible through one API: The existing on-device model (fast, private, capability-constrained) A new Private Cloud Compute model (bigger, reasoning-capable, 32K token context window) Third-party models including Claude and Gemini, cal...

Gemini 3.5 Pro Missed June. Four Researchers Left for Anthropic. Here's What I'm Watching.

Google made a specific promise at Google I/O on May 19: Gemini 3.5 Pro would be generally available by June. Sundar Pichai, when pressed on the timeline, said "give us until next month." The audience groaned audibly. They were right to. As of June 29, Gemini 3.5 Pro is still in limited Vertex AI enterprise preview. The public launch has been pushed to July . And in the same week the June deadline slipped, four senior Gemini researchers announced they were leaving for Anthropic . Google's AI coding teams have lost six researchers in five months. That combination is not damning by itself. But it is a signal, and I think it's the more interesting story here. What a Month's Slip Actually Costs A one-month delay sounds minor. It rarely is when you're managing a product roadmap around it. If your team planned feature launches, customer commitments, or integration timelines around Gemini 3.5 Pro hitting GA in June, you just ate a planning hit. The areas Google ...