Skip to main content

Posts

Showing posts with the label Google

Google's Gemini Managed Agents Just Got Cron Triggers, Budget Caps, and Hooks

Google pushed an update to Gemini API Managed Agents on July 28 that I've been waiting for. It's not a new model. It's the production scaffolding that makes running agents as background workers actually viable: budget caps, cron triggers, environment hooks, and a proper API to manage the sandboxes. None of this is conceptually new, but it fills the gaps that made me nervous about committing to the platform for anything beyond a demo. What They Shipped Six things landed in this release. Gemini 3.6 Flash is now the default model for managed agents, with no code changes required. But the parts that matter operationally are the new production primitives. Cron triggers. You bind an agent, a prompt, an environment, and a cron expression into a single persistent resource. The trigger fires on schedule without you doing anything. Each run reuses the same sandbox, so files written in one execution are visible to the next. If a run fails five times in a row, the trigger pauses i...

Gemini 3.6 Flash Is Out: What the Token Efficiency Gains Actually Mean for Agent Loops

Google dropped three new models into its Gemini Flash tier on July 21. Most coverage fixated on the naming confusion (what is a 3.6 Flash versus a 3.5 Flash-Lite anyway?) and on the Gemini 4 name-drop buried in the announcement. But the story that actually matters for teams running production agents is more concrete: the economics of calling Flash just got meaningfully better, and in a way that compounds harder than the headline numbers suggest. What actually changed with 3.6 Flash Gemini 3.6 Flash replaces 3.5 Flash as the default workhorse in the family. Same general job: coding, knowledge work, multimodal tasks at production scale. Two things changed: it scores higher on every major benchmark, and it uses fewer tokens to get the same work done. On DeepSWE , 3.6 Flash scores 49% versus 37% for 3.5 Flash. MLE-Bench: 63.9% versus 49.7%. OSWorld-Verified (computer use): 83.0% versus 78.4%. For a Flash-tier model to widen the gap that much while staying in the same cost bracket is a ...

What Noam Shazeer Leaving Google for OpenAI Actually Means

On June 18, Noam Shazeer posted on X that he was joining OpenAI. His title: lead for AI architecture research, confirmed by OpenAI's chief research officer Mark Chen. If that name doesn't mean anything to you, here's the context that makes it matter. Shazeer is one of eight co-authors of "Attention Is All You Need," the 2017 paper that introduced the Transformer architecture. The architecture every major language model runs on today. GPT-5.5, Claude Fable 5, Gemini 3.5 Flash, Llama, Mistral, all of them. The ideas in that paper are as foundational to modern AI as UNIX was to operating systems. He left Google after that paper, co-founded CharacterAI, and built it into a consumer AI product with enormous scale. In August 2024, Google paid approximately $2.7 billion (structured as a technology license from CharacterAI) to bring Shazeer and a cohort of researchers back into Google DeepMind. His role: VP of engineering and co-lead of Gemini, specifically owning the ...

Gemini 3.5 Pro Missed June. Four Researchers Left for Anthropic. Here's What I'm Watching.

Google made a specific promise at Google I/O on May 19: Gemini 3.5 Pro would be generally available by June. Sundar Pichai, when pressed on the timeline, said "give us until next month." The audience groaned audibly. They were right to. As of June 29, Gemini 3.5 Pro is still in limited Vertex AI enterprise preview. The public launch has been pushed to July . And in the same week the June deadline slipped, four senior Gemini researchers announced they were leaving for Anthropic . Google's AI coding teams have lost six researchers in five months. That combination is not damning by itself. But it is a signal, and I think it's the more interesting story here. What a Month's Slip Actually Costs A one-month delay sounds minor. It rarely is when you're managing a product roadmap around it. If your team planned feature launches, customer commitments, or integration timelines around Gemini 3.5 Pro hitting GA in June, you just ate a planning hit. The areas Google ...

OKF: Why Your Agent's Context Layer Is the Problem, Not Your Retrieval Strategy

Every agent project I've built that touches internal data hits the same wall. The agent needs context: what is this BigQuery table, what do the columns mean, how does it join to the orders table, what's "monthly active users" in your org and not the textbook definition. You end up dumping SQL schemas into the system prompt, pointing at Confluence pages, writing a bespoke context builder that assembles fragments before each request. It works, barely, and it doesn't travel. Move to a different team's data, start a new project, and you're rebuilding it from scratch. Google Cloud published a spec on June 12, 2026 that addresses exactly this: the Open Knowledge Format (OKF), v0.1. It formalizes what Andrej Karpathy called the "LLM wiki" into a portable, interoperable format. What OKF Is (and What It Isn't) OKF is not a service or a platform. It's a file format. The spec fits on a single page. Your knowledge base is a directory of markdow...

Gemini 3.5 Flash and the End of 'Use the Biggest Model' for Agents

I've been defaulting to Opus-tier or GPT-5.5 for anything agent-related because that felt like the safe call. Better reasoning, better tool use, better outcomes. Flash-tier models were for batch jobs, summaries, things where you didn't care that much about output quality. That calculus broke for me after spending time with the Gemini 3.5 Flash benchmarks . The model went GA on May 19 at Google I/O. The number that got my attention: 83.6% on MCP Atlas, a benchmark specifically for multi-step tool orchestration using Model Context Protocol servers. That puts it 8.3 points ahead of GPT-5.5 (75.3%) and 4.5 points ahead of Claude Opus 4.7 on the same eval. "Flash" doesn't mean what it used to. What MCP Atlas Is Actually Measuring MCP Atlas tests whether a model can chain together multiple tool calls across MCP servers, recover from partial failures, and complete multi-step tasks without going off-script. It's not a writing or reasoning benchmark. If you're ...