Skip to content

A new model won’t save you

Three coding flagships shipped in three weeks, and the tempting move is to wait for the next one. Here's why that's usually the wrong call - and the one situation where it isn't.

A new model won’t save you

Three frontier coding models landed in three weeks, and every operator I know had the same first thought: maybe this is the one that finally makes the agents work. North Wayne is here to talk you off that ledge. Her point in today’s First Opinion isn’t that the new models are weak - they’re genuinely better - it’s that “wait for the next release” is a decision, and usually a losing one. A stronger model raises the return on knowing your own system; it never supplies that knowledge for you. Read this before you defer another roadmap. — Muximus


There’s a decision sitting on a lot of desks this month, and most people are about to get it wrong. Three frontier coding models shipped in about three weeks - Anthropic’s Claude Fable 5, Claude Sonnet 5 on June 30, and Z.ai’s open-weights GLM-5.2. If your AI agents have been underwhelming, the obvious move is to hold the roadmap, wait for the next release, and hope it clears the bar. Don’t. Waiting for a better model is the right call in exactly one situation, and for most of you this isn’t it.

Here’s the framework I’d use before you defer anything. The question was never which model. It’s how much the model knows about your business - and no release changes that answer for you.

When waiting is actually the right call

Give the wait-and-see position its due, because part of it is sound. On small, self-contained work - a narrow task, clear scope, a result you can check at a glance - a stronger model closes the gap on its own. One person fixing one well-bounded problem, a job with an obvious pass or fail: here each new generation genuinely helps, and you feel it immediately. If that’s the shape of most of your work, buy the upgrade and move on. You don’t need the rest of this.

But that’s not most operators. Most of you run large systems with history, conventions, and consequences - and that’s where the reasoning flips.

On a real system, the model’s IQ isn’t the bottleneck

The deciding factor on a real system isn’t how smart the model is. It’s how much it knows about your work: where to start, the conventions that hold in this corner of the business, the past choices that still bind you, the parts of the system nobody is allowed to break, the checks that actually mean something, and the numbers you watch after you ship. A smarter model guesses at all of that better. But if you haven’t defined the job, it’s still guessing - and a confident guess on a system you can’t afford to break isn’t an upgrade. It’s a liability in a nicer suit.

This is what every launch makes easy to forget. You watch the benchmark, the demo, the video of the model working for hours, and you conclude you can hand it the assignment and walk away. Usually you can’t - not because the model is weak, but because the assignment was never the whole job.

The better the model, the more expensive the near-miss

This is the part I most need you to hear, because it runs against instinct. A stronger model makes mediocre work more dangerous, not less.

Weak output looks weak. You catch it because it reads like a rough draft, and your people slow down and check it. Strong output looks almost finished. It touches the right places, adds its own checks, summarizes cleanly, and explains its choices with confidence - and it can still miss the customer edge case, the product decision nobody wrote down, the rule that lives only in one senior person’s head, or the step that has to be reviewed by hand before it goes out. The more polished the work looks, the less your team scrutinizes it, at exactly the moment the misses are the ones that surface in production and cost you. Capability raises how good the work can be and, in the same motion, raises how convincingly wrong it can be.

The sticker price is not the invoice

There’s a second version of the same trap, and it’s the one that’ll show up on your budget. A cheaper model is not automatically a cheaper bill.

Sonnet 5 launched at an introductory $2 per million input tokens and $10 per million output through August 31, then $3 and $15 - against Opus 4.8 at $5 and $25, per Anthropic’s own figures. Cheaper on the page. But in the same announcement Anthropic notes that Sonnet 5 changed how it counts the units you’re billed for: the same input now maps to roughly 1.0 to 1.35 times as many tokens, and the introductory price is set so the switch comes out “roughly cost-neutral” against the model it replaces. The headline rate dropped. The bill was engineered to hold.

Open weights don’t get you out of it either. Independent testing by Artificial Analysis and by Simon Willison found GLM-5.2 - the strongest open-weights model on Artificial Analysis’s independent index - to be token-hungry, burning noticeably more output per task than the model it replaced, which shrinks its per-call price edge on exactly the long-running work you’d adopt it for. The price on the vendor’s page is a rate. Your invoice is that rate multiplied by how your workflow uses the model - and that second number is yours, not theirs. Swap the model and you change the first number while leaving the one that actually moves your bill untouched.

What the upgrade actually changes

None of this means the new models don’t matter. They raise the ceiling, and that’s real: a model that holds a longer task without drifting, checks its own work, and reads your system more accurately is worth having. It just moves where the work sits. The new model doesn’t replace your workflow - it increases the return on a good one. Less of your effort goes into producing the work; more of it goes into defining the job: what good input looks like, what context has to go in, what proves a task is actually done, and where the model has to stop and hand back to a person.

So don’t onboard this month’s model by handing it your biggest, messiest problem on day one. Run it like a pilot with real success criteria - the discipline most teams skip. Pick one narrow, well-understood slice of work. Decide what “done well” means before you start: the checks, what a reviewer should expect, the metric you’ll read after it ships. Write that down so the model can follow it. Prove the model on that slice. Then, and only then, widen the scope. The upgrade compounds that groundwork; it doesn’t substitute for it. This is the argument made at TomerCode when Fable 5 landed, and every release since has only confirmed it.

The verdict

Don’t wait for the model that saves you. It isn’t coming, because the thing holding your agents back was never on the vendor’s roadmap - it’s on yours. A smarter model is better raw material: it can help more, run further, and save real manual effort. But if you can’t explain to it what “good” means inside your business, it will not work that out on its own. Spend the next sprint defining the work, not shopping for the release that makes defining it unnecessary. The next new model won’t save you from that job. It will only punish the places you skipped it - faster, and more convincingly than the last one did.

Add VarOps on Google