Release Wave
A Runtime, Not a Chatbot
This week’s frontier releases reframe the model as an execution engine and the chat window as a control plane. One flagship splits into three tiers — Sol, Terra and Luna — and adds programmatic tool-calling and parallel subagents; a companion “live” mode runs full-duplex voice that listens and speaks at once. A new “work” product sustains projects for hours across connected apps and files to emit editable documents, spreadsheets and decks, while a rival’s million-token model bundles computer-use and multi-agent orchestration behind a metered API.
Fast & Cheap
The Cheap-and-Fast Challenger
A new challenger model lands at roughly 80 tokens per second and about twice the token efficiency of the leaders, priced at $2 and $6 per million input and output tokens. Independent reviewers rate it “Opus-level” — around 4.6–4.7 on their internal scale — “not state of the art, but pretty good, and very fast,” earning a slot for long, multi-step work even from people who keep a pricier default for prose.
Cost Discipline
Route by the Job, Not the Brand
As the premium specialist model moves to pay-per-use at $10 / $50 per million tokens — double the daily-driver tier at $5 / $25 and the volume tier at $2 / $10 — the discipline is to send each task to the cheapest model that honestly clears it. Max the effort dial before switching models at all, and write the routing rules straight into the project config so the decision is automatic rather than a habit.
Middleware
The Router Boom
“API bankruptcy” has birthed a middleware category. Capable open weights now absorb 70–80% of routine traffic; premium models handle the rest. One fusion router scored 64.7% on a deep-research benchmark — beating a standalone frontier model at 60.0% and another at 58.8% — at roughly half the token cost, while a self-hosted proxy with a 10-millisecond classifier claims to cut bills up to 70%. The durable asset isn’t the router; it’s the private eval you tune it on.
Open Weights
The Frontier, Caught in Public
Open-weight models have closed the gap fast enough to reset strategy. One hits ~59% on a hard software-engineering benchmark cheaply and under a runnable licence; others handle the majority of everyday workloads outright. The uncomfortable coda is dependency risk: a June episode in which export rules forced a lab to pull its two strongest models offline worldwide is a reminder that a rented brain can be switched off by someone other than you.
Policy
A Six-Month Clock
A rumoured regulatory threshold would delay or block any open-weights model whose capability meaningfully exceeds today’s strongest closed systems — a line an open model, most likely from China, is expected to cross within about six months. Open models lack a central economic champion to lobby for them, so any restriction is likely to loosen far more slowly for open weights than for closed. The thing democratising capability is the thing most exposed to being legislated.