The Flagship
A Model That Plans Better Than Your Instructions
The first of a new top-tier family — sitting a rung above the previous flagship — has upended a habit its power users spent years perfecting. A first-day teardown reports that carefully staged, step-by-step prompts produced worse results, because the model now sets its own direction, allocates effort across the parts of a task, and kills its own weak assumptions before returning an answer.
The practical advice that follows is a genuine shift in craft: stop hand-writing reasoning and instead build the apparatus around the model — a memory it can draw on, a verifier that checks its work, explicit boundaries, and a deliberate control over how much effort a given task deserves. The prompt is no longer the product; the loop is.
The Open Frontier
China's Weights Get Heavier
A 1.6-trillion-parameter model, trained on custom domestic chips and released under an MIT licence, now runs agentic coding workloads under an open harness with, its makers say, stable repository-level edits and task execution. It arrives beside a no-code voice-agent builder and a crop of local-first models winning developer mindshare.
Underneath the announcements, a budget-minded engineering note delivered the reality check: a new four-bit build of a 27-billion-parameter open model is useful, but quantizing its attention path pushes it toward endless "thinking" and lower accuracy — the usable variants land between 20 and 29 gigabytes depending on which layers stay in higher precision. A rival speculative-decoding method points the same way: the open race is now fought on the economics of inference.
The Stack Fills In
Plumbing, Everywhere You Look
The week's smaller releases sketched a stack maturing above the models themselves: always-on agents arriving on phones, hosted tool-protocol servers shipping from a major social platform and a browser engine alike, voice agents folded into a deployment gateway, a foundation model aimed at tabular data, memory reframed as something a model can be trained to use, and a fresh arena for benchmarking how well agents orchestrate their own sub-agents. A multi-billion-dollar "frontier" venture and a lively debate over hidden messages in coding-agent output rounded out a week whose center of gravity has clearly moved from the model to the harness around it.