Within the same hour, the two leading laboratories shipped new models and touched off an open price war. The flagship tier fell to four and twenty dollars per million input and output tokens — a fifth cheaper than a week earlier — while matching the quality of a costlier sibling and running better than thirty percent faster. Cache reads were cut by sixty percent.
The answer came almost immediately: a pair of smaller models priced at two and ten dollars, and at ten cents and fifty cents, per million — permanent rates, not promotions, each carrying a million-token context window and a ninety-percent discount on cached input. One makes half as many mistakes as its predecessor; the other matches last season's premium model at roughly one-hundredth of the cost.
Strip away the launch theatre and one figure organises the rest: the cost of a given level of machine capability has fallen about forty-seven percent every quarter for three years. That compounding is why both houses can slash prices and still widen margins — and why the interesting question is no longer which model is best, but what becomes valuable once the model itself is nearly free.
The answer, pursued across the inside pages, is that value is migrating off the model and into the machinery around it — the harness, the decision layer, and the humans who verify the work.
The new flagship is cheaper per token and quicker, and one tester migrated a 680,000-line codebase in under a day. But the upgrade ships with four breaking changes to existing integrations — check them before swapping model identifiers, not after.
Migration · verify before deploy
The challenger's two new models split the labour: one for complex, multi-step work; one for high-volume, simple tasks. Both were trained like the flagship, share a million-token context with 128k output, and take a 90% cached-input discount. The larger halves its error rate; the smaller undercuts last season's premium by ~100×.
The commoditisation, continued
"The model is becoming the commodity input. The moat is the machinery around it."
The most consequential launch of the week may not be a larger language model but a smaller, stranger one — a "decision model" that generates nothing at all. You hand it text and a question framed as a yes/no, a multiple choice, or a rating, in plain English, and it returns a calibrated probability over a fixed set of answers. It reads the input once and scores every option at ordinary human-glance speed.
The economics are the shock. A single decision costs about 4.2 cents per million input tokens, with output free — roughly eighteen times cheaper than the cheapest small chat model and more than two-hundred times cheaper than a frontier flagship. Latency runs between seventy and five-hundred milliseconds, most calls near a tenth of a second. The maker emerged from stealth in mid-September on a forty-million-dollar seed, founded by one of the original architects of instruction-tuned chat, and named the model for the nineteenth-century economist who observed that when a resource gets cheaper, we consume vastly more of it, not less.
The category error worth avoiding is treating this as a faster chatbot — "like calling a calculator a slow typewriter." A generator writes free text and can wander; a classical classifier is rigid and needs retraining; a decision model reads like the former and answers like the latter, returning a label and a confidence figure software can act on directly. Its primitives are simple: choose among options, score against a rubric, or estimate whether a statement is true. The questions run independently, in parallel, against shared state, and are recombined in ordinary code — "an if-statement whose condition understands what the customer meant."
The technique underneath is the part any engineering leader can steal today. Instead of asking a model the fuzzy question directly — "Is this email urgent?" — you decompose it into discrete checks: was it sent by a human, do you know the sender, would ignoring it cost you? Each is scored separately and combined in code. Editors who ran this pattern across ninety days of correspondence found it did the triage a person would; one demonstration had the model read 384 news stories to brief fifteen brands in a single pass.
For a decision-maker the fine print matters. The "confidence" number describes the shape of the output distribution, not an independently measured probability that the answer is correct — and a schema-valid reply can still be semantically wrong, naming a real department but the wrong team. Calibrated classifiers are not new; the fair benchmark is a fine-tuned encoder, not a chat model. An open-source rival already exists — a 421-million-parameter model, laptop-runnable under a permissive licence — though it needs fine-tuning where the paid one works from a cold start. The signal for a technology chief is less "buy this vendor" than "an entire class of your AI features is secretly classification, and classification just became almost free."
The sleeper result of the week pointed an AI at the scaffolding rather than the model — the tools, context, files and feedback loops that surround it. It fed 152 optimisation ideas into a selection loop; four mechanisms survived, validated on a held-out fifty-one-task benchmark. Across two different frontier models, token traffic fell between forty-five and forty-nine percent and real API cost dropped by half against the standard developer harnesses, at unchanged quality — an estimated eight to thirteen dollars saved for every hour of agent work.
The four survivors remove waste, not intelligence: a model reading a four-thousand-line log to use six lines of it; finished sub-tasks still sitting in context and billed every turn; the same large tool output replayed for the tenth time; an edit and its obvious validation sent as two separate round-trips. None of it makes the model smarter. All of it makes the bill smaller. A parallel technique from another lab regularised self-improving harnesses and improved out-of-distribution results across eight benchmarks while spending fewer tokens.
The counter-intuitive lesson on scale is to run fewer agents, more carefully. The widely-shared "thousand agents overnight" fantasy matters far less than five settings — dependency structure, file isolation, the tiers of persistence, refusing to let an agent grade its own work, and enforcing hard limits as rules rather than polite suggestions. The real wall is not a thousand agents; it is four or five sessions hammering one repository until review can no longer keep up.
A pipeline that converts research papers and their code into callable tools reached 91.2% accuracy against 80.3% for the same model working from the raw repository — an eleven-point jump from better scaffolding, not a better model. It converted 74 of 100 papers into working tools.
Evidence · the harness pays
A veteran design engineer's phrase for models that seem expert one day and fail the identical task the next. Her response: stop reading the code the agent writes, and instead write detailed specifications for how the agent must prove itself. Verification, not generation, is the job.
The craft shifts to specs
Web-search interfaces have long hit a wall of about sixty-five percent accuracy on questions that need second-by-second facts — prices, weather, scores. A new "knowledge" layer injects live, structured feeds directly into search and lifts a 155-question benchmark from 64.5% to 84.2% accuracy, at the same cost, by adding a single parameter to the existing call. Each answer carries a source and a timestamp, so a model can finally quote a live exchange rate without hallucinating one. The lesson for anyone building on retrieval: the ceiling was never the model's reasoning; it was the freshness of what you fed it.
A container platform launched an admin console that governs how agents execute, which networks and credentials they may touch, and which external tools they can call — defined once and enforced on every developer's machine, every action logged and exportable to a security monitor, each session sealed in a hardware-isolated sandbox. A separate vendor asked the blunt question every board eventually will: "can you prove what your agent did?" The tooling to constrain autonomous systems is arriving faster than the norms to use it — which means the moment to write the policy is now, while it is cheap and hypothetical.
At the frontier's edge, biosecurity is being framed as an offence-versus-defence arms race. Genomic language models treat DNA as a four-letter language over very long sequences; shown only mediocre examples ranked low-to-high, one model spontaneously produced better ones it had never seen. A predecessor was used to generate whole viral genomes later synthesised into functional viruses. The founder's uncomfortable thesis — defence is losing, so push the frontier harder — arrived the same week two lab chiefs were slated to brief the UN Security Council on AI safety, and newsletters asked who, exactly, gets to decide.
The model's price falls ~47% a quarter, and a research AI rewriting only the plumbing captured half of agent cost with the model untouched. The model is becoming a swappable commodity; the durable moat is a proprietary, self-improving harness plus a decision layer you own. The scarce 2027 role isn't the clever prompter — it's the person who audits context waste and owns the tool-call graph.
A four-cent decision model and a ten-cent generator are signals that most real workloads — triage, routing, "does this need a human?" — never needed a frontier model. The winning architecture pushes 99% of volume onto cheap classifiers and generators, reserving flagship reasoning for the 1% that compounds. Routing everything through the best model is a 100-to-240× tax on decisions a classifier answers in a tenth of a second.
Three unrelated sources converge: the highest-leverage agent setting is refusing to let an agent grade its own work; a senior engineer now writes specs for how the agent proves itself rather than reading its code; and an eleven-point accuracy gain came from scaffolding, not a better model. As generation nears free, the appreciating human skill is writing crisp, adversarial acceptance criteria.
Per-machine agent consoles, "prove what your agent did" auditing, and a biosecurity arms race are the same story at three scales: autonomous systems now touch credentials, networks and — in the extreme — synthesisable genomes, and the tooling to constrain them is outrunning the norms. Write the agent-action policy now, while it's cheap and hypothetical, not after the first incident makes it adversarial.
Everyone is racing the same two directions — more concurrent agents, cheaper tokens — so the non-consensus move is the opposite: cap concurrency below the point where review breaks down, and pour every dollar the price war hands you into an independent, well-paid verification layer whose only job is to prove your agents wrong. In a world of near-free generation, the firm that does less, more carefully out-ships the one running a thousand unchecked agents. Throughput was never the constraint. Trustworthy attention was — and the whole industry is optimising the axis that just turned abundant while starving the one that stayed scarce.