largest open model
A new open model matched the frontier overnight at a fraction of the cost — and the labs that built the frontier are quietly changing which game they are playing.
A single model release reset the industry's centre of gravity this week. The new system is an open-weight mixture-of-experts of 2.8 trillion parameters with a one-million-token context window, and a decoding method that reprocesses only the changed slice of that context to run 6.3× faster at long range. It took first place on a live front-end coding leaderboard, ahead of the strongest Western flagship, and placed third on a broad intelligence index — the sort of result that used to take a year and a nine-figure budget.
What makes it a reckoning rather than a headline is the price. Within eight days, three separate labs shipped frontier-class capability at a fraction of prevailing cost, one at under half the going rate for a top proprietary model. Independent trackers now put the best open weights roughly four months behind the absolute frontier, at output prices as much as 150× cheaper. Raw intelligence, in other words, is turning into a utility — and the firms that trained the frontier are pivoting from selling tokens to selling the products, skills and connectors wrapped around them, where the switching costs actually live.
For anyone building atop these systems the calculus flips overnight: capability that cost a fortune in spring is now a commodity to be shopped by the task. Which surfaces the week's quieter revelation — the scarce input is no longer the model at all. It is the human attention required to verify what a fleet of tireless machines now produces before dawn.
A major lab's next family arrives as a trio — a flagship that reportedly beats the best Western model while spending 54% fewer output tokens, a balanced mid-tier at a lower price, and a fast, cheap runt. Event-triggered workspace agents and parallel coding sub-threads ship alongside it.
A celebrated founder raised $2B at a $12B valuation — the largest seed ever — then released a 975-billion-parameter model that is, by admission, below state of the art. The tell: it is a loss-leader for selling fine-tuning as a practical stand-in for the learning that frozen weights cannot do.
“When generating work is nearly free, the only thing that doesn't scale is a human's capacity to check it.”
The week's landmark model earns its speed through architecture, not brute force. It activates just sixteen of eight hundred and ninety-six experts per token, and its attention variant caches the unchanged context so that only the new fragment is recomputed — the source of its 6.3× long-context advantage. It writes and repairs code directly from live screenshots, is already served through a public API and coding tool, and its full weights are scheduled for public release on the 27th, held back only through a trial window.
The release did not arrive alone. In little over a week, a social-media giant, a rocket-and-AI venture and the open-weight newcomer each shipped a model competitive with the leading proprietary system — one at under half its price. Optimists count six or more viable model builders forming; pessimists foresee two or three survivors clinging to ninety-percent inference margins. Either way, the floor has fallen out of the token, and the argument has moved from who is smartest to who is cheapest per finished task.
The much-discussed 975-billion-parameter model, trained with an unusual optimiser pairing across tens of millions of asynchronous reinforcement rollouts, is not meant to win benchmarks. It is meant to prove a business: host a thousand cheaply fine-tuned variants for roughly the cost of one, and sell that customisation as the answer to a hard truth — a shipped model's weights are frozen and cannot learn from the conversation in front of them.
The scaling logic is escaping software. The uncomfortable lesson that hand-coded human knowledge eventually loses to compute and data is now reaching the wet lab: candidate drugs surfacing from massive-dataset pattern-matching rather than elegant theory, instruments reading thousands of blood proteins at once, and dexterous robots poised to run around-the-clock automated laboratories in tight, agent-driven experiment loops.
Read beside a head-of-state address positioning one nation as the developing world's AI partner, the free release looks less like charity and more like strategy. When you trail by a season, you champion sharing — capturing the mindshare, tooling and standards of everyone priced out of frontier bills. The prize is not margin; it is becoming the default substrate on which the next billion builders stand.
The most experienced practitioners have stopped writing most of their own code. The creator of a leading coding agent now describes supervising fleets of parallel agents — often from a phone — that repair failing continuous-integration runs, review pull requests on separate branches, and are set to argue against one another's work before a human ever looks. The scale is already concrete: one music company runs agents across more than twenty million lines of code; a delivery giant has handed its assistant company-wide tools.
The mechanism that makes this reproducible is deceptively small. A "skill" is a short structured file — a name, a description, a workflow — that the system can test on its own before packaging; connectors let an agent reach outward into mail, chat and the web; and plugins bundle both so a single prompt can chain research, drafting and delivery. The practical craft is learning what not to load: one emerging pattern builds a separate agent per model, so lean open-weight models pull in a heavyweight capability pack while frontier models are spared the wasted context.
The frontier of the craft is now orchestration and restraint, not syntax. The winning skill is no longer typing the function; it is designing the loop, choosing the model, and deciding which of a dozen tireless workers to trust.
A newly published top-ten catalogues the attack surface that opens once skills can act rather than merely answer: malicious skills, supply-chain compromise, over-privileged access, poisoned metadata, untrusted external instructions, weak isolation, silent update drift, poor scanning, absent governance and unchecked cross-platform reuse.
A survey of platform-engineering leaders found 93% had already suffered an AI-caused incident, yet only 30% had any formal AI policy. In the same week, a runaway billing alert projected $140 billion in a month on an account that normally spends a few dollars.
The AI-semiconductor complex shed more than a trillion dollars of market value in under two months even as usage data kept breaking records — a memory maker off thirty percent in three weeks, a cloud upstart nearly halved, the bellwether chip designer alone losing a trillion. Yet the four largest cloud buyers are tracking roughly $725 billion of capital spending this year, up seventy-seven percent, with forecasts of big-tech capex crossing a trillion in 2027. The dissonance — collapsing equity, exploding spend — is either the first crack in the bubble or a textbook case of falling prices detonating demand.
The subtler read came from the software tape, where the sell-off proved selective rather than broad. Cash-flow multiples have sunk to decade lows, but a fifty-point gap has opened between the winners — security, observability, vertical specialists — and the losers in horizontal software and ad-tech. The verdict traders are pricing: software alone is no longer a moat.
A clinical-answers startup is likely to pass on a round that would value it near $20 billion — despite revenue approaching $300 million annualised, ninety-percent gross margins, and less than five percent of its ad inventory sold. Growth that fast makes fresh capital look like unwanted dilution.
A three-year-old service that routes developers across four hundred models is fielding takeover interest at a figure in the billions — a steep premium on a recent $1.3 billion mark — after annualised revenue leapt five-fold to $50 million.
A mature streaming leader now manufactures growth internally: quarterly revenue up thirteen percent to $12.6 billion, an ad tier past 250 million monthly users chasing some $3 billion in ad revenue, and live events — barely five percent of content spend — driving six of its ten biggest sign-up days. Ninety-seven billion hours were watched in six months; the shares still fell forty percent on the year.
The quiet lesson for anyone buying intelligence: sticker price is a vanity metric. One flagship's new tokenizer emits about thirty percent more tokens on identical code than its old one — a silent price rise hidden beneath the per-token quote.
Stack the week's facts — open weights four months behind, output tokens 150× cheaper, three cheap frontiers in eight days — and one conclusion holds: raw intelligence is becoming a utility. The counter-intuitive twist is that this rewards the labs that see it coming. Pivoting from selling tokens to selling products, skills and connectors is not a retreat; it is a flight to where switching costs and data actually live. The firms that lose are the ones still treating "access to a good model" as a moat.
Every instinct says the binding constraint is compute or model quality. Invert it. Inference just fell roughly 150-fold; the best practitioners already run fleets of agents; a schema harness hit 99% on a hard reasoning benchmark by writing each puzzle's mechanism as executable code. When producing work is nearly free and nearly infinite, the thing that does not scale is a human's capacity to check it — and the governance data (93% hit by an incident, 30% with a policy) shows the gap already open. So stop optimising for the smartest model and optimise for the highest ratio of self-verifying output — work that arrives with its own proof, test or receipt. The next defensible companies will not sell intelligence; they will sell trust per human-minute. Whoever makes verification as cheap as generation owns the decade.
The most valuable fact of the week was not a model release; it was that a leading tokenizer now emits thirty percent more tokens on the same file than its predecessor. Fold that into a five-multiplier stack and "price per million tokens" becomes meaningless — a "cheaper" model can be dearer per finished job. The edge, still unclaimed by most teams: benchmark on your own prompts in dollars-per-completed-task, not published rates.
A leading lab selling a $70 log-off basketball in the same week it ships workday-automating agents, while a streaming giant wages a "battle for attention," rhyme more than they seem to. When the companies harvesting attention begin productising rest, human focus and wellbeing are becoming the contested, monetisable frontier. Expect "off-switch," "focus," and "verified-human-time" to graduate from features to premium tiers. The contrarian builds for the backlash, not the binge.
Markets read the free 2.8-trillion-parameter release as bearish for Western labs and dumped chip stocks. But paired with a national pitch to be the developing world's AI partner, it reframes as standards capture, not a fire sale. The uncomfortable question for Western policy: a three-quarter-trillion-dollar capital race optimises for the wrong prize if the world ends up building on the free thing.