The Race Is Ten Months Wide, If You Measure It Honestly
A researcher stripped the scaffolding off the week's most celebrated model and asked it nine hundred maths questions. What was left tells a different story from the leaderboards.
The release landed on a Thursday and the leaderboards fell over by Friday. A 2.8-trillion-parameter open-weight model activating 16 of 896 experts, carrying a million-token context and native vision, took fourth place of 187 tracked models on the headline intelligence index at 57 against a tracked average of 31. It took first on the frontend code arena, ahead of both closed leaders, and first on a legal reasoning benchmark at 94.6% criterion pass against 93.6%.
Then somebody built a better ruler. The problem with modern benchmarks is that they measure the scaffolding as much as the model — tool use, retry logic, prompt harnesses, parallel sampling. So one researcher asked the model more than 900 maths questions demanding a single-number answer with no reasoning permitted, which strips the harness away entirely and leaves only what was learned in pretraining. The model landed between two closed releases from May and November of last year. Call it ten months behind, against headline scores that had it level with the frontier.
An independent national safety institute reached the same neighbourhood by a different route. Across 70 narrow cyber evaluations, the open-versus-closed gap has narrowed to 4–7 months from an average of 6–10 a year ago, with one Chinese open model sitting level with a closed flagship at a measured 4.3-month distance. Real convergence, then — but convergence of a different size than the headlines claimed.
The most useful artefact of the week was neither score but an accounting. One analyst decomposed the model's apparent gains into 10% genuine out-of-distribution capability, 25% benchmaxxing, 20% usemaxxing, 20% cheating and reward hacking, 10% induced innovation and 15% distillation from a Western frontier model. Only one of those six lines is the thing everyone thought they were reading about.
The meter, not the model, is the product.