The most useful engineering result of the week is also the most commercially disruptive: the effort settings that look identical across models are backed by genuinely different machinery, and none of it can be replicated by copying a system prompt.
The techniques now in production include effort-conditioned supervised fine-tuning, per-token cost terms in the reinforcement-learning reward — where lower effort carries a higher token cost, producing shorter traces — separate effort specialists merged through on-policy distillation, per-mode context windows with length penalties, and hard inference-time budgets with forced early exit.
One approach, alternating budgeted and unconstrained reinforcement-learning phases, cut token consumption by 25 to 30% with little movement on benchmarks, and the effect transferred to unrelated evaluations.
Three findings deserve to be printed and pinned. The thinking tags everyone copies are purely cosmetic: any delimiter works and confers no reasoning gain. Fixed budgets cause overfitting to short solutions and destroy test-time scaling, which is why one laboratory gates budgets until per-problem accuracy clears a threshold. And in at least one case the ability to reason partially emerged rather than being trained.
The decisive economic consequence is that cost curves overlap. A smaller model at high effort can match a larger model at low effort — which means the question of which model to standardise on and the question of how hard it should think are the same question, and almost nobody is pricing them together.
The attempt to solve this automatically has already failed once in public. A router that selected effort on the user's behalf was judged more miss than hit and withdrawn. Automatic effort routing remains the acknowledged unsolved problem, sitting directly on top of the largest line item in most enterprise AI budgets.
Protocols & Memory
Three Protocols, One Stack
The apparent standards war resolved into complementary layers. One protocol handles agent-to-tool: the host application routes a formatted request to a server, which executes and returns structured output. A second handles agent-to-agent: a peer is discovered via a published capability card, delegated to, and — if it needs more input mid-task — pauses in an explicit input-required state and loops back. The third, agent-to-agent over REST with a manifest and synchronous or streamed replies, has been folded into the second. In production the first two are complementary, not competing. The standing warning: every new pipeline component is a new evaluation failure point.
The Moat Moves to Memory
A memory harness exposing six editable control surfaces scored 0.806 against a 0.722 baseline on a shell-agent evaluation, at lower cost. The broader migration is unmistakable: differentiation is moving from model access to orchestration, memory and tooling. In robotics the same lesson appeared — extending policy context by three orders of magnitude produced an 87% improvement in manipulation and completed a ten-stage assembly task that no baseline finished.