Your Agent Has Too Much Context
The counter-intuitive lesson of the week is that a coding agent that keeps erring usually needs less context, not more. Bolt on three or four specifications at once and it treats a detail from one as fact about another, over-anchors on half-read material, and burns far more runs. Specs themselves have drifted from describing outcomes to smuggling in implementation detail — one ballooned past a hundred lawyerly bullet points — and every candidate “source of truth” fails in its own way: tests get gamed by the agent that writes them, code honestly shows how a system is (bugs and all) but never how you want it, and the human holding the real intent goes on holiday.
The fix is turning out to be old. The agent world is rediscovering the state machine, wrapping bounded-autonomy “circuits” in which each node is a mini-agent with only as much freedom as its slot deserves — AI on rails. Frontier harnesses point the same way, baking in fewer fixed instructions each generation and pulling skills into the window on demand.
Nowhere is the strain clearer than review. At one large engineering org, lines of code per human-landed change are up 106% and changes per developer per month up 51% in a year — with over 80% of that growth from agentic AI — even as the share reviewed within a day falls and some teams sit on thousands of pending reviews. Median change size is up 64%; developers say they want to spend barely 7% of their time reviewing; historically only a quarter write a detailed description of what they changed. Generation is now free and adjudication is the tollbooth. The workable answers all narrow the machine’s rope: split giant changes into stacked, independently reviewable layers, and write a testable finish line before the goal — “every face one colour,” not “make it look organized.”
The Software FactoryOne framework handed 200+ open issues to four specialised agents that separately reproduce, diagnose, verify and fix — cutting a five-year backlog toward zero. The split matters more than the count: no agent owns an issue end-to-end.
Don’t Be a Meat ProxyIf your reply is only “the model said” plus 800 words, you didn’t save the team work — you made everyone review the model for you.
Routing Goes PublicA cloud gateway now accepts standard requests and routes them to whichever frontier or open model fits, with rate-limiting and token tracking built in.
Standards On RailsEngineering guardrails delivered as a governed set of rules agents retrieve at the point of work — code review, design review, incident review.
Craft Notes
■Sandboxes for Agents You Can’t Trust
A container platform shipped a 4.0 release with a new default Rust runtime that shrinks footprint and startup latency, positioning itself as the isolation layer for agent workloads — each unpredictable agent gets its own lightweight VM rather than sharing a host kernel. Named backers already building on it include a payments giant, a hyperscaler and a leading chipmaker.
■Teacher and Student
The primitive worth knowing cold: distillation trains a small “student” to copy a big “teacher,” yielding a genuinely separate, cheaper model that can match or beat its source — unlike mere quantization or pruning. One popular open family is distilled straight from its larger sibling, then quantized for the device.