A one-year-old startup pointed coding agents at a rival AI chip and reached 92 percent of its theoretical peak before the working day was out. The number that should unsettle Nvidia isn't ninety-two. It's ten.
Every competitive advantage carries an invisible denominator: the number of engineer-years a rival would need to reproduce it. We almost never write the number down, yet we price it into everything — valuations, roadmaps, the flat confidence with which we tell a board that nobody can catch us.
This week a company barely a year old divided that denominator by several orders of magnitude, in public, against the most storied moat in computing. It did not build a faster chip. It pointed coding agents at somebody else's chip and had them write the software layer that makes silicon usable — the thing Nvidia has been accreting, kernel by kernel, since February 2007.
The instinct is to argue about whether the demonstration was real. Resist it; that argument is cheap and we run it on page six anyway. The lens for this issue is more uncomfortable. Sort your own advantages into two piles. Pile one is expensive because it was hard to build. Pile two is expensive because it is hard to verify. Only the second pile is still a moat.
Everything else in these ten pages — a humanoid that learned to walk and grasp with one brain, a reactor reclassified from construction project to manufactured unit — is the same story wearing different clothes. The unit of accumulation is changing. Most of us have not repriced.
A note on method, since it bears on how much weight the reader should put on any of it. The central claim in this issue is self-reported by an interested party, and we say so on page six at greater length than the claim itself gets on page four. That is deliberate. A magazine about innovation that only prints the things it can be certain of would be a very thin magazine, and a magazine that prints them without their error bars would be a worse one.
The useful posture, as ever, is neither credulity nor dismissal. It is to ask what would have to be true for the claim to matter — and then to go and look at your own balance sheet with that question in hand.
How a $15 million startup turned the deepest software moat in the industry into a benchmark — and why the interesting number is the clock, not the score.
Nvidia's chips are not the reason Nvidia is Nvidia. The reason is a software layer called CUDA, first released on 16 February 2007, and the roughly six million developers who have since built their working lives on top of it. Buy an accelerator from anyone else and you inherit a blank sheet: no kernels, no libraries, no profiler, no nineteen years of answered questions. That blank sheet — not clock speed, not memory bandwidth — is what has kept a trillion-dollar company safe.
Every challenger has walked into it. AMD, Google, Amazon, Cerebras, Groq, d-Matrix, Rebellions — each with silicon that looks competitive on a spec sheet and a software stack that looks like a to-do list. Amazon's own internal documents, according to reporting on the company's chip programme, once flagged CUDA as a major roadblock to adopting the AI processors Amazon had designed itself. When the largest cloud on earth cannot escape a competitor's software, the word "moat" stops being a metaphor and becomes an accounting entry.
The blank sheet has always had a price, and until this year that price was denominated in headcount and calendar. Standing up a usable inference stack on new silicon means writing a kernel for every operator a model touches, mapping a memory hierarchy nobody has documented publicly, building a scheduler that keeps every compute unit fed, and then — the unglamorous majority of the work — proving that the numbers coming out match the numbers a GPU would have produced. Specialist team. Multiple years. Most companies do not have the team and cannot buy the years.
Two things happened in close succession. On 20 July an inference-software startup called Infinity announced a $15 million seed round. On 3 August its founder told Business Insider that his team had used AI coding agents to rebuild CUDA-like software for the chip company d-Matrix in about ten hours — and framed it, pointedly, as proof that one of Nvidia's biggest moats is being crossed.
Ten hours is not a performance claim. It is a claim about the cost of an entire category of work — and if it holds even approximately, it reprices a large part of the semiconductor industry's competitive map.
Notice which figure the industry reached for and which it skipped. Everyone repeated 92 percent, because 92 percent is a scoreboard and scoreboards are comfortable. But peak utilisation on a well-understood operation is a solved kind of problem; teams have been grinding toward those numbers for years, and getting there is an achievement of degree.
Ten hours is an achievement of kind. It says the work has moved from a capability you must possess to a resource you can rent — and the history of technology is largely a history of things making that particular crossing.
The claim is self-reported and the workload is narrow, and people who build chip-agnostic software for a living think the excitement is overdone. We give them page six rather than a parenthesis. Hold both readings: the demonstration is thinner than the headline, and the implication is larger than the demonstration.
Infinity was founded in 2025 by Jeremy Nixon, formerly of Google Brain, around an idea it calls automated invention: aim agents at a hard technical problem and let them iterate against reality rather than against a human reviewer. Its agent, Ignition, writes the low-level code that lets a model run on a chip, measures what it wrote, and rewrites whatever is slow.
Pointed at d-Matrix's Corsair accelerator, Infinity says Ignition reached as much as 92 percent of the chip's theoretical peak performance within ten hours — getting there by learning to distribute matrix-multiplication work across all thirty-two of Corsair's compute units. Keeping every unit busy is precisely the problem that separates a chip's datasheet from its delivered throughput, and it is normally where months disappear.
The ten-hour figure is the headline. The ten-day figure is the argument. Infinity says it had Qwen3, Qwen3.5 and Gemma 4 running fully on the chip inside ten days, and that after a single day of automated optimisation Ignition was delivering 34 percent more inference throughput on Qwen3-8B than vLLM — the open-source serving framework most of the industry treats as its floor.
That last comparison matters more than the peak-performance number. Theoretical peak is a physics question. Beating vLLM is an economics question, and it is the one buyers actually ask.
The seed round closed on 20 July: $15 million at a reported $100 million valuation, led by Touring Capital, with researchers from OpenAI and Anthropic participating personally. The people closest to frontier coding agents are funding the thing that turns coding agents on the hardware layer beneath them.
Infinity's own framing is not "a better CUDA." It is automated invention: point agents at a technical problem, let them iterate against the physical world rather than against a reviewer, and treat the resulting artefact as a by-product. Chip software happens to be an unusually clean test case, because the scoreboard is unambiguous and the silicon never flatters you.
CUDA's first advantage is the software. Its second is everything built on top of it: millions of lines of customer code, internal tooling, hiring pipelines and habits that make switching slow even when the alternative is free. Agents attack the first advantage directly. They do very little to the second, which is organisational rather than technical.
Nvidia, for its part, does not dispute the trend so much as claim it. The company says developers lean on CUDA's libraries more every year, and that its own engineers use coding agents to extend and validate CUDA faster. The tool is symmetric — and Nvidia has by far the largest corpus of CUDA for an agent to learn from.
Which is the part of the story that should give a challenger pause. Every argument that agents dissolve incumbency is also an argument that agents compound it, because incumbents own the training data. The asymmetry that decides this is not who has better agents. It is who has more examples of being wrong.
Training rewards peak performance; you buy the fastest thing available and eat the lock-in. Inference rewards cost per token, forever, at volume — and that flips the calculus toward software that runs anywhere. Marshall Choy of the Korean chip startup Rebellions puts it bluntly: on the inference side, CUDA "is no longer a factor." He calls it an open-source play.
The strategy only works if porting is cheap. Alibaba's T-Head has shipped an open-source CUDA alternative; optical, networking and edge-inference challengers are all raising against the same thesis. Agents are the mechanism that turns that thesis from a wish into a schedule.
The question a CTO asks a challenger vendor has been "do you have a software stack?" — a question with a yes/no answer that almost always came back no. The useful replacement is two-part and much harder to fake: how fast can you generate a stack for my workload, and how do you prove it is numerically correct at my scale?
Note what that does to negotiating leverage. Even if no buyer ever switches, a credible second source resets the price of the first — and credibility now costs ten hours of compute rather than a three-year engineering programme.
The practical move for a large buyer this quarter is unglamorous and cheap: ask a challenger vendor to port one real workload, and time them.
If this thesis is right, the observable signal will not be a challenger publishing a faster number. Numbers are marketing and always have been. The signal is a challenger publishing a schedule: a public commitment to have a customer's arbitrary model running on its silicon within days, with numerical parity stated as an SLA rather than a footnote.
If that becomes normal by year end, the moat has genuinely moved. If porting quietly stays a quarter-long project with a professional-services invoice attached, then what we saw this week was a very good demo of the easy 5 percent — and page six is the page to reread.
The question everyone asked this week was whether CUDA's moat is gone. It is the wrong question, and its answer — no, not yet — is worth nothing to anybody.
Here is the reframe. A moat's depth has only ever been measured in one unit: the human-years an adversary would need to reproduce it. Which means every moat in the economy has been quietly indexed to the price of engineering labour, and that price is now the fastest-falling input we have. It follows, uncomfortably, that the more of your advantage is made of accumulated code, the more of it has become a depreciating asset. You still pay to maintain it. Your rival has stopped paying to match it.
So here is the read that looks wrong until it doesn't. Nvidia's nineteen years of CUDA were never the asset. The asset was nineteen years of things going wrong.
Every silent numerical mismatch. Every kernel that was brilliant on one generation and catastrophic on the next. Every corruption that only appears above ten thousand devices, at hour forty of a run, on one batch size. None of that lives in the source code. It lives in regression suites, known-bad configuration lists, and the institutional memory of engineers who have been paged at 3 a.m. An agent can regenerate the code in an afternoon. It cannot regenerate the scars.
If that is right, the strategic instruction inverts. A challenger should stop spending to build a stack — agents will do that, and cheaply. It should spend to build evidence: the largest public corpus of "this kernel is known-correct at this scale." And Nvidia should stop selling CUDA as a body of software. It should sell it as a body of proof.
And for the reader who runs neither a fab nor a foundation model, the transfer is exact, and it is not comfortable. Your decade-old codebase is not the moat you have been telling the board it is. Your decade of post-mortems might be — assuming anyone wrote them down.
It is one founder's self-reported claim, on one chip, on one class of workload. Ninety-two percent of theoretical peak on matrix multiplication is the easiest region of a software stack, not a representative sample of it.
Chris Lattner, whose company Modular builds chip-agnostic AI software, calls the hype "very overblown." Writing code, he argues, is a small part of building software; tuning it for production is the hard part — and chip software is a niche field with few public examples for an agent to have learned from.
Bing Xu's line cuts both ways. If verification is the bottleneck, generating code faster mainly generates unverified code faster.
Switching costs are organisational. Amazon had unlimited engineers and still logged CUDA as a roadblock to its own chips.
And the tool is symmetric: Nvidia is pointing the same agents at the largest CUDA corpus in existence. Rivals must close the gap faster than an incumbent that is not asleep can open a new one.
Three moves from outside the compute stack — each one, on inspection, the same story about what an accumulated advantage is actually made of.
On 30 July, Google DeepMind released Gemini Robotics 2 — the first model in the family to drive legs, torso, arms and fingers under a single learned policy rather than bolting a locomotion controller onto a manipulation controller. It shipped in three forms: a vision-language-action model, an embodied-reasoning planner for multi-step tasks, and an on-device variant; it was demonstrated on Apptronik's Apollo 2 humanoid, tested with both Sharpa and Inspire five-fingered hands. Splitting walking from grasping was always an artefact of how humans organise engineering teams. A body that can brace, lean and step into a reach is recovering physics that the split threw away.
Valar Atomics announced a $1 billion Series B led by Sequoia on 3 August, roughly tripling its valuation to $6 billion months after its last round, alongside a $200 million credit facility. Sequoia's Shaun Maguire joins the board. The Hawthorne company's stated shift is from proving one reactor works to producing fleets of small ones, with AI data centres named as a principal target market. The repricing here is categorical, not technical: a reactor built as a construction project has no learning curve, and a reactor built as a manufactured unit has one.
On 30 July an FDA advisory committee voted 10–3 that efficacy data for Replimune's RP1, an oncolytic immunotherapy given with nivolumab, are evaluable and clinically meaningful in advanced melanoma that has progressed after anti-PD-1 therapy — ahead of a 2 August target action date. The application had already drawn two complete response letters, in July 2025 and April 2026. The IGNYTE data underneath: a 33.6 percent objective response rate, median duration of response of 24.8 months, and median overall survival of 32.9 months. Persistence is not a scientific virtue, but in regulatory science it is frequently the decisive one.
Nineteen years against ten hours is the ratio the industry noticed. Six million developers against both is the one it should have. Software can be regenerated overnight; a profession cannot.
Performance figures are vendor-reported and not independently reproduced.
Three stories in these pages, one shape. A software layer that took nineteen years, regenerated to 92 percent of a chip's ceiling in ten hours. A robot control problem that occupied two separate specialisms for two decades, collapsed into a single policy. A reactor reclassified from a construction project into a manufactured unit. In none of these cases did the physics change. What changed was the unit of accumulation.
For most of industrial history, advantage compounded through irreversibility. You built the thing; the building was hard; the difficulty of the building was the defence. Patents, tooling, codebases, tacit process — all of it rested on the quiet assumption that reproduction costs roughly what production cost. That assumption has been the load-bearing wall under a great deal of corporate strategy, and it is being removed while the roof is still on.
What replaces it is an asymmetry worth memorising. Anything whose difficulty lived in the assembly is getting cheaper — fast, and probably faster than your depreciation schedule assumes. Anything whose difficulty lives in knowing whether the assembly is correct is getting more expensive, because correctness is the one thing you cannot generate faster than you can check. Generation scales with compute. Verification scales with consequence.
This is not a story about chips, and it is not a story about Nvidia, which will very likely be fine. It is a story about where any organisation's durable value has migrated to while nobody was recataloguing. Your ten-year-old codebase is not a moat; a competitor's agent can approximate it over a long weekend. Your ten years of incident reports, failure modes, regression suites and near-misses might be the only thing that is.
So ask it before someone else does, and ask it about your own house rather than Nvidia's: which of our advantages would survive an adversary who could write everything we have written, but knew none of what we have learned?