Here is a lens worth carrying into any technology briefing. Ask not what got better — ask what got cheap, because the thing that gets cheap immediately promotes whatever sits next to it in the chain. Railways made freight cheap and made land valuable. Compilers made instructions cheap and made programmers valuable. Cloud made servers cheap and made data valuable. The winner is rarely the party that cheapened the step. It is the party holding the step that suddenly binds.
This week the step that got cheap was the scientific hypothesis. On Thursday, a Stanford team published a "Virtual Biotech" in Science: 37,075 software agents arranged into divisions under a chief-scientist agent, reading the results of more than 55,000 clinical trials and proposing targets. Three days earlier, a consortium of five drug companies reported that pooling 20,167 protein structures they had never published — under a scheme where none of them had to hand the data over — produced a folding model that beats every public one.
Read together, they say something sharper than "AI is coming for the lab." They say the generation of plausible ideas has collapsed in price, and therefore two things have become precious: the capacity to test an idea, and the record of tests that failed. Neither is a model. Both are physical, slow, and expensive. That is the inversion this issue is about — and it is where the returns in AI-for-science will accrue for the rest of the decade.
Somewhere in the design of the Virtual Biotech there is a decision that should unsettle anyone who runs an engineering organisation. James Zou's group at Stanford did not build a bigger model. They did not fine-tune one. They took an off-the-shelf frontier model — versions of Claude, though Zou is careful to say any capable model would do, including an open one you could run yourself — and then they spent their creativity on something else entirely. They drew an org chart. A chief scientific officer agent at the top. Divisions beneath it: target identification, trial design, and the rest. Thirty-seven thousand and seventy-five employees, each handed exactly one late-stage clinical trial to read. The intelligence was a commodity input. The structure was the contribution.
Drug discovery is not one hard task. It is several hundred loosely coupled ones that each require a different kind of looking, and the failure mode of a single powerful model asked to do all of them is not stupidity — it is averaging. Ask one system to simultaneously hold genomics, medicinal chemistry, regulatory precedent and trial statistics in mind and you get competent mush. Every judgement gets diluted by the others.
Division of labour is the oldest known fix for that, and it works on software for the same reason it works on people: it narrows what any one worker has to be right about. Give an agent one trial and one question and its answer is auditable. Aggregate thirty-seven thousand such answers and you get something no single pass over the same corpus produces — a distribution, with structure in it. Which is how this experiment found its most valuable result, and why almost every account of it has buried that result under the drug.
Three days before the Science paper, a quieter thing happened. AbbVie, Astex and three other drug companies — organised as the AI Structural Biology Network, coordinated by the federated-learning firm Apheris — took OpenFold3, the open replication of AlphaFold 3, and fine-tuned it on 20,167 protein-ligand structures that have never been public.
These are pictures of proteins gripping candidate drugs, made by X-ray crystallography and cryo-electron microscopy inside commercial programmes, most of which failed. The Protein Data Bank holds more than 200,000 structures. It holds perhaps 10,000 with a drug-like molecule bound — the exact regime that matters for medicine, and the exact regime where public data runs out. "The data that's missing from the PDB," AbbVie's John Karanicolas told Nature, "is exactly the data that's present in our internal data." The pooled model beat public OpenFold3 by a wide margin. It also beat each company's own privately fine-tuned model. Pooling, not privacy, was the edge.
Zou's agents were pointed at the graveyard: the published results of more than 55,000 trials across a wide range of conditions. What came back was not a compound. It was a rule about which compounds are worth trying. Drugs aimed at proteins that are active in specific cell types — rather than expressed broadly across the body — were found to be nearly 50% more likely to reach market than the rest.
Sit with the economics of that for a moment. A phase-III failure costs a pharmaceutical company hundreds of millions of dollars and years of option value. A prior that shifts the odds by half, applied at portfolio scale before any money is committed, is worth more than any single successful asset. The agents did not discover a drug. They discovered a filter.
The headline result is more modest than the headlines. Given a human tip-off that CD276 — a protein that dampens immune response and is heavily expressed in lung tumours — might be a target, the system confirmed it against existing data and designed a strategy: an antibody that recognises CD276, tethered to an anticancer payload. An antibody-drug conjugate. External reviewers called the avenue promising.
Note the shape of the win. The hypothesis was seeded by people. The validation was retrospective. Nothing was synthesised. Nothing was dosed. What the system supplied was the connective work — the weeks of literature, data and design reasoning that normally sit between a hunch and a proposal.
The structural-biology result is the more commercially radical of the two, and it hinges on a mechanism rather than a model. Five competitors contributed training signal without any of them shipping their crystallography to a rival or to a cloud they don't control. The structures were supplied such that proprietary data stayed private; the gradients travelled instead.
On a held-out benchmark of 1,056 protein-ligand structures, the pooled model predicted more than half to high accuracy. Public OpenFold3 managed a third. Boltz-2, the strongest open competitor, reached roughly 40%. The consortium has not published a peer-reviewed paper; the result lives in a blog post; the model is not available to anyone outside. All three caveats matter, and none of them changes the direction of the arrow.
“You add all this data, and you get a pretty big bump in performance.” Mohammed AlQuraishi, computational biologist, Columbia University
A result published in Nature Structural & Molecular Biology earlier this year explains why the vaults were worth raiding at all: the accuracy of co-folding models does not degrade gracefully. It falls off a cliff the moment they are asked about molecules substantially unlike anything in training. For a scientist exploring known chemistry that is tolerable. For a medicinal chemist whose entire job is novelty, it is disqualifying.
Governments have noticed. OpenBind, backed by up to £8 million (US$10.8 million) of UK funding, released hundreds of new structures last month with thousands more in train — an attempt to rebuild in the open what the consortium assembled behind glass.
If your AI-for-science roadmap is a list of models, it is priced for the last regime. Agent swarms make candidate hypotheses close to free; the queue that forms is at the assay, the crystallography beamline, the animal facility, the recruiting site. Capital should follow the queue. In practice that means robotic wet-lab throughput, standardised assay protocols, and the unglamorous data engineering that makes a failed run machine-readable rather than a PDF in a drawer.
The consortium's result is a valuation event for something almost nobody carries on a balance sheet: records of experiments that did not work. Those records are what the public corpus structurally lacks, because nothing incentivises anyone to publish them. Whatever your field, the question to ask this quarter is which of your failures are legible, and to whom they would be worth something.
The most transferable finding of the week is that pooled training beat siloed training even for the firms that owned the silos. That makes federated architecture a competitive instrument rather than a compliance technology. The organisation that can credibly hold the middle — writing the governance, running the enclave, splitting the upside — captures value without owning a single datum. This is a legal and diplomatic capability being sold as an infrastructure one.
The pattern to watch next: whether trial sponsors begin federating clinical data the way these five federated structural data. That is a far larger vault, and the same argument applies to it.
A shared starting model is copied to each company. Nothing has left the building yet — the model travels to the data, which is the reversal that makes all of this possible.
Each copy trains locally on structures that never move. What it sends back is not the structures but the adjustments it wants to make to the weights — a summary of what it learned, not what it saw.
The coordinator averages those adjustments into one improved model and sends it round again. After enough rounds every participant holds a model that has effectively read all five vaults.
Predicting the shape a protein takes while gripping another molecule — the version of the problem that matters for drugs, and much harder than folding alone.
The small molecule that binds a protein. Most drugs are ligands; most ligands are not drugs.
Antibody-drug conjugate. A guided missile: an antibody for targeting, a toxic payload for effect.
Many model instances given narrow roles and a way to report to each other, so that the aggregate is more auditable than one model doing everything.
Here is the move that looks wrong. If agent swarms can generate credible drug hypotheses at near-zero marginal cost, the rational response is not to buy more inference. It is to short intelligence and go long on verification — to spend the next three years building the most boring thing in the building: throughput at the bench.
The logic is a constraint argument, not an AI argument. A system that produces a thousand plausible hypotheses per day is worth nothing to an organisation that can test four per week. Every additional hypothesis past the fourth is inventory, and inventory with a shelf life, because a hypothesis you cannot check is indistinguishable from a hypothesis you invented. The binding resource stopped being cognition some time ago. This week we got the evidence.
The second half of the trade is weirder. If verification is the bottleneck, then the most valuable data in the world is the record of verifications that came back negative — precisely the data no journal wants and no company publishes. And the way to monetise a hoard of negative results is not to hoard it harder. It is to convene: to let four competitors train against your failures in exchange for training against theirs. Five firms did exactly that this week and every one of them ended up with a better model than its own private version. Cooperation strictly dominated secrecy, among parties who would happily litigate each other on a Tuesday.
So the contrarian position, stated plainly: the scarce asset in AI-for-science is not a model, it is a robot that can be wrong cheaply — and the second scarcest is a lawyer who can get rivals to share their mistakes.
And there are good reasons it might. Start with the evidence quality. The Virtual Biotech's outputs were never experimentally validated — no synthesis, no assay, no dosing. Its lung-cancer target came with a human tip-off attached. The structural-biology result is a non-peer-reviewed blog post describing a model nobody outside five companies can inspect, benchmarked on structures drawn from those same five companies, which is exactly the condition under which a held-out set flatters a model.
Then the architecture. Thirty-seven thousand agents sounds like collective intelligence; mostly it is embarrassing parallelism — one trial per worker, then aggregation. That is a map-reduce with a language model in the mapper. It is genuinely useful and it is not a new kind of mind, and anyone extrapolating from the headcount to emergent reasoning is extrapolating from a billing metric.
Federation has its own honest problems: gradients leak information, membership inference is a live attack, and the incentive to contribute your best data collapses the moment the pooled model becomes a product. The cooperative equilibrium held for five firms and one round. It is not obvious it holds for fifty.
And the strongest counter is simply that compute has beaten cleverness for a decade, and betting against it has been a reliable way to look foolish. Perhaps a model good enough in 2029 makes wet validation a formality rather than a gate. That would invert this column completely — and it is the specific claim worth watching, because everything here depends on it being false.
OpenAI claims its most advanced model has settled a Millennium Prize Problem — one of the handful of open questions carrying a US$1 million purse — by proving that the Navier-Stokes equations can produce a singularity, a physically impossible infinite fluid speed. The satisfying part is the physics. Brown University's George Karniadakis calculates that in air the blow-up arrives when a stretching vortex thins to roughly 70 nanometres: about the distance one air molecule travels before hitting another. The equations assume fluid is continuous. At 70 nm that assumption simply stops being true, so of course they break. The proof did not reveal a flaw in physics; it located the edge of a model. Meanwhile the credit fight it started — over whether the system absorbed unpublished human work — drew a Nature editorial within two days.
Physicists have driven the famously delicate low-energy transition in thorium-229 using a continuous-wave, narrow-bandwidth solid-state laser at 148 nanometres — at sub-nanowatt power — and read the resonance out in absorption rather than waiting for the nucleus to fluoresce. Both changes matter more than they sound. Pulsed lasers wasted almost every photon on frequencies the nucleus ignored; fluorescence detection forced you to wait for a slow decay. Continuous excitation plus absorption readout means fast signal acquisition from a millimetre-sized room-temperature crystal. A clock referenced to a nucleus rather than an electron shell is far better shielded from its environment — which is why this is a navigation, geodesy and fundamental-constants story, not only a timekeeping one.
An analysis in Science of Chinese patent filings and corporate research from 2010 to 2022 finds that firms placed on the US Entity List responded by citing scientific literature in their patents 72% more often, and by participating more in producing that literature themselves. The mechanism is uncomfortable and obvious in hindsight: cut a company off from licensed technology and you have not removed its demand for the knowledge — you have redirected it from the commercial channel to the open one. Controls designed as a brake on capability functioned as a subsidy for indigenous science. For anyone modelling the compute supply chain, this is the coefficient to revise.
AI agents deployed inside the Virtual Biotech, one per late-stage clinical trial, all reporting to a single chief-scientist agent.
Published clinical trials read, across a wide range of conditions, to hunt for predictors of eventual market approval.
Higher likelihood of reaching market for drugs targeting proteins active in specific cell types. The week's most valuable sentence.
Proprietary protein-ligand structures from five drug companies, pooled without any of them handing over the files.
Structures in the entire public Protein Data Bank that include a drug-like molecule — out of more than 200,000 total. The shortage in one figure.
Held-out structures in the benchmark. Pooled model: over half predicted accurately. Public OpenFold3: one third. Boltz-2: about 40%.
UK government funding (up to £8 million) behind OpenBind, which is trying to build the missing public dataset in the open.
Water molecules in the record 2024 supercomputer simulation — enough for a cube a few micrometres across. Why brute force will not replace fluid equations.
The Millennium Prize purse on the Navier-Stokes problem — a rounding error beside the compute bill, and the cheapest publicity any lab bought this year.
For four hundred years the rate limit on science was imagination. Finding the right question was the achievement; Newton's genius was not arithmetic. Every institution we built — the professorship, the grant, the journal, the patent — is an apparatus for identifying and rewarding the person who had the idea. The whole edifice assumes ideas are scarce.
This week, in three separate places, that assumption cracked. Thirty-seven thousand agents produced a portfolio-level insight in a domain where such insights take careers. Five rivals discovered that their most valuable shared asset was a pile of things that hadn't worked. And a model returned a proof of a problem that had defeated mathematicians for a century — immediately triggering a fight over who deserved credit, which is precisely the fight an institution has when its core scarcity assumption stops holding.
If ideas become abundant, then everything downstream of an idea becomes the constraint, and everything upstream becomes a commodity. The prestige gradient inverts. The person who can run the decisive experiment in a week outranks the person who can think of it. Reproduction — long the least glamorous activity in science, the thing nobody funds — becomes the highest-leverage function in the building.
We have been here before in a smaller way. When compilers made instructions cheap, the industry did not conclude that programming was over; it discovered testing, and then spent thirty years learning that the test suite was the real asset. The analogue for science is a wet lab you can address like an API, and a culture that treats a negative result as revenue rather than embarrassment.
Which suggests the question to carry forward. Not how smart will the models get — that one has a trend line and a lot of people watching it. The better question is whether we will build institutions that can absorb abundant hypotheses at all, or whether we will spend the decade generating a thousand excellent ideas a day into an organisation that can still only check four.
Callaway, E. How a team of AIs discovered a promising lung-cancer drug. Nature, 17 Sep 2026.
nature.com/articles/d41586-026-02954-y
Zhang, H. G., Eckmann, P., Miao, J., Mahon, A. B. & Zou, J. Science (2026).
doi.org/10.1126/science.aeg6779
Callaway, E. Drug firms' secret data supercharge AI protein models. Nature, 14 Sep 2026.
nature.com/articles/d41586-026-02882-x
Škrinjar, P. et al. Evaluating generalization in protein–ligand cofolding methods.
Nat. Struct. Mol. Biol. 33, 782–794 (2026).
doi.org/10.1038/s41594-026-01797-5
Jones, N. AI cracked the Navier–Stokes challenge. What does that mean for physics?
Nature, 18 Sep 2026.
nature.com/articles/d41586-026-02922-6
Continuous-wave laser absorption spectroscopy of the thorium-229 nucleus.
Nature, 16 Sep 2026.
nature.com/articles/s41586-026-11011-7
Chinese companies doubled down on science after US tech restrictions.
Nature, 18 Sep 2026 · Science, doi 10.1126/science.aed7991.
nature.com/articles/d41586-026-02946-y
AI companies must work with the research community to protect attribution.
Nature editorial, 16 Sep 2026.
nature.com/articles/d41586-026-02886-7
Who gets credit in the AI era? OpenAI maths bombshell sparks debate.
Nature, 17 Sep 2026.
nature.com/articles/d41586-026-02910-w