← The stand
Cover of The Collection, Volume 1, Number 20: The Review. Wednesday 9 September 2026, Melbourne. Skip cover

Vol. 1  ·  No. 20  ·  Wednesday 9 September 2026  ·  Melbourne


The Collection

The Review

Collected and edited by Newsletter World for AK

Contents

A letter, four pieces, standing orders, and a colophon. Monday and Tuesday stayed empty on the rack. The mail was thick enough for a real paper. Today asks who still reads the code.

  1. iiiEditor’s LetterAgents write the pull request. Bots review the bots.03
  2. ivThe AvalanchePRs up fivefold. Humans grade the AI review.04
  3. vThe StampMath may be real. The stamp war is the story.05
  4. viThe StudentThe teacher is expensive. Safety does not transfer cleanly.06
  5. viiThe GradeRAG stores facts. Fine-tuning stores behaviour.07
  6. viiiStanding OrdersFour rules for this issue.08
  7. ixColophonThe letters, named.09

03  ·  Editor’s Letter

Agents write the pull request. Bots review the bots.

Sunday’s Hire stays on this week’s rack. We do not reprint it. Monday and Tuesday went unsold; the morning run did not fire. Overnight the letters filled again. The Pragmatic Engineer asked CTOs what to do when agents open more pull requests than any team can read. Gary Marcus counted nine OpenAI misconduct reports in a week, including an AGI-era blitz that bullish users still do not feel. The Algorithmic Bridge watched labs fight over a Millennium Prize stamp while kids’ PISA scores keep falling. An AI governance letter explained distillation: the student copies the teacher, and the guardrails often do not come along. Cobus Greyling put humans back in the preference loop and said it plainly: RAG stores facts; fine-tuning stores behaviour.

Agents write the pull request. Bots review the bots. The engineer grades the review.

The code is cheaper than attention. Attention is the scarce desk.

The Hire asked who you put on the payroll. The Review asks who still opens the diff.

The Collection · The Review03

04  ·  The Desk

The Avalanche

GitHub PRs up fivefold in three years. Growth nearly doubled again from late 2025. Nobody reviews every line by hand.

Since late 2025, Gergely writes, the era of developers writing most code by hand looks over at startups and in Big Tech. Agents open more pull requests, and the diffs are larger. GitHub’s own chart: PRs opened up about fivefold over three years, with a sharp second climb after late 2025. The question haunting CTOs is not whether review matters. It is what review even is when the pile never clears.

The most common pattern: AI review tools comment first; humans review the review. Bun runs CodeRabbit, GitHub Code Review, and Claude Code Review on the same PRs. Weaviate’s CTO likes an adversarial agent, a human scope call, an agent fix, then a human exit. Noise still kills adoption; WeTravel stayed off AI review after a noisy trial, even after a June retest improved. Uber built uReview to grade and cull low-confidence bot comments before a human sees them. OpenAI and Anthropic triage by blast radius: low-risk can ship on AI alone; high-risk still needs a person. Anthropic’s Jarred Sumner says a human still merges low-risk today, with a goal of another Claude doing that later. Duckbill Group, five people, hit sixty open PRs, then switched to risk labels (public API, auth, schema, agent skills) plus harder tests. Merges nearly doubled; no-human median merge time fell to about an hour. Some teams now review the plan, the tests, or the database schema, and skip the implementation. Others just try to make agents ship smaller PRs. Talk of dropping human review is louder than evidence of shops that already did.

Ninety percent agents. Humans keep the exit criteria.

The avalanche is not optional. The desk that still reads every line is already behind.

The Collection · The Review04

05  ·  The Floor

The Stamp

A Millennium Prize path, AGI marketing, and a week of alleged coverups. The name on the paper is the fight.

Marcus lists nine fresh OpenAI stories in about a week: a German site hack besides Hugging Face; reports they knew weeks earlier and did not disclose; coy testimony to Congress; Brockman’s Astra-as-AGI campaign that bullish users and Artificial Analysis both refuse to crown; ARC-AGI-3 at 99.9% with an in-house harness ARC themselves could not match out of the box; and a fight with mathematicians on a Millennium Prize path that, if the account holds, smells like credit pressure bordering on extortion. OpenAI says a fuller response is coming. Marcus wants Altman and Brockman out until the top changes. He also notes at least sixteen senior exits since January despite the IPO clock.

Alberto Romero’s Algorithmic Bridge letter lands on the same bruise from another angle. There appears to have been real AI-augmented progress toward Navier-Stokes. The public story became corporate drama: who stamps the result. Romero’s claim is sharp: when stakes rise, the top labs drop cooperative coordination. Altman and Amodei treat each other’s gains as losses. Agents learn hive deference; their builders revert to tribe. In the same week, PISA math, reading, and science scores keep sliding for about fifteen years. Phones, COVID, now generative AI each take the blame in turn. Models climb the highest math. New humans fall out of the lowest.

The equations may hold. The stamp war is the news.

Review the claim. Label the response unread until it arrives. Do not confuse a scoreboard with a verdict.

The Collection · The Review05

06  ·  The Sheet

The Student

Distill the teacher into something cheap. Retest the student as if it were new.

Anthropic accused Moonshot of industrial-scale distillation against Claude: millions of exchanges through fraudulent accounts, pulling agentic reasoning, tool use, and coding into rival models. Similar charges have hit DeepSeek, MiniMax, and others. The same technique, the letter insists, is ordinary inside the labs themselves: banks, pharma, retailers, and governments distill proprietary or third-party models for cost, latency, and compliance.

Teacher to student. The student mostly sees outputs, not weights or original training data. Capabilities transfer unevenly. Safety and refusal often degrade. Fairness can amplify under output mimicry. A cheaper model widens the attack surface because jailbreaks get cheaper to run at scale. Governance advice in the free strip: do not assume Haiku or Flash inherits the teacher’s safety sheet; run independent audits; treat distillation as dual-use. The paid playbook on when to greenlight sits unread behind the wall. We do not invent it.

Cheaper to run is not safer by inheritance.

Review the student on its own. The teacher’s report card does not travel.

The Collection · The Review06

07  ·  The Bench

The Grade

Generate, grade, train, generate again. Preference data is the behaviour you keep.

Greyling splits the stack cleanly. Context and RAG store facts for inference. Fine-tuning stores behaviour. Distillation with a teacher model is one path: prompt the large model, keep the pairs you want, train the small one on that slice. Preference tuning is another: generate two replies, have a human or a consistent rubric pick a winner, curate chosen and rejected pairs, train with DPO or GRPO, then generate again from the new policy.

He runs the loop agentically on Apple silicon with small Qwen models, MLX-LM and LoRA, plus a simple grading UI. Grok Build writes the scripts in plain language. Honest limit: full RL on a Mac is slow; iterate small, then move scale to a cheap GPU. A few hundred tight preference pairs can move narrow behaviour. Broad capability still wants more data. Fifty great pairs beat five hundred noisy ones. Inconsistent graders bake noise into the policy.

Facts in the retrieve. Behaviour in the grade.

The review desk is not only for code. It is how the model learns what good looks like.

The Collection · The Review07

08  ·  Standing Orders

Four rules for this issue

  1. I

    Review the review, then keep the exit.

    Let agents comment first. Keep a human on scope and exit criteria. If the bot noise is louder than the signal, fix the filter before you trust the stamp.

  2. II

    Triage by blast radius.

    Auth, public APIs, schema, and agent skills still need a person. Low-risk paths can lean on AI if the guardrails and tests are real. Do not pretend every diff is equal.

  3. III

    Retest the student.

    Distilled models are not safer by inheritance. Safety, refusal, and fairness often fail to transfer. Audit the cheap model as if it were new.

  4. IV

    Grade behaviour on purpose.

    RAG holds facts. Preference pairs hold behaviour. Inconsistent graders train noise. Label unread paywalls instead of inventing the missing playbook.

The Collection · The Review08