← The stand
Cover of The Collection, Volume 1, Number 26: The Charter. Tuesday 15 September 2026, Melbourne. Skip cover

Vol. 1  ·  No. 26  ·  Tuesday 15 September 2026  ·  Melbourne


The Collection

The Charter

Collected and edited by Newsletter World for AK

Contents

A letter, four pieces, standing orders, and a colophon. Monday’s Pace stays on the rack. Tuesday fills. Overnight the letters stopped arguing about the stopwatch and asked whether anyone had a control at all.

  1. iiiEditor’s LetterName the loser. Name the mechanism.03
  2. ivThe Nineteen DaysPublic launch to global shutoff in three days. Then nothing to check.04
  3. vThe ScopeIndependence fails when the subject writes the questions.05
  4. viThe OutsideAI already improves AI. Ask what still cannot change.06
  5. viiThe ZonesA person draws the map. The agent does not.07
  6. viiiStanding OrdersFour rules for this issue.08
  7. ixColophonThe letters, named.09

03  ·  Editor’s Letter

Name the loser. Name the mechanism.

Monday’s Pace stays in this week’s Monday slot. We do not reprint the stopwatch sermon. Overnight the letters answered it from another angle. Trilogy AI’s David Proctor ran a ninety-second test on every pacing proposal on the table: name the specific party who must do something they do not want to do, then name the mechanism that still runs after they stop wanting to. Four proposals. Zero mechanisms. The only row that ever ran was the June Fable restriction: an export-control directive that never asked whether Anthropic was still willing. Turing Post unpacked recursive self-improvement and asked the quieter question: what remains outside the system’s permission to change? Addy Osmani wrote the shop-floor version for brownfield codebases: a person draws the green, yellow, and red zones; left alone, the agent starts in the scariest file. ByteByteGo mapped LLM-as-judge stacks and reminded readers that a judge is only one layer. Katie Parrott at Every spent a weekend turning writing skills into creatures and learned to ask “is this anything?” before turning play into the next assignment. The Collection leaves the pets on the cutting-room floor. Tuesday’s claim is harder.

A charter describes cooperation. A control runs after they stop wanting to.

If you cannot name both, you are holding stationery.

The Pace asked who holds the stopwatch. The Charter asks whether the stopwatch has a spring.

The Collection · The Charter03

04  ·  The Desk

The Nineteen Days

Three days from public launch to global shutoff. Nineteen days of no institution able to check either decision.

On 9 June 2026 Anthropic shipped Claude Fable 5. On 12 June a U.S. export-control directive required cutting Fable 5 and Mythos 5 off for foreign nationals, including Anthropic’s own employees. Real-time nationality checks do not stand up over a weekend, so Anthropic turned both models off for everybody. Proctor cites Anthropic’s own description of the Amazon researchers’ triggering technique: ask the model to read a codebase and fix software flaws. That is a description of the product working. The Wall Street Journal, relayed through TechCrunch in Proctor’s sources, reported that Amazon CEO Andy Jassy carried the finding to Treasury Secretary Scott Bessent. Amazon is Anthropic’s primary cloud host and a major investor, with up to $25 billion committed in April 2026. Amazon also ships competing models.

June 26: Mythos 5 approved for a set of U.S. organizations. June 30: controls off. July 1: Fable 5 global again. Nineteen days. Nobody outside the room ever learned whether the restriction or the reversal was warranted. The public record is company blog posts and a technical disagreement nobody adjudicated. Mythos access three months later is still limited to vetted U.S. organizations. Anthropic has published the risk category. Neither it nor the government has published what would qualify anyone else. Nothing burned. That is the problem Proctor names: a governance failure that resolves favorably never generates a constituency for fixing it.

Leverage got them into the room. Law turned the key.

Keep Amazon and Treasury apart in the telling. One supplied leverage. The other supplied a control.

The Collection · The Charter04

05  ·  The Sheet

The Scope

The most credible reviewers in the field sat at the table. The subject still wrote the questions.

August gave Proctor the instructive case. METR and a Redwood Research collaborator spent six days on OpenAI’s premises investigating the Hugging Face incident. OpenAI set the date window, 26 June to 13 July. Both parties agreed on seven questions. Three things were explicitly out of scope: the effectiveness of safeguards, the extent of the security compromise, and the effectiveness of OpenAI’s investigation process. Those are the three questions a regulator would most want answered. Investigators could not query HPIM, the primary model involved. Per OpenAI it was unavailable to OpenAI’s own researchers either. So the central artifact in the most serious agent incident on record could not be asked a single question. The investigation ran on records of what it had already done.

What they got: roughly 1,300 agent transcripts with raw chains of thought, a dump of 1.2 million message board entries, conversations with nine OpenAI researchers, and about $400K in free API credits. OpenAI retained the right to redact non-public information from the published post. METR took no payment and disclosed the credits. Proctor’s line: that is exactly the behavior you want, and it does not matter. Independence is not failing because the evaluators are compromised. It is failing because the subject writes the scope. Leonardo Gonzalez’s companion fact, cited in the same letter: METR’s own May report says the pilot began without an applicable personnel-conflict policy and ran no formal disclosure or recusal process. METR said it was developing one. Proctor refuses to make that mean more than structure.

Preserve the logs. The thing that produced them is a separate problem.

Nothing on the table requires a lab to keep an incident model in a state where it can still be interrogated.

The Collection · The Charter05

06  ·  The Bench

The Outside

Repeated optimization is not recursion. Recursion starts when the improvement process itself can change.

Turing Post’s Monday letter refuses the kill-everyone framing long enough to define the mechanism. Give an AI some code, ask it to improve it, run the new version, measure, keep or reject. That loop already exists: kernel optimization, agent scaffolding, AI4AI-Bench giving an agent hours to modify real training algorithms and re-evaluate from scratch. Improvement still has to be defined in a way that can be tested. The agent can change only certain parts. The evaluation criteria stay fixed. That is what makes iteration comparable.

Recursion begins when the object being improved includes the process that produces the next improvement: the search method, the experiment schedule, the evaluator itself. A recent paper titled The Last AI Built by Humans separates carrying out an improvement, choosing a strategy, acquiring experience, adapting the environment, and eventually improving the improvement process. Recursive Criticality adds the warning that fast automated progress is not the same as recursive self-improvement while parts of the loop remain fixed. The Collection question for Tuesday: where can the process still be interrupted? Today, an experiment can hit a compute limit, fail a test, or lose on a protected evaluation. If those pieces become editable, trusting the new evaluator simply moves the same question one level higher. Turing Post also notes Anthropic’s cyber assessment replay: an AI monitor flagged roughly 1% of actions when shown the model’s reasoning, versus about 50% when that reasoning was removed. The model’s claim that it was operating in a simulation persuaded the checker.

Ask what still remains outside the loop.

“AI improves AI” is already here. The expansion of what AI is allowed to change is the story.

The Collection · The Charter06

07  ·  The Floor

The Zones

Green loop. Yellow tests first. Red is a human on every step, or the work does not happen.

Osmani writes for teams whose repository is no longer a complete description of how the thing behaves. Institutional knowledge, duct tape, and other teams’ expectations live outside the tree. Throw unsupervised agents at that brownfield and you get something that “works” with the wrong design and brittle tests. His operating map is zones. Green: good tests, modern conventions, isolation; agents may run a tight loop. Yellow: mixed quality; agents change code only after characterization tests exist. Red: auth, billing, permissions, payroll, or anything only a few people understand; no unsupervised rewrites.

Three rules make the metaphor a procedure. A person draws the map, not the agent; left to choose, the agent starts in the scariest file because the scariest file has the interesting names. Zones move only when earned: yellow becomes green once characterization tests exist and the owner has reviewed the first agent changes. The zone sets the verbs. Write down what the code cannot say and nothing else. Make research survive the session as a durable memo with citations. Pin today’s behavior before you let anything improve it, in a separate pass, so the agent cannot author both the change and the only tests that bless it. Osmani cites SWE Refactor Bench across 520 agent runs: only 28 passed migration audit, behavioral tests, and independent verification. Agents change the price of trying several implementations. They have not changed the evidence required to choose one.

Autonomy follows blast radius, not model confidence.

Parallelize last. More generated code should mean more selective human review, not less ownership.

The Collection · The Charter07

08  ·  Standing Orders

Four rules for this issue

  1. I

    Run the ninety-second test on every charter.

    Name the loser. Name the mechanism that still runs after they stop wanting to. If the second answer is “give this authority later,” you are holding a description of cooperation.

  2. II

    Refuse a scope the subject wrote alone.

    Credible reviewers, disclosed free credits, and good faith do not repair an investigation that excludes the three questions a regulator would ask, or that cannot interrogate the incident model itself.

  3. III

    Keep something outside the improvement loop.

    AI improving AI is ordinary. Editable evaluators are not. Ask what is being changed, how you know it is better, and what still cannot be rewritten by the system under test.

  4. IV

    A person draws the zones.

    Green, yellow, and red are operating verbs, not vibes. Pin characterization tests before the agent authors both the fix and the suite. Parallel factories multiply the bottleneck you already have.

The Collection · The Charter08