← The stand
Cover of The Collection, Volume 1, Number 21: The Harness. Thursday 10 September 2026, Melbourne. Skip cover

Vol. 1  ·  No. 21  ·  Thursday 10 September 2026  ·  Melbourne


The Collection

The Harness

Collected and edited by Newsletter World for AK

Contents

A letter, four pieces, standing orders, and a colophon. Wednesday’s Review stays on this week’s rack. The mail was thick enough for a real paper. Today asks what sits around the model.

  1. iiiEditor’s LetterThe harness runs one step ahead of the model.03
  2. ivThe CrutchesCodex keeps the harness ahead. Then it sheds.04
  3. vThe SpendStop token maxing. Start harness maxing.05
  4. viThe ProbeWriting a trap and proving it is fair are different skills.06
  5. viiThe ForestWork in progress goes dark. Wallfacers only.07
  6. viiiStanding OrdersFour rules for this issue.08
  7. ixColophonThe letters, named.09

03  ·  Editor’s Letter

The harness runs one step ahead of the model.

Wednesday’s Review stays on this week’s rack. We do not reprint it. Overnight the letters shifted from who grades the diff to what wraps the model. The Pragmatic Engineer sat with Tibo Sottiaux of Codex: the harness is always slightly ahead; the crutches get discarded when the weights catch up. Cobus Greyling called the opposite habit by name (token maxing) and pointed at a study where the same six models did the same twenty-two tasks for fewer tokens, less money, and less time because the software around them changed. State of AI ran LivingArena: ten frontier models writing traps for each other, and the best trap-writers were not the best answerers. Erik Hoel said culture has become a dark forest; unfinished work gets hunted.

The harness runs one step ahead of the model. When the model improves, they throw the crutches away.

The weights are not the product. The wrap is.

The Review asked who still opens the diff. The Harness asks what the model is leaning on.

The Collection · The Harness03

04  ·  The Desk

The Crutches

Rust for scale. Open source so the fork can change the model line. The harness shrinks as the weights grow.

Tibo Sottiaux helped build Codex and now runs Core Products & Platform at OpenAI, the org that still includes it. On Orosz’s podcast he walks the build: Codex CLI in Rust even when models were stronger at Python and TypeScript, because the vision was millions of cloud machines and a rewrite later would cost more than the early friction. Open source, he says, buys trust and contributors; the sting is watching work land in competing tools first. That same openness is why Codex can run models that are not OpenAI’s. Lock the harness to one vendor and anyone can still fork a few lines. Claude Code, by contrast, stays closed and Anthropic-only.

The line that belongs on this cover: the Codex harness stays slightly ahead of OpenAI’s latest model. Guardrails, safety, efficiency, steerability, the developer message injected at the start of each turn: crutches. As models improve, some crutches get discarded and the harness shrinks. That has been the cycle. Inside OpenAI, Codex is plugged into Slack, documents, and code by default; new joiners are told to ask it before they ask a person. Correctness checks and security review are headed the same way. Maintenance that used to take months can take hours if the abstractions and tests are real. Tibo’s personal shift is blunt: being “in the zone” is history; he still opens an editor because it feels nice, then fires an agent for the data and decides.

Harness first. Then shed.

Do not confuse the model card with the product. The wrap is doing the work you notice.

The Collection · The Harness04

05  ·  The Floor

The Spend

Same models. Same panel. Better wrap. Fewer tokens, less money, less time.

Greyling names the default habit: token maxing. Buy capability with longer traces, more agent runs, more parallel tool calls, until spend compounds and the value of each token falls. Frontier labs, he notes, have been harness-maxing behind the API for years. A reverse-engineering study claimed most of Claude Code is not the model. OpenAI’s deep research path runs an augmented harness you do not see. The new paper he cites held twenty-two tasks and one judging panel across six foundation models. Change the software around them, not the weights, and the work completed with 38% fewer tokens, 41% less money, and 44% less time.

Quality per dollar rose about 82%. Task completions per million tokens rose from 54.9 to 92.0. The harness also added a net-new capability: delegated sub-agents. Efficiency improved across every model. Quality did not. Stronger models convert harness structure into better answers; weaker ones can choke on the extra orchestration. Greyling calls that harness leverage. The cost lever is large: jumping from the most expensive model to the cheapest under a baseline saved about 36%; keeping any model and adopting the harness saved about 33% to 61%. HarnessBench keeps reminding the same lesson: a better wrap is not a free upgrade for every weight file.

Buy less model. Build more wrap.

If the bill is rising and the answers are not, look at the harness before you buy another frontier seat.

The Collection · The Harness05

06  ·  The Sheet

The Probe

Three thousand six hundred rounds. Ten models writing exams for each other. The skills split.

LivingArena skips the frozen question bank. Ten frontier models take turns as questioner and answerer. The questioner studies the opponent’s history, writes a targeted trap with a reference answer and a verification procedure, and submits it to a panel of non-participating judges before the opponent ever sees it. Only validated questions count. Forty-five pairs, eight directional matches each, ten rounds: 3,600 rounds. Contestants include GPT-5.5 and 5.2, Gemini-3.5-Flash and 3.1-Pro, Claude Opus 4.7 and 4.6, DeepSeek-V4 previews. Judges: Kimi-K2.6, MiniMax-M2.7, GLM-5.2.

Scoring is asymmetric. A rejected question costs the questioner. A validated trap the answerer fails earns the questioner a point, with discounts after the first hits in a domain so one soft spot stops paying forever. The paper’s finding is the Collection line: finding a real weakness and proving the test is valid do not travel together in the same model. Fixed benchmarks expire the day they stop distinguishing anyone. LivingArena asks whether a model can see another model’s blind spot and still write a fair exam.

Best trap-writer is not best student.

A harness that only flatters its own model is not a probe. It is a mirror.

The Collection · The Harness06

07  ·  The Bench

The Forest

Share the draft and the hunters arrive. Wallfacing is the new default.

Hoel borrows Liu Cixin’s dark forest: reveal yourself and something shoots. He says scientists, writers, and mathematicians now live there because of AI. His case study is unfinished math. Two researchers, one at Anthropic, one at NYU, had been pushing a Navier-Stokes program for about a year with models as tools, not as autonomous solvers. Their starting ideas, they insist, came from earlier open work by Córdoba and Martínez-Zoroa. Then the credit race went loud. Buckmaster’s statement says they were first told “very little human input” went into a rival claim; later it emerged a whole team had been on it. He asked whether their Codex sessions had trained or informed the model. He was told user data was not looked up. On training, he says he got no answer.

Terence Tao’s worry, as Hoel reads him: automated proofs can be verifiable and still “odorless,” stripped of the insight humans used to trade in public. Open problems become a non-renewable resource when solving them no longer seeds the next generation of ideas. Conferences, drafts, floating a novel plot online, those commons assumed effort had a floor. Now a prompt can mine the almost-as-good version. Hoel’s prescription is wallfacing: finish in silence, leak nothing. Slash-and-burn culture instead of crop rotation. Yesterday’s Review covered the stamp fight. Today’s letter is the quieter damage: the draft that never leaves the tree.

The commons closes. The forest goes dark.

Gary Marcus, same morning, argued for a boycott until incorrigible systems get a real plan. Low p(doom), high p(dystopia). We note it; we do not turn this paper into a manifesto.

The Collection · The Harness07

08  ·  Standing Orders

Four rules for this issue

  1. I

    Budget the wrap, not only the weights.

    If spend climbs and quality does not, change the harness before you buy another frontier seat. Token maxing is a habit, not a strategy.

  2. II

    Keep the crutches temporary.

    Guardrails, injected developer messages, and scaffolding should shrink as the model improves. A harness that only grows is a product debt.

  3. III

    Probe with a fair exam.

    Finding a weakness is not the same skill as proving the test is valid. Do not let your own stack grade itself without an outside panel.

  4. IV

    Assume the draft is hunted.

    Work-in-progress shared into model-shaped channels can become training signal or competitive intel. Wallface what you cannot afford to lose. Label unread paywalls instead of inventing the missing chapter.

The Collection · The Harness08