← The stand
Cover of The Collection, Volume 1, Number 27: The Factory. Wednesday 16 September 2026, Melbourne. Skip cover

Vol. 1  ·  No. 27  ·  Wednesday 16 September 2026  ·  Melbourne


The Collection

The Factory

Collected and edited by Newsletter World for AK

Contents

A letter, four pieces, standing orders, and a colophon. Tuesday’s Charter asked whether the stopwatch has a spring. Overnight the letters answered from the shop floor. The factory runs. The job that remains is naming the goal.

  1. iiiEditor’s LetterThe factory runs. Name the goal.03
  2. ivThe TakeoverCodex went from nice-to-have to backbone.04
  3. vThe HandholdJudgment moved to the goal. The agent walks the change to prod.05
  4. viThe JudgmentA linter for judgment, not another essay.06
  5. viiThe TransferOnly the agent that could explore transferred.07
  6. viiiStanding OrdersFour rules for this issue.08
  7. ixColophonThe letters, named.09

03  ·  Editor’s Letter

The factory runs. Name the goal.

Tuesday’s Charter asked whether the stopwatch has a spring. Overnight the letters answered from the shop floor. Gergely Orosz at The Pragmatic Engineer walked OpenAI and found Codex has become the backbone of almost everything: seven leaders on the record, non-engineering orgs from near zero to nine-tenths usage in about four months, long-running agents that spin off other agents. Mike Taylor at Every timed TypeSafe’s Jev, a judgment model that returns probabilities fast enough to sit inside an agent loop while the work is still happening. Latent Space’s swyx interviewed Good Start Labs on game environments that transfer to finance work, and which harness designs actually carry the skill. Semafor’s Echoes of 2007 and Gary Marcus translating Sam stay on Monday and Tuesday’s rack; we do not re-harvest the pacing fight. On the Melbourne street, Cyber Daily’s Auto-IT ransomware note is weather, not a piece.

Everything after the goal is an agent. Naming the goal is still a person.

The factory runs. The job that remains is naming the goal.

Wednesday fills. Thursday and Friday are still empty on this week’s rack.

The Collection · The Factory03

04  ·  The Desk

The Takeover

Codex went from nice-to-have to backbone. Almost everyone uses it weekly. The phone store is still slow.

Orosz sat with seven OpenAI leaders: Venkat Venkataramani (VP Engineering, Applied Infrastructure), Sulman Choudhry (Head of Engineering, ChatGPT), Andrew Ambrosino (Lead, Desktop), Joe Gershenson (Lead, Core Agent), Akshay Nathan (Engineering Lead, Productivity), Ahmed Ibrahim (Codex), and Steve Coffey (Responses API). The through-line is not a demo. Non-engineering orgs in finance, recruitment, and legal went from roughly zero percent to about ninety percent Codex usage in about four months. Almost all employees use Codex and ChatGPT Work weekly. The Mac app landed in February, Windows in March, ChatGPT Work as a Codex harness in July. Adoption surged even while the app still put code on screen, hostile to non-engineers; roughly forty percent non-engineering adoption then.

The /goal setting keeps the agent working until the outcome is complete. April to May usage moved from about sixty percent to about ninety percent. Threads last days. Long-running agents spin off other agents. Akshay Nathan names the “awareness overhang”: the capability is already there; people discover uses by word of mouth. Role-specific plugins and domain experts embedded in ChatGPT Work engineering teams close the gap. Full dependence shows up as minor outages spotted by internal messages as fast as, or before, automated alerts. IDE usage has fallen since January as Codex surged. The team hesitated in December 2025 about releasing the app versus CLI and IDE paths; an Antigravity VS Code fork existed; they bet IDEs would matter less. Pull requests per engineer took a hockey-stick shape. Venkat reports roughly a tenfold load on some build-test-deploy systems in about six months.

Everything is now a coding agent.

Native mobile still waits on App Store and Google review measured in hours and days. Sulman Choudhry’s line: code generation is minutes; shipping to phones is still slow. Facebook’s 2010s weekly-plus-feature-flag model remains the prior breakthrough. Codex mobile-first makes the gap painful. Waiting days to get code onto a phone looks absurd.

The Collection · The Factory04

05  ·  The Sheet

The Handhold

Nine stations from outcome to production. The human still names the goal and still says yes to the mitigation.

The agentic software factory is a pipeline, not a slogan. A human defines the desired outcome; engineers sound more like product managers. Codex gathers context: docs moved into source, plus GitHub, Slack, Notion, Databricks, Datadog, logs, and internal skills. It implements until the goal is met and verifies. Build, test, and CI follow; the agent babysits the pull request to green; a perf harness sends problematic PRs to Synthetics A/B. Agentic code review stacks multiple domain-specialist agents, classifies risk, auto-approves the low-risk ones, and keeps high-risk changes stricter or human. After human approval, agentic deploy “handholds” the change to production, including feature-flag rollouts, and builds its own monitoring dashboard; the long-term shape is per-change autonomous SRE. Agents now make the per-change dashboards humans used to. Perf Factory has agents sift alerts, de-dupe, root-cause latency regressions, and propose fixes. Sevbot is the incident agent on Codex: it collects context and suggests mitigations, and never executes without an engineer’s say-so. The goal is autonomous routine outages. Oncall is not gone yet.

Handhold this change until it is safely rolled out.

Judgment moved to the goal. Everything after that line is machinery that still needs a named person at the hard gates: approve to prod, App Store, execute the mitigation.

The Collection · The Factory05

06  ·  The Bench

The Judgment

Probabilities in fractions of a second. Cheap enough to sit inside the loop while the work is still wet.

TypeSafe launched Jev: a model that returns probabilities for yes/no or defined categories, structured answers for code rather than chatbot essays. Training is Reinforcement Learning for Calibrated Decisions (RLCD); confidence should match accuracy. Cofounder Diogo Almeida, who coauthored InstructGPT in 2022 at OpenAI, puts the problem bluntly: the problem is the text itself. The manifesto line: “We’re building prod, not God.” System One architecture is not token-by-token; it answers parallel questions in fractions of a second. Pricing is $42 per billion tokens, against most LLMs priced per million, with no charge for output tokens (“too cheap to meter”).

Mike’s test: twenty-seven Every articles plus ten AI-styled counterparts, twenty-one AI-tell questions run concurrently, 777 judgments in under 0.7 seconds for about a quarter of a cent. Across eleven experiments: 1,709 judgments for less than a cent. Against Dan Shipper’s comparison stack, Jev median 0.35 seconds per passage versus Fable 5.1 high effort at 8.83 seconds, about twenty-five times faster and about 580 times cheaper. Jev caught six of seven intended defects; Fable caught seven of seven; the miss was an unexplained “teach a calendar” action. The intended use is a code-linter-for-knowledge-work inside Codex and Claude loops while work is in progress. Accuracy for production is still an open question. Useful as early warning.

We’re building prod, not God.

A linter for judgment, not another essay at the end of the shift.

The Collection · The Factory06

07  ·  The Floor

The Transfer

Diplomacy to support. Railroads to finance. Only the terminal agent that could explore carried the skill out of the game.

Good Start Labs spun out of Every in October 2025 with $3.6 million from General Catalyst, Inovia, Every, and angels. The idea came from a 2025 Twitch Diplomacy stream: o3 planned betrayal and won; Opus 4 refused to lie and got destroyed. Fine-tuning on Diplomacy improved customer support and industrial operations benchmarks, per Duffy and Every. They trained a 30B model in 1830: The Game of Railroads and Robber Barons, almost no luck except initial play order, with a stock-market mechanic and tasks that mirror finance workflows from database to Excel to functions.

Single-turn Q&A versus a multi-turn terminal agent with tools: both improved in-game; only the terminal-agent design improved the Finance-Agent benchmark. Harness design changes what is learned: pictures versus text versus Python. More capable base models need less handholding for the same task, but for “environment as curriculum” the harness matters more. Astra does less chain-of-thought; a harness can force using code so the result is trustworthy. The company sells trajectories and learning environments to frontier labs, anonymized. The evidence so far: goal-directed execution and reasoning transfer (1830 to finance; Diplomacy to support; environments improve tool use). Breadth and reliability of transfer are still open.

The environment is the curriculum.

Only the agent that could explore transferred. Quiz scores stayed in the game.

The Collection · The Factory07

08  ·  Standing Orders

Four rules for this issue

  1. I

    Name the goal before you open the factory.

    Everything after the goal is an agent. If you skip the naming, you are running machinery without a product.

  2. II

    Count what still requires a human name.

    Approve to production. App Store and Google review. Execute a Sevbot mitigation. Those gates are still people. Measure them.

  3. III

    Put a judgment linter in the loop.

    Do not wait for the essay at the end. Probabilities while the work is wet beat a chatbot postmortem after ship.

  4. IV

    Prefer environments that transfer.

    Multi-turn, tools, verifiable rewards. Single-turn quiz scores improve the game and leave the finance bench cold.

The Collection · The Factory08