← The stand
Cover of The Collection, Volume 1, Number 37: The Invoice. Sunday 27 September 2026, Melbourne. Skip cover

Vol. 1  ·  No. 37  ·  Sunday 27 September 2026  ·  Melbourne


The Collection

The Invoice

Collected and edited by Newsletter World for AK

Contents

Saturday left The Butler in the window: a fee that decides whose. Overnight the letters measured the agent’s bill. The AI Corner printed GitHub’s cost engineering and watched shorter tool calls raise the finished invoice. Exponential View named safety that sands edges. Generative Programmer mapped contracts at seven agent boundaries. DealBook asked what happens to the billable hour when the work halves. They shortened the call. The invoice grew.

  1. iiiEditor’s LetterThey shortened the call. The invoice grew.03
  2. ivThe UnitTokens per call is the wrong meter. The finished job is the unit.04
  3. vThe NumbnessSafety that sands edges is numbness. Disagreeing with an LLM is a good sign.05
  4. viThe BoundarySeven contracts at the edges. A map is not a mandate.06
  5. viiThe HourIf the work takes half the time, is the bill half?07
  6. viiiStanding OrdersFour rules for this issue.08
  7. ixColophonThe letters, named.09

03  ·  Editor’s Letter

They shortened the call. The invoice grew.

Saturday closed The Butler on a fee that decides whose. Sunday’s mail opens the ledger. The AI Corner’s Ruben Dominguez walks through GitHub’s published cost engineering: a Rust utility shortened shell output for a coding agent; responses got shorter; when the omitted text mattered, the agent recovered with extra turns that dragged the full history; average tasks used more tokens and took longer while completion rates held. The unit of measure must be the finished job, not tokens per call. Exponential View’s Azeem Azhar reads MFA homogenization as a preview of LLMs that sand edges into numbness, and treats disagreement with a model as a good sign. Bilgin Ibryam’s Generative Programmer maps seven boundary contracts (Open Responses, MCP, A2A, AG-UI/A2UI, AGENTS.md and Skills, OpenTelemetry GenAI still in Development, ACP) and says you compose them independently. A map is not a mandate. DealBook asks whether a lawyer who finishes in half the time should bill half.

They shortened the call. The invoice grew.

The finished job is the unit.

Sunday fills the second slot on this week’s top rack. Saturday’s Butler stays beside it. Monday through Friday wait on letters. Last week’s shelf (Unsold through Latency) does not move.

The Collection · The Invoice03

04  ·  The Desk

The Unit

GitHub shortened the tool reply. The finished job got dearer. Measure the task, not the call.

Ruben Dominguez’s letter, filed under shorter prompts making agents more expensive, keeps the arithmetic plain. Agentic workloads land somewhere around a thousand times the token consumption of ordinary chatbots, because every tool turn rewrites the context and the attention cost scales with the square of sequence length. Peter Steinberger, who built OpenClaw, is cited for $1.3 million on tokens in one month: 603 billion of them, across 100 coding agents run by three people. A Harvard, MIT, and Northeastern token-reduction paper names OpenClaw and Codex as the systems worth studying for that burn.

At roughly $2.50 per million input tokens, a run carrying 30,000 tokens of context costs seven and a half cents. Ten thousand runs a day is $750, call it $22,500 a month, for input alone. Uber reportedly exhausted its entire 2026 AI budget in four months. Microsoft is reported to have ended Claude Code licences after a pilot that began in December 2025. In a survey of 2,500 decision-makers, the rollback rate reached 81% at firms with mature governance.

A cheaper call can raise the invoice.

GitHub evaluated Rust Token Killer against its own agentic coding benchmarks. Tool responses got shorter. When omitted text mattered, the agent reopened the original or reran the command. Each recovery added a turn that dragged the accumulated history. On average the task consumed more tokens and took longer. Completion rates held. So each call got cheaper, and the whole task got more expensive.

Waste concentrates in four reservoirs: prompts and tool output, history and thinking. Caching with static content first can drop the stable portion toward a tenth of the normal input rate; one research agent saw input costs fall 87% on that change alone. GitHub’s removal of dead line numbers from a file-reading tool cut inference cost around 5% offline and about 3% per user per day in production. A 42-page report cost 84,000 tokens per call as a PDF and 9,500 after conversion to plaintext. The shipped compressor stayed conservative: diffs and source were left alone after agents reopened originals. An automated prompt rewrite halved size, then broke parallel subagents into serial work; GitHub stopped the experiment and added a test. Sunday keeps those figures as printed. The invoice records what nobody decided.

The Collection · The Invoice04

05  ·  The Sheet

The Numbness

Workshop protocols sand peculiarity. LLMs amplify the middle. Disagreement is a vital sign.

Azeem Azhar’s Exponential View, subject-lined Safety in numbness, opens on Condorcet’s hope that shared notations would propel knowledge forward, then turns to Erik Hoel’s essay on how MFA programmes swallowed literary fiction. The protocol is group workshop after group workshop. Peer review sands away peculiarity, weird ambition, and the unfamiliar edges most likely to attract an objection. Out the other end comes writing homogenised to survive the room: “wan little husks of ‘auto fiction’,” as Joyce Carol Oates put it. There were 15 MFA programmes in the US in 1975 and 250 by 2012. Hoel argues the craft and the programme that teaches it became identical.

Azhar reads LLMs as standardised systems trained against particular incentives, outputting against a statistical distribution with its own shape. Applied as a rubric, they amplify sameness. They parallel workshops that sand sharp edges. They norm by finding the middle. Outliers are unwelcome. Homogenisation, which he treats as safety institutionalised, makes the next Rushdie harder to find.

Safety that sands edges is numbness.

The letter’s working theme, kept as printed in the free matter and the paid tease, is that disagreeing with an LLM is a good sign. Sunday does not invent MFA enrolment beyond Hoel’s 15-to-250 arc, and does not invent the paid depth past what the open letter carries. The Collection files the numbness claim beside the invoice: a system that cuts what looks expensive can also cut what made the work distinct.

The Collection · The Invoice05

06  ·  The Bench

The Boundary

Seven contracts at the edges of an agent. Compose them independently. A map is not a mandate.

Bilgin Ibryam’s Generative Programmer maps the emerging standards behind AI agents as contracts at boundaries, not as a single industry creed. An agent must reach a model, tools, remote agents, users, organisation-specific instructions, operational systems, and, for coding work, an IDE. Each connection is a boundary. Vendors and working groups are converging on a different contract at each one.

The seven kept on Sunday’s sheet: Open Responses for calling models; MCP for tools (closest to a de facto standard); A2A for delegating work to other agents; AG-UI and A2UI for interacting with users (AG-UI 1.0 still unratified); AGENTS.md, Agent Skills, and plugins for packaging instructions; OpenTelemetry GenAI semantic conventions for tracing (the GenAI conventions remain marked Development); ACP for connecting coding agents to IDEs (ACP v2 still experimental). You can compose these independently. Use MCP without adopting A2A. Most teams start with model access and add a contract when the agent crosses a new edge.

A boundary contract is not a whole-system standard.

Sunday’s angle is the contract at the boundary, not a claimed industry consensus. Thursday already filed that agreement is not adoption. Today the map stays a map: adopt a contract where you need to replace what sits on the other side without rewriting the agent, and leave the rest alone until you hit that point.

The Collection · The Invoice06

07  ·  The Floor

The Hour

If a lawyer uses AI and finishes in half the time, should the bill be half? Big firms are eager to show embrace.

Sarah Kessler’s DealBook item, Beyond the billable hour, puts the invoice question in a profession that still meters time. If counsel uses AI and does the work in half the time, should the fee fall with the clock? The letter notes big firms eager to show they embrace AI. Sunday keeps that tease as printed and does not invent firm names, fee schedules, or partnership votes that the open matter does not carry.

Time cut is not value cut.

The Collection files The Hour beside The Unit. A shorter call that forces recovery turns raises the agent’s invoice. A shorter matter that still carries the same liability may not halve the lawyer’s. Both ask what the meter is measuring.

The Collection · The Invoice07

08  ·  Standing Orders

Four rules for this issue

  1. I

    The finished job is the unit.

    Tokens per tool call measure the wrong thing. Count the task that ships, not the reply that looks cheap.

  2. II

    A shorter call can raise the invoice.

    Omit what the agent must recover, and every recovery turn drags the history. GitHub watched the average task grow dearer while completion held.

  3. III

    Safety that sands edges is numbness.

    Workshop protocols and LLMs both pull toward the middle. Disagreement with a model is a vital sign, not a defect.

  4. IV

    A boundary contract is not a whole-system standard.

    Compose MCP, A2A, AG-UI, and the rest at the edges you cross. A map of contracts is not a mandate to adopt them all.

The Collection · The Invoice08