The Cost of Thinking Is the New Line Item
The clearest artefact of the week is a free calculator that prices one workload across ten models and every route — pay-per-token, subscription, and self-hosted GPU. It exposes how reasoning tokens dominate: a 250-token answer can bill as 1,050, a 4.2× multiplier hidden inside "thinking."
The knobs matter more than the vendor. Caching drops input from $1.40 to $0.26 a million; batch APIs save about half with a day's patience; self-hosting only breaks even near 3.8 million articles a month. FinOps, not benchmarks, is becoming the discipline that separates a viable product from an expensive demo.