Google DeepMind returned to the top of the leaderboard this week with Gemini 4 Argon, a model that raises the output ceiling from 64,000 to one million tokens through a new "Long Decode Continuation" interface that pauses and resumes long generations across calls. On independent scoring it sits level with GPT-6 Astra at 53 on the Artificial Analysis index, and Google claims first place on 13 of 19 benchmarks, including 77.9% on DeepSWE.
The headline number is reliability: a 15% hallucination rate against Astra's 51% — though accuracy on the same test falls to 50% from 63%, and Argon spends 62,000 output tokens per task against Astra's 27,000. List price is $4/$20 per million tokens, halved indefinitely to $2/$10, with cached input 95% off.
The catch is access. Argon is restricted to government users and trusted defenders in a cyber programme while guardrails are refined. Google is already using it internally to migrate more than 800,000 lines of C and C++ to Rust and to reclaim over 300 TiB of memory. Sceptics point to a weak legal benchmark and whisper "benchmaxxing" — but the larger signal is that the frontier now arrives gated.