In short
Argon matched GPT-6 Astra in one independent evaluation, while Google claims to lead on 13 out of 19 tests. But long outputs are costly in terms of token usage, and access is still limited.
One million output tokens—and for those who give models long tasks, it looks like a gift. But Gemini 4 Argon has a catch: in one evaluation, it used more than twice as many output tokens per task as GPT-6 Astra.
Google claims that Argon took first place on 13 of 19 tests. In an independent evaluation by Artificial Analysis, it scored 53 points, the same as Astra. The one-million-token limit is achieved through Long Decode Continuation: a long response is paused and then continued through subsequent API calls.
According to Artificial Analysis, with the introductory discount, a task on Argon cost $1.99 versus $3.26 on Astra. However, Argon used an average of 62,000 output tokens per task, while Astra used about 27,000. The savings came from the token price, not from using tokens sparingly. Without the discount, the cost of an Argon task in this evaluation would have risen to $3.98.
For now, Argon is available only to participants in the Fairwind government program and vetted cybersecurity professionals. Google promises to expand access after refining its safety restrictions. There are also questions about the results themselves: some observers doubt certain figures, and on one legal benchmark, Argon was outperformed by Muse Spark 1.2. So the record for response length is still easier to confirm than its reliability for everyday work.
What task would a model’s long response save you time on, even if it used more tokens?
Source: Latent.Space