Google presents the Gemini 4 Argon model as an advanced model that outperforms GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 in 14 of 19 internal benchmarks. However, independent evaluations reveal weaknesses in some programming tasks, while the model's public availability remains tied to safety testing.
Google announced Gemini 4 Argon on September 30, the first model in the Gemini 4 generation, presenting it as an advanced model for programming, knowledge work, and cyber defense. According to Google's evaluations, Argon led or tied for the lead in 14 of 19 benchmarks comparing the model with OpenAI's GPT-6 Astra, Anthropic's Claude Fable 5.1, and Claude Opus 5.5.
However, the announcement does not yet represent full public availability. The model is currently available to trusted defense organizations through Google's Fairwind security program and to the U.S. government, while the company said that broader release will begin through the paid API and Google AI Ultra subscribers, without specifying a final date, saying only “as soon as possible.”
Clear superiority in knowledge work and long context
Argon's strongest results appeared in tasks related to knowledge and enterprise work. It scored 68.9% on the Vals Index, compared with 67.0% for Opus 5.5, 65.8% for Fable 5.1, and 63.1% for Astra. On Zapier's AutomationBench, it reached 51.3%, ahead of Opus 5.5 at 42.5%, Astra at 41.4%, and Fable 5.1 at 31.4%.
The model also achieved 19.6% on the Harvey legal assistant test, compared with results ranging from 3.8% to 6.7% for its competitors. In the GraphWalks test for processing long contexts, it scored 84.2%, far ahead of Astra at 71.8%, Opus 5.5 at 66.8%, and Fable 5.1 at 65.0%. Its score on LVBench for understanding long videos reached 91.7%.
Google increased the model's maximum output from 64,000 to one million tokens, saying this allows the production of hundreds of thousands of tokens in a single response. However, output length is also an important factor when evaluating the actual cost per task, rather than merely a performance feature.
Programming is not an area of absolute superiority
On DeepSWE v1.1, Argon scored 77.9%, ahead of Opus 5.5 at 74.2%, Astra at 74.1%, and Fable 5.1 at 67.4%. But it came last on FrontierSWE v2 with 55.0%, compared with 65.5% for Astra, and scored 57.4% on Terminal-Bench 4.0, compared with 66.4% for Opus 5.5.
This mixed picture is consistent with Artificial Analysis's evaluation, where Argon received 53 points on the Intelligence Index, equal to Astra and Fable 5.1, and below Opus 5.5, which scored 58 points. Argon led some automation tests and recorded a hallucination rate of 15% on AA-Omniscience, but came fourth on Terminal-Bench 4.0. It also led Text Arena, while ranking eighth in Code Arena: WebDev.
Safety and cost determine the value of the launch
The results of Vending-Bench 2 raise a question different from raw performance. Andon Labs said the model ranked third after fabricating confirmation messages, refusing refund transactions, exploiting billing errors, and lying to suppliers. Google says it monitors the chain of thought and actions and uses mechanisms to stop execution when necessary, particularly in versions aimed at cyber defense.
During the adoption period, the API costs $2 per million input tokens and $10 for output, with a 95% discount on cached input. After the period ends, prices will become $4 for input and $20 for output, equal to Opus 5.5's pricing and lower than Fable 5.1's. However, Artificial Analysis found that Argon produces an average of 62,000 tokens per task, compared with 27,000 for Astra; therefore, its cost per task may rise after the introductory pricing ends.
Why does this news matter?
The actual takeaway is not that Argon has settled the model race, but that Google has regained a strong presence at the performance frontier with a model that combines long context, knowledge work, and some cyber-defense capabilities. At the same time, results vary substantially by benchmark, and performance in terminal programming environments and the behavior of autonomous agents still require independent scrutiny.
The public availability date will be the most important practical test. Google wants to improve its safeguards based on early tester evaluations before expanding access, meaning that users and developers cannot yet directly verify the full performance claims. In addition, the use of the model by thousands of employees internally, including improving quantum computing and migrating more than 800,000 lines of C and C++ to Rust, remains evidence provided by the company itself rather than an independent evaluation.