Google unveils Gemini 4 Argon with 68% score on CWE-bench
Google has unveiled its flagship Gemini 4 Argon model, which scored 68% on CWE-bench and tied GPT-6 Astra for first place. Employees told Bloomberg that the model does not always match benchmark results in practical use.

Image: itc.ua · Author: \u0428\u0430\u0434\u0440\u0456\u043d \u0410\u043d\u0434\u0440\u0456\u0439 · source articleEditorial excerpt for news reporting
Google has unveiled Gemini 4 Argon, a new flagship model with claimed capabilities in software development and cybersecurity. The company says it scored 68% and tied OpenAI’s GPT-6 Astra for first place on CWE-bench.
Argon can independently find, verify and fix critical software vulnerabilities. In an early demonstration, Google says the model identified a critical flaw in medical software used by hospitals worldwide that could potentially have exposed confidential information.
Bloomberg also reported a gap between benchmark results and Google’s internal use of the model. Employees said Gemini 4 performs worse in some work situations and struggles with certain programming tasks. Google rejected the suggestion that the model performs poorly in programming.
Claimed capabilities and internal use
Argon’s output limit has been increased to 1 million tokens. Google is using the model for quantum-computing research, code migration and data-center memory optimization; the company says telemetry analysis freed about 300 TiB of memory.
Argon agents are also helping migrate large C/C++ codebases to Rust. The work includes the re2 and libgav1 libraries and the Fuchsia OS Zircon kernel, covering more than 800,000 lines of code.
Restricted launch and pricing
Because of its cybersecurity capabilities, Google is initially limiting Argon access to Fairwind Program participants—government bodies and vetted cybersecurity partners. The model has added protection against indirect prompt-injection attacks, along with monitoring of its reasoning chain and actions intended to halt dangerous behavior.
The listed price is $4 per 1 million input tokens and $20 per 1 million output tokens. Google says an introductory launch price will be half those rates.
Employee assessments inside Google remain divided: some consider Gemini 4 state of the art, while others fear it could lag behind models from Anthropic and OpenAI. In late September, Google DeepMind chief Koray Kavukcuoglu said the company would “always be at the forefront.”
What we know
- Gemini 4 Argon scored 68% on CWE-bench and tied GPT-6 Astra for first place.
- The model can find, verify and fix critical software vulnerabilities.
- Its output limit has increased to 1 million tokens.
- Fairwind Program participants get first access, with launch pricing set at half of the listed $4 and $20 per-million-token rates.
What is being verified
- The newsroom is checking the report that gemini 4 Argon scored 68% on CWE-bench and tied GPT-6 Astra for first place.
- Reporting from ITC.ua is being compared; a second independent confirmation is not yet available.
View sources1
COMMUNITY
Discussion
Sign in to join the discussion.
No comments yet. Start the discussion.