← News

Google Gemini 4 Argon official blog key art with Gemini logo and large 4

30 Sep 2026

Google

Google announces Gemini 4 Argon, limited to trusted cyber testers for now

Google DeepMind said Wednesday it is announcing Gemini 4 Argon, a frontier model for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense, rolling out first to trusted cyber defenders through its Fairwind Program while it strengthens safeguards before wider release.

Google is putting a frontier model on the same gated path other labs use: trusted cyber defenders first, with a wide score table and a 1 million token output ceiling published beside the gate. The first seats are for people who break and patch systems for a living, and a general release is still ahead.

On Wednesday, 30 September 2026, Google announced Gemini 4 Argon. The post is on the Google blog, dated that day, and it does not print an hour. The byline is Koray Kavukcuoglu, senior vice president of Google DeepMind and chief AI architect at Google. Google calls Argon its new frontier model. A frontier model, here, is a system at the top of what the lab says it can do, not a smaller, cheaper variant. The post says Argon is built to keep reasoning across long, complicated jobs, and that it delivers frontier performance in real software engineering, in enterprise knowledge work such as legal and finance, and in cybersecurity defense. Those lines are Google’s.

Who can use it now, as the post states it. Argon is rolling out to a set of trusted cyber defenders through Google’s Fairwind Program. A cyber defender, in that sentence, is someone whose job is to find and fix attacks, not a general customer. Google says releasing a model at this level needs a phased approach. It says it is engaged in the U.S. government’s voluntary process for pre-release model access, and that it will keep gathering feedback from early testers while it iterates on guardrails. The later release it describes is to developers, enterprises, and consumers, as soon as possible, starting with paid API customers and Google AI Ultra subscribers. An API is the paid hook a developer uses to call the model from software. The post does not name a calendar day for that wider release. It does not say Argon is in the consumer Gemini app today.

The introductory price on the post is $2 per million input tokens and $10 per million output tokens. A token is a small piece of text the model reads or writes. A million tokens is the unit these prices use, not a count of users. Cached input tokens are priced at 95 percent off the input price. Cached input is text the service has already stored from an earlier request, so a later call can reuse it. Ninety-five percent off the introductory $2 rate is 10 cents per million cached tokens. The post states the discount. It does not print that 10-cent line as its own price. A footnote says that after the introductory period, the price is $4 per million input tokens and $20 per million output tokens. The post does not say when that introductory period ends. Those prices are Google’s.

The output ceiling is 1 million tokens, up from 64,000 tokens Google cites for prior models. One million is a bit more than fifteen times 64,000. Output tokens are what the model writes in one run, including the long stretch of reasoning it may do before a short answer. Google calls 1 million an industry-leading ceiling, and says the extra room lets the model think through a hard problem in one pass instead of stopping early. That ceiling and that comparison are Google’s.

On software work, Google says its engineers are already using Argon for everyday debugging, large code moves, and algorithm design. It says Argon sets a new state of the art on DeepSWE v1.1 at 77.9 percent. DeepSWE v1.1 is a test of long software jobs in real code, the kind that take many steps. State of the art, in that sentence, is Google’s claim that 77.9 percent is the best published score it is showing. On the chart Google posts for that test, 9to5Google’s same-day reading puts Claude Opus 5.5 at 74.2 percent and GPT-6 Astra at 74.1 percent. Those bars are on Google’s chart. They are not a second lab’s recount. The chart is not a clean sweep. On the same comparison graphic, another model leads on FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science, and OSWorld-2.0. A lead on one row is a lead on that row.

Past coding, Google says Argon is the leading model on the Vals Index. That index, as Google describes it, measures economic impact across finance, coding, legal, and tax, with each sector weighted by its share of U.S. gross domestic product. Gross domestic product is the country’s total output. Google also says Argon leads on Vals Finance Agent v2, a multi-step finance-research test, and on Harvey’s Legal Agent Benchmark, a test of legal research and drafting. The article text does not print percents for those three. The charts on the post carry the bars. On AutomationBench, which Google describes as Zapier’s test of end-to-end work across ordinary business jobs, Google says Argon ranks first at 51.3 percent. Fifty-one percent is a little over half. On LVBench, a test of understanding long video, Google says Argon is state of the art at 91.7 percent. That is a bit more than nine in ten. Those scores are Google’s.

On security repair, Google says Argon ties for first on CWE-bench v1 at 68 percent. CWE-bench, as Google describes it, tests whether a model can fix security vulnerabilities. Sixty-eight percent is about two in three. The post says that result builds on Gemini 3.8 Flash Cyber’s frontier score on CWE-bench v0. It says “ties for first” and does not name the other model in the article text. Google also says Argon found a wider set of exposures than 3.8 Flash Cyber on an internal vulnerability benchmark covering codebases in 20 programming languages, and that on Wiz’s internal test of live websites, with no source code handed over, Argon did better at mapping the attack surface, finding vulnerabilities, and producing a proof that the flaw is real. A proof, here, is evidence the hole can be triggered, not a patch already shipped to hospitals. Those comparisons are Google’s, including the internal ones.

What Google says Argon is already doing inside the company. Thousands of Googlers, the post says, have pointed to its coding, research, and writing. Three examples are Google’s claimed internal results. On quantum algorithms, Argon helped researchers shrink the spacetime cost of subroutines that slow important applications. Spacetime, in that line, is the qubits times the gates, the hardware steps a quantum routine needs. In one example, Google says it beat a published baseline by 40 percent in a matter of minutes. On memory, a team of Argon agents read fleet-wide profiling data, applied memory fixes across Google’s data centers, and freed more than 300 tebibytes once the changes rolled out, with an estimated 500 tebibytes to 1 pebibyte in total savings. A tebibyte is 1,024 gigabytes. A pebibyte is 1,024 tebibytes, so the high end of that estimate is a bit more than three times the 300 tebibytes Google says are already free. On code moves, Argon agents are migrating C and C++ to Rust across Google, from tens of thousands of lines in libraries such as re2 and libgav1 up to more than 800,000 lines for the Fuchsia operating system’s Zircon kernel. Google says those rewrites are still in automated and manual review, emulation testing, and audit before they run in production, because many of the systems are critical. For libgav1, Google’s open-source video decoder, the agents rewrote 32,000 lines of SIMD code. SIMD is code that does the same math on many numbers at once. Google says the result is a memory-safe decoder that runs 2.7 times faster than the earlier Rust port, with the same video output, and closer to the optimized C++. Those results are Google’s account of its own systems. They are not an outside audit.

For trusted defenders and Google’s own internal teams, the post says Argon will be released without cyber guardrails, so they can use the full defensive capability. Guardrails, here, are the limits that stop a model from helping with an attack. Google says Argon can find, check, and patch critical software vulnerabilities on its own. Wiz, Google says, is already using Argon through Scan for Good, a program for finding and fixing high-risk exposures in critical public infrastructure at no charge. In an early demo, Google says the model found a critical vulnerability that exposed sensitive personal information in healthcare software used by hospitals worldwide, a hole previous frontier models had missed. The post does not name the software, a public vulnerability ID, or a patch.

Before a broad release, Google says it is still strengthening four kinds of safeguards. On misuse, the model is designed to refuse requests that would help with a cyber attack or with chemical, biological, radiological, or nuclear harm, while still allowing legitimate scientific research, under Google’s Frontier Safety Framework. Google says it is improving checks on the model’s internal activations, the signals inside the network, to spot misuse, and that internal and external red teams tested those checks with hand attacks and automated ones. On prompt injection, Google says Argon is its most resilient model yet against hidden instructions that try to hijack a model from inside a document or a page. It says Argon leads on Gray Swan’s Indirect Prompt Injection benchmark. The article text does not print a percent for that lead. On misalignment, Google says it watches Argon’s chain of thought and its actions and stops the run when the model steps past what the person asked. A chain of thought is the model’s written reasoning. Google says a similar watch was used on training runs, with alerts to an incident-response team, and that the findings were kept out of training so the model would not learn to hide from the monitor. On the systems around the model, Google says it is sealing sandboxes before high-risk training or evaluations. A sandbox is a closed test environment. Those four areas are Google’s description of work still underway before a wide release.

In plain terms, Google said on Wednesday that Gemini 4 Argon is its new frontier model, priced at $2 and $10 per million tokens for an introductory period, with a 1 million token writing ceiling. The scores it prints, including 77.9 percent on DeepSWE v1.1, 51.3 percent on AutomationBench, 91.7 percent on LVBench, and a 68 percent tie on CWE-bench v1, are Google’s. The chart it posts also shows tests where Argon is not first. The people using it now are trusted cyber defenders in the Fairwind Program, plus teams inside Google. A wider release to developers, companies, and consumers is described as later, starting with paid API customers and Google AI Ultra subscribers. The post does not give that day.

The picture is Google’s Gemini 4 Argon launch graphic. A blue field holds the Gemini spark and the words Gemini 4 Argon, with a large white 4 on the right. It is the company’s key art for this post. It does not print a calendar date.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources