Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense


Google DeepMind has just announced Gemini 4 Argon, its new frontier model and the first model of the Gemini 4 generation. It targets long-horizon software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The biggest technical change is output length. Argon can generate up to 1M tokens in a single response, up from 64K on earlier Gemini models.

What Google Announced

Google DeepMind described Argon as built for complex workflows across coding, enterprise knowledge work and cybersecurity defense.

Google is taking a phased approach. It is participating in the U.S. government’s voluntary process for pre-release model access. It will gather feedback from early testers and iterate on guardrails before a wider release.

Pricing is already public. Argon launches at an introductory $2 per 1M input tokens and $10 per 1M output tokens. Cached input tokens get a 95% discount, which works out to $0.10 per 1M. After the introductory period, pricing moves to $4 input and $20 output. Logan Kilpatrick confirmed the introductory $2 in and $10 out pricing.

Why the 1M Output Limit Matters

Current frontier APIs cap a single response far lower. Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra each allow 128K output tokens.

Google team states that Argon can think deeply and generate hundreds of thousands of tokens in one trajectory. For developers, that means large refactors or long reports without splitting work across turns. The cost is real, though. A full 1M output tokens costs $10 at introductory pricing and $20 after.

Google has not disclosed Argon’s input context window.

Benchmarks: Where Argon Leads and Where It Trails

Google compared Argon against GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1. Argon leads outright on 12 of 18 benchmarks and ties for first on 1.

Where it leads:

  • DeepSWE v1.1 (long-horizon software engineering): 77.9%, a new state of the art. Opus 5.5 scores 74.2% and GPT-6 Astra 74.1%.
  • Vals Index (economic impact across finance, coding, legal and tax): 68.9%, ranked first.
  • AutomationBench (Zapier, end-to-end business execution): 51.3%, ranked first. Opus 5.5 scores 42.5%.
  • Harvey Legal Agent Benchmark: 19.6%, against 5.4% for GPT-6 Astra.
  • LVBench (long video understanding): 91.7%, a new state of the art.

Where it trails:

  • FrontierSWE v2: 55.0%, behind GPT-6 Astra at 65.5%.
  • Terminal-Bench 4.0: 57.4%, behind Claude Opus 5.5 at 66.4%.
  • OSWorld-2.0 (computer use): 69.2%, behind GPT-6 Astra at 72.6%.

Artificial Analysis reported that Argon equals GPT-6 Astra on its Intelligence Index at 60% of the cost per task, using discounted prices.

Cyber Defense: Find, Validate, Patch

Google trained Argon to autonomously find, validate and patch critical software vulnerabilities. Trusted defenders and internal Google teams receive it without cyber guardrails.

On CWE-bench v1, which tests vulnerability remediation, Argon ties for first at 68%. The rival models on that leaderboard run inside their own agent harnesses.

Wiz is already using Argon through its Scan for Good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide. Google says previous frontier models had missed it.

Before broad release, Google is strengthening safeguards in 4 areas:

  • Misuse defenses for cyber and CBRN risks, including activation monitoring, under its Frontier Safety Framework.
  • Indirect prompt injection resistance, where Argon leads Gray Swan’s IPI benchmark.
  • Misalignment monitoring of chain-of-thought and actions, with the ability to stop execution.
  • Sealed, isolated sandboxes for high-risk training and evaluations.

‘)});
var tm;
function playOut(){var bs=O.querySelectorAll(‘.bar’);bs.forEach(function(b){b.style.width=”0″});setTimeout(function(){bs.forEach(function(b){b.style.width=b.dataset.w+’%’})},80);
var c=R.querySelector(‘#mx-count’),t0=null;cancelAnimationFrame(tm);function st(ts){if(!t0)t0=ts;var p=Math.min((ts-t0)/1600,1);c.textContent=Math.round(1000000*(1-Math.pow(1-p,3))).toLocaleString(‘en-US’);if(p<1)tm=requestAnimationFrame(st)}tm=requestAnimationFrame(st)}
R.querySelector(‘#mx-play’).onclick=playOut;

var ST=[[‘Fairwind Program’,’Live now’,1,’Trusted cyber defenders and internal Google teams. They receive Argon without cyber guardrails for defensive work.’],[‘Paid API + AI Ultra’,’Next, no date’,0,’Paid Gemini API customers and Google AI Ultra subscribers are first in line once Google finishes iterating on guardrails.’],[‘Enterprises + consumers’,’Later’,0,’Broader availability to developers, enterprises and consumers “as soon as possible”, per Google. No free tier announced.’]];
var P=R.querySelector(‘#mx-pipe’),I=R.querySelector(‘#mx-info’);
ST.forEach(function(s,i){var d=document.createElement(‘div’);d.className=”stage”+(s[2]?’ live’:”);d.innerHTML=’‘+s[0]+’‘+s[1]+’‘;d.onclick=function(){P.querySelectorAll(‘.stage’).forEach(function(x){x.classList.remove(‘sel’)});d.classList.add(‘sel’);I.textContent=s[3];report()};P.appendChild(d)});
P.firstChild.click();

var M=[[‘Argon (intro)’,2,10,1],[‘Argon (standard)’,4,20,1],[‘Claude Opus 5.5’,4,20,0],[‘Claude Fable 5.1’,10,50,0],[‘GPT-6 Astra’,10,50,0]];
var fmt=function(v){return v>=1?v.toFixed(1)+’M’:Math.round(v*1000)+’K’};
function calc(){var i=+R.querySelector(‘#mx-in’).value,o=+R.querySelector(‘#mx-o’).value;R.querySelector(‘#mx-in-v’).textContent=fmt(i);R.querySelector(‘#mx-o-v’).textContent=fmt(o);
var cs=M.map(function(m){return i*m[1]+o*m[2]}),mx=Math.max.apply(null,cs);
R.querySelector(‘#mx-cost’).innerHTML=M.map(function(m,k){return ‘

‘+m[0]+’$’+cs[k].toFixed(2)+’

‘}).join(”)}
R.querySelector(‘#mx-in’).oninput=calc;R.querySelector(‘#mx-o’).oninput=calc;

var B={‘DeepSWE v1.1’:[‘Long-horizon software engineering. Argon sets a new state of the art.’,[77.9,74.2,67.4,74.1]],’Vals Index’:[‘Economic impact across finance, coding, legal and tax, weighted by U.S. GDP share.’,[68.9,67.0,65.8,63.1]],’AutomationBench’:[‘Zapier benchmark for end-to-end business execution.’,[51.3,42.5,31.4,41.4]],’LVBench’:[‘Long video understanding.’,[91.7,83.7,79.7,87.5]],’FrontierSWE v2′:[‘Agentic coding. Argon trails here; GPT-6 Astra leads.’,[55.0,62.3,56.3,65.5]],’Terminal-Bench 4.0′:[‘Command-line problem solving. Claude Opus 5.5 leads.’,[57.4,66.4,57.9,58.2]]};
var NM=[‘Gemini 4 Argon’,’Claude Opus 5.5′,’Claude Fable 5.1′,’GPT-6 Astra’],bk=’DeepSWE v1.1′,BT=R.querySelector(‘#mx-bt’);
Object.keys(B).forEach(function(k){var b=document.createElement(‘button’);b.className=”tab”;b.textContent=k;b.onclick=function(){bk=k;drawB()};BT.appendChild(b)});
function drawB(){BT.querySelectorAll(‘.tab’).forEach(function(t){t.classList.toggle(‘on’,t.textContent===bk)});var d=B[bk],top=Math.max.apply(null,d[1]);R.querySelector(‘#mx-bdesc’).textContent=d[0];
var W=R.querySelector(‘#mx-bars’);W.innerHTML=NM.map(function(n,k){return ‘

‘+n+(d[1][k]===top?’ ★’:”)+’‘+d[1][k].toFixed(1)+’%

‘}).join(”);
setTimeout(function(){W.querySelectorAll(‘.bar’).forEach(function(b){b.style.width=b.dataset.w+’%’})},60)}

var CY=[‘Find: Argon uncovered a critical flaw in hospital healthcare software via Wiz Scan for Good, missed by earlier frontier models.’,’Validate: on Wiz\’s black-box pentest benchmark, Argon produces proof-of-concept evidence for the vulnerabilities it finds.’,’Patch: on CWE-bench v1 (vulnerability remediation) Argon ties for first place at 68%.’],ci=0;
setInterval(function(){for(var k=0;k<3;k++)R.querySelector(‘#mx-n’+k).classList.toggle(‘act’,k===ci);R.querySelector(‘#mx-cinfo’).textContent=CY[ci];ci=(ci+1)%3},1500);
go(0);window.addEventListener(‘resize’,report);setTimeout(report,400);
})();



Source link

  • Related Posts

    Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

    Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval…

    NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

    Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues…

    Leave a Reply

    Your email address will not be published. Required fields are marked *