Gemini 4 Argon coding results come with a split verdict: 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5, but behind rivals on Terminal-Bench 4.0 and FrontierSWE v2. Almost nobody can test it yet. Google announced the model on September 30 with a 1 million token output limit, up from 64,000, and opened access only to cyber defenders in its Fairwind program.
Every score below comes from Google or outside evaluators. Saganote has not run its own tests, and Google has not published a public model ID to run them against.
Gemini 4 Argon Coding Scores: Where It Wins and Where It Trails
Argon's coding record depends on who ran the test. Google computed the DeepSWE v1.1 score itself, on a mini-SWE agent harness at the highest thinking setting, then took Opus 5.5 numbers from Anthropic's system card and GPT-6 Astra numbers from the public leaderboard. Terminal-Bench 4.0 followed the same pattern, with Google evaluating Argon and pulling other models' scores from the public leaderboard.
| Benchmark | Gemini 4 Argon | Best rival | Who ran it |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | Opus 5.5 at 74.2% | Google, self-computed |
| Terminal-Bench 4.0 | 57.4% | Opus 5.5 at 66.4% | Google for Argon, public leaderboard for rivals |
| FrontierSWE v2 | 55.0% | GPT-6 Astra at 65.5% | Proximal leaderboard |
| Code Arena WebDev | Eighth place | Seven models rank higher | Arena leaderboard |
We use the highest scoring thinking level for each model.
Google DeepMind, Gemini 4 Argon model evaluation methodology

Read together, the scores point one way. Argon leads where a single attempt must fix a real repository, and it trails where an agent grinds through terminal work or runs for hours - FrontierSWE allows up to 20 hours per trial. A model that wins the benchmark its maker ran itself and loses the one an outside lab ran is worth testing, not worth switching to. Google's coding results have fallen short before: Gemini 3.5 Pro arrived months late after its coding benchmarks missed expectations.
1 Million Output Tokens Raises the Ceiling for Ports and Migrations
Output used to be the bottleneck. Earlier Gemini models stopped at 64,000 output tokens, so developers had to split a large port across many calls. Gemini 4 Argon launched with a 1 million-token output limit, a jump of roughly 15 times, and Google's examples show the scale it targets.
- Language ports: Google says Argon agents replaced 32,000 lines of SIMD code in the libgav1 video decoder, reaching a 2.7x speedup over the Rust port with identical video output.
- Large migrations: Google cites C/C++ to Rust conversions of up to 800,000 lines in Fuchsia's Zircon kernel.
- Vulnerability patching: Google says Argon finds, validates, and patches security flaws without a human in the loop.
Longer output carries costs. Artificial Analysis measured about 62,000 output tokens per task for Argon against 27,000 for GPT-6 Astra, a 2.3x gap. Google's long-context score on GraphWalks falls from 99.7% up to 128,000 tokens to 84.2% between 256,000 and 1 million, and no evaluation reviewed here measures output quality near the 1 million mark.
Verbosity Erases Argon's Price Edge Once the Discount Ends
| Model | Input per 1M tokens | Output per 1M tokens | Cache read per 1M tokens |
|---|---|---|---|
| Gemini 4 Argon (introductory) | $2 | $10 | $0.10 |
| Gemini 4 Argon (standard) | $4 | $20 | $0.20 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| GPT-6 Astra | $10 | $50 | $1.00 |
Price per token tells half the story. Artificial Analysis puts the cost of running its Intelligence Index at $1.99 per task for Argon at the 50% introductory rate, against $3.26 for GPT-6 Astra. Once the discount ends, Argon's figure rises to $3.98, about 1.2 times Astra's, because Argon writes more than twice as many tokens per task. Google has not confirmed when the introductory rate ends, and Opus 5.5 and Astra prices come from a DEV Community comparison; current rates for every major provider sit in the AI pricing comparison.

Long-horizon runs show a wider spread. On Proximal's FrontierSWE v2 leaderboard, as reported by OrcaRouter, an Argon trial averaged $129.36 and 10.6 hours, against $98.87 for Opus 5.5 and $1,029.65 for Astra. Proximal footnotes that cache hit rates ran low because of a non-production setting, so costs in production could land lower.
Which Model to Pick for Which Coding Job
A split decision fits the evidence better than a winner. Each row below maps a job type to the model the published numbers favor, and every pick could change once independent testers get hands-on time with Argon.
| Coding job | Model the evidence favors | Why |
|---|---|---|
| Fixing bugs in a real repository in one attempt | Gemini 4 Argon, test first | Leads DeepSWE v1.1 at 77.9% against 74.2%, though Google ran the test |
| Long terminal and agent sessions | Claude Opus 5.5 | Leads Terminal-Bench 4.0 at 66.4% against 57.4% and beats Argon on FrontierSWE v2 |
| Longest-horizon engineering and ML research | GPT-6 Astra | Tops FrontierSWE v2 at 65.5%, at about $1,030 per trial |
| Front-end and web UI work | A model ranked above Argon | Argon sits eighth on Code Arena WebDev |
| Cost-sensitive batch runs | Gemini 4 Argon at introductory pricing | $1.99 against $3.26 per task, until the discount ends |
Access Starts With Fairwind and Has No Public Model ID Yet
Access runs in three phases. Only the first has started, and Google gave no date for the other two beyond saying broader access will come as soon as possible.
- Phase one, live now: trusted cyber defenders in the Fairwind program
- Phase two, no date: paid API customers and Google AI Ultra subscribers
- Phase three, no date: developers, enterprises, and consumers broadly
Valletta Software reports that Argon has no public model ID and does not appear on Gemini's API pricing page. A test harness that hard-codes a guessed name will break on release day.
Seven Test Rules for Judging Argon Once Access Opens
Benchmarks cannot predict a specific codebase. A seven-rule test plan turns a public scorecard into a private one that reflects a team's own repository, task mix, and cost limits.
- Pull 10 to 20 tasks from the team's own repository, such as bug fixes with existing tests, a module migration, and a refactor. Public benchmarks never saw that code.
- Run every task five times and report the mean. Proximal reports Argon's FrontierSWE v2 score as a five-run mean with a margin of plus or minus 9.9 points, which shows how far single runs can swing.
- Fix the harness. Use the same prompts, tools, time limit, and thinking level for every model, because Google's DeepSWE number uses its own harness at the highest thinking setting.
- Track cost per solved task, not price per token. Argon wrote 62,000 output tokens per task in Artificial Analysis testing against 27,000 for Astra.
- Test long output in steps at 100,000, 500,000, and 1 million tokens, and compile and run tests on each chunk. No evaluation reviewed here measures output quality near the top of the range.
- Gate merges on the full test suite plus a security scan. Nobody can review a million tokens of generated code line by line.
- Log wrong answers separately from refusals. Argon's 15% hallucination rate in AA-Omniscience came with 50% accuracy, so a cautious model can look cleaner than it is useful.
# Test-plan skeleton. Replace the MODELS entries once Google publishes a model ID.
import statistics
TASKS = ["fix-auth-bug", "port-parser-to-rust", "add-pagination"] # tasks from your own repo
RUNS_PER_TASK = 5
MODELS = ["ARGON_MODEL_ID", "OPUS_MODEL_ID", "ASTRA_MODEL_ID"] # placeholders, not real IDs
def run_task(model, task):
# 1. Send the same prompt, tools, and thinking level to every model.
# 2. Apply the returned patch in a clean checkout.
# 3. Run the full test suite.
# Return (passed: bool, output_tokens: int, cost_usd: float)
raise NotImplementedError
for model in MODELS:
runs = [run_task(model, t) for t in TASKS for _ in range(RUNS_PER_TASK)]
solved = sum(1 for passed, _, _ in runs if passed)
total_cost = sum(cost for _, _, cost in runs)
print(model, "solve rate:", solved / len(runs),
"cost per solved task:", total_cost / max(solved, 1),
"mean output tokens:", statistics.mean(tokens for _, tokens, _ in runs))Dates are still missing. Google has given no timeline for paid API or Ultra access and has not said when the introductory price will end.
