Gemini 4 Argon: 1M Output Tokens at $2/$10 — AI Digest

Google shipped Gemini 4 Argon: first on 13 of 19 benchmarks, up to 1M output tokens and $2/$10 at the launch discount. OpenAI nears $70B in annualized revenue.
Today's top stories
Google ships Gemini 4 Argon: first on 13 of 19 benchmarks and 1M output tokens
Google DeepMind introduced Gemini 4 Argon, a model for coding, enterprise knowledge work and cyber defense that Google says beats GPT-6 Astra and Claude Opus 5.5 on 13 of 19 published benchmarks. The launch was announced by Google DeepMind and on the Google blog. On DeepSWE, Gemini 4 Argon scores 77.9% versus 74.2% for Opus 5.5 and 74.1% for Astra, The Rundown reports. Access starts with government users and trusted cyber defenders in the Fairwind Program; Google says developers, enterprises and consumers come next, once guardrails are refined (Google, Demis Hassabis). Standard pricing is $4/$20 per million input/output tokens; a 50% introductory discount with no announced end date brings it to $2/$10, and cached input is 95% cheaper (Philipp Schmid).
Gemini 4 Argon writes up to 1M tokens per answer, nearly 16x the old 64K cap
Google cites a 1-million-token output limit for Gemini 4 Argon, up from 64,000 (Google AI). Long Decode Continuation is an API feature that pauses a long response and resumes it across subsequent calls. That is how Artificial Analysis reached 1M output tokens, while Vals lists a 262K maximum — so the "1M" figure depends on how the model is called. In the Artificial Analysis run, Gemini 4 Argon scores 53 on the Intelligence Index, level with GPT-6 Astra and one point above GPT-6.1 Sol. A task costs $1.99 with Gemini 4 Argon versus $3.26 with Astra at discounted pricing, or $3.98 at standard pricing. The saving comes from price, not frugality: Argon averages 62,000 output tokens per task against Astra's 27,000.
Gemini 4 Argon agents freed 300 TiB of memory at Google and are moving 800K lines of C/C++ to Rust
Google says agents built on Gemini 4 Argon freed more than 300 TiB of data-center memory and are migrating more than 800,000 lines of C/C++ kernel code to Rust (summary of the report). In a video decoder, the agents replaced 32,000 lines of SIMD code with safe Rust, making the existing Rust port 2.7x faster with identical output. Google's team also says internal agent loops on Gemini 4 Argon helped complete the CK conjecture. These are Google's own claims and have not been independently verified. Ready-made agent workflows for your own tasks live in the AI SKILLS automation templates catalog.
OpenAI nears $70B in annualized revenue and talks about raising $30B at a $1.4T valuation
OpenAI is close to $70 billion in annualized revenue and in talks to raise $30 billion at a $1.4 trillion valuation, with its IPO pushed to next year, according to The New York Times, as relayed by reporter Sri Muppidi. The same week, OpenAI attributed the core of a campaign to extract its models' hidden reasoning to individuals linked to Moonshot AI: it recorded 16,000 attempts from more than 4,000 users in two days, with related activity across more than 15,000 users (summary). Outside researchers say their attacks kept working on GPT-6 Astra until this week and that patches were hard to propagate across product versions and third-party hosts (Jonas Geiping). Nathan Lambert argues the vulnerability is the API provider's responsibility.
Ethan Mollick: agents learned to work as a group on their own — the Bitter Lesson strikes again
On October 1, 2026, Ethan Mollick admitted a forecasting miss: he expected that getting AI agents to work as a group would require building something like a company, but newer models handle it themselves. In his essay The Dot and the Swarm he writes that he "fell prey to The Bitter Lesson" — the pattern in which elaborate human-designed rules keep losing to the scale of machine learning. Mollick draws the parallel with prompting: elaborate templates and prompt chains that walked a model through a task step by step lost their value once models learned to plan the steps themselves. The practical takeaway for agent builders: less rigid orchestration, more effort on stating the goal and checking the result. How to brief models today — see the AI SKILLS prompt library.
Numbers and facts
- +14.4 points of answer recall@10 for Perplexity's open embedding model pplx-embed-v2-context-9b-preview over voyage-context-4 on turbopuffer's private benchmark, using 1 KB int8 vectors instead of 8 KB (turbopuffer); the weights are open on Hugging Face.
- $12 per minute is the price of video from Wan 3.0, #1 on the new AA-Video-T2V v2.0 leaderboard (1080p, 68,000+ human votes); Seedance 2.5 is #2 at $34.12 per minute and MiniMax H3 is statistically tied at $4.80 (Artificial Analysis).
- 4.8x the token throughput of GB200 at the same decode speed is what Cognition reports as the first customer on Vera Rubin via CoreWeave (Cognition).
- 648 ms p50 time-to-interactive for Cloudflare's rebuilt Containers for agents, 6x faster than before, with snapshots in beta (Cloudflare).
- 416 cards on RAG, vector search and embeddings sit in the AI SKILLS catalog as of October 2, 2026: 238 automation templates, 120 Claude Code skills and 58 open-source projects (platform catalog data).
Different views: is Gemini 4 Argon a new leader or tuned to the tests?
Independent evaluations confirm that Gemini 4 Argon puts Google back among the leaders, but not alone at the top: Artificial Analysis scores it level with GPT-6 Astra (53 and 53), while Vals ranks it first on its index at 68.9%. Benchmaxxing is tuning a model to public tests so that leaderboard numbers run ahead of real-world usefulness. That is what some observers suspect of Gemini 4 Argon.
- For. Vals says Gemini 4 Argon built 30 Vibe Code Bench apps perfectly versus 25 for Opus 5 and 24 for Astra, and its Terminal-Bench 4.0 score rose from 19.0% to 57.6%. It is #1 in Text Arena at 1525.
- Neutral. Per Artificial Analysis, Gemini 4 Argon hallucinates on 15% of AA-Omniscience answers versus 51% for Astra, but its accuracy is also lower — 50% versus 63% (breakdown). The model is more cautious rather than more knowledgeable.
- Against. On Harvey's legal benchmark Gemini 4 Argon scores 19.6% against 25.42% for Muse Spark 1.2 (BlackHC); commentators raise possible tuning through preference data and dispute the DeepSWE figure. On Reddit, users note that DeepSWE is close to saturation and that parity with Astra on Terminal Bench 4 says more.
Tools and techniques
- Edit an image ten times in a row without drift. Ideogram 4.5 is an editing model built for artifact-free multi-turn edits, with open weights promised. Over ten consecutive edits, 94–99% of untouched content stays identical. The AI SKILLS prompt generator can draft the edit prompt.
- Index with one model, search with another. Cohere Embed 5 puts its Pro and Fast variants in a shared embedding space: index your corpus with the more accurate Pro and serve queries with the cheaper Fast. Cohere says Fast beats other fast-tier models by at least 6 points at a third less cost than Pro.
- Encode the whole document, not the chunks. Perplexity's pplx-embed-v2-context encodes the full document once and pools chunk vectors afterward, so every chunk carries the document's context. Open-source tools for RAG and document search are in the AI SKILLS Open Source section.
In brief
- Upstage released Solar Mini 4 — 35B total and 3B active parameters, 24 on the Artificial Analysis Intelligence Index at $0.10/$0.40 (Artificial Analysis); more on open models in our open-source LLM guide.
- Ling-3.1-flash, a 500B model reported close to GPT-5.6 Sol and Opus 5, ranks #2 among open-weight models in Mobile App Arena.
- DeepSeek released an open-source Huawei Ascend toolkit with TileLang optimized for Ascend 950.
- Greg Brockman dropped a promised second $25M donation to the Leading the Future super PAC.
- Flow, which builds AI tooling for hardware engineers, raised a $50M Series B at a $750M valuation.
- Factory removed advisor Chris Degnan, alleging he was confiding in Cognition; Cognition named Degnan its CRO the same day, and its CEO denies any information was shared.
- A Nature paper presents the first superhuman AI for Stratego, built on reinforcement learning and test-time compute under imperfect information.