GPT-6.1 Sol, Argon: $2/$10 per 1M Tokens — AI Digest

OpenAI prices GPT-6.1 Sol at $2/$10 per 1M tokens, Google opens Gemini 4 Argon to a narrow group, and Claude Sonnet 5.5 holds $2/$12 while ranking #3 in Agent Arena.
Top stories
GPT-6.1 Sol: $2/$10 per 1M tokens, $0.72 per task against Astra's $3.26
OpenAI's GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, one fifth of GPT-6 Astra's $10/$50. At its DevDay on September 29, 2026, OpenAI pitched the model as near-Astra intelligence for a fifth of the price. Artificial Analysis then measured $0.72 per Intelligence Index task at maximum effort, against $1.04 for GPT-6 Sol and $3.26 for Astra. Its explanation: fewer turns and cheaper cache reads, not fewer generated tokens. OpenAI also fixed an image-encoding bug, which added one Index point to GPT-6 Luna. Sam Altman called Sol the company's fastest-growing model.
OpenAI Dots: always-on agents with their own cloud computer
OpenAI's Dots are persistent personal agents that message you first and connect to more than 4,000 apps. Each dot runs on GPT-6 Astra with its own cloud computer, and users set what it may do alone, what needs approval and what is off limits. Ben's Bites describes a Dot as a never-ending orchestrator chat that spins up Codex and ChatGPT Work tasks and reports back; Pro plans get one. Ethan Mollick, writing in One Useful Thing, groups Dots with Meta's Muse as "Clawlikes", successors to OpenClaw, and says his own agents increasingly catch his mistakes rather than the reverse. He adds that Muse currently tops the App Store.
Gemini 4 Argon: Google claims 13 of 19 benchmark wins, access stays restricted
Google DeepMind's Gemini 4 Argon lists at $4/$20 per 1M tokens, halved to $2/$10 by an introductory discount with no announced end date. Access starts with government users and vetted cyber defenders in the Fairwind Program, according to Google DeepMind and The Rundown AI (October 1, 2026). Google says Argon leads 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5 and scores 77.9% on DeepSWE versus 74.2% for Opus 5.5, with output of up to 1M tokens. Artificial Analysis gives it 53 on its Intelligence Index, level with Astra.
Claude Sonnet 5.5: unchanged $2/$12 pricing, 30% faster, #3 in Agent Arena
Anthropic released Claude Sonnet 5.5 on September 28, 2026 at Sonnet 5's price of $2/$12 per 1M tokens, claiming over 30% more speed and up to 30% lower cost per task. Anthropic calls it a clear upgrade over Sonnet 5. Vals AI lists a 1M-token context window and up to 128K output. In Agent Arena the model debuted at #3 and first in the Chat category, at $2.74 per task against $1.58 for Opus 5.5. One caveat: at maximum effort Artificial Analysis recorded about 193K output tokens per task, the most it has measured. Anthropic staff advise against running Sonnet at max effort.
Decision models: Cloudflare, Perplexity and OpenAI turn classification into its own layer
Three vendors shipped decision models in one week: fast models that pick one answer from a fixed list. A decision model is a small model that takes a question plus fixed answer options and returns probabilities over them instead of free text. OpenAI launched the Decisions API on GPT-6 Luna for near-instant classification and routing. Perplexity released pplx-decider-v1-27b, built on Qwen3.8-27B, at $0.04 per million input tokens, free output and open weights. Cloudflare's two Clef models come with Apache 2.0 weights. Skeptics such as @mervenoyann call the category a rebrand of zero-shot classifiers.
Numbers and facts
Our price table for this week's models (per 1M input / output tokens, from the sources cited in this issue):
| Model | Input / output | Note |
|---|---|---|
| GPT-6.1 Sol | $2 / $10 | $0.72 per task (Artificial Analysis) |
| Claude Sonnet 5.5 | $2 / $12 | 1M context, 128K output |
| Gemini 4 Argon | $4 / $20 | $2 / $10 with the 50% promo; $1.99 per task |
| GPT-6 Astra | $10 / $50 | $3.26 per task |
| GPT-6 Astra Ultrafast | $60 / $300 | six times the normal price, up to 300 tokens/s |
| Solar Mini 4 | $0.10 / $0.40 | 35B total / 3B active, closed weights |
- ChatGPT plans: Pro 100 gives 5x Plus usage, Pro 200 gives 10x and the new Pro 500 gives 25x. Pro 200 subscribers lost roughly half their old value and pushed back loudly.
- Vals Index: Argon ranks first at 68.9% and $15.68 average cost per task. On Vibe Code Bench it built 30 apps perfectly, against 25 for Opus 5 and 24 for Astra.
- Hugging Face: identical weights score 62% in one harness and 33% in another, so agent results depend on the harness as well as the model.
Different views: is Gemini 4 Argon the new frontier leader?
Argon tops the Vals Index at 68.9% and places in the top five on 20 of 22 Vals benchmarks, yet Epoch AI ranks Opus 5.5 first on its ECI at 167, with Sonnet 5.5 roughly level with Fable 5.1 at 165. Those are different test suites, not interchangeable leaderboards.
For. Vals puts Argon first on its aggregate index, and Google says new Gemini revisions now go through weeks of testing by thousands of internal engineers. Argon also took first place in Text Arena.
Neutral. Artificial Analysis scores Argon level with Astra at 53, but the savings come from price: Argon spends 62K output tokens per task against Astra's 27K.
Against. A circulated Bloomberg report blamed weak coding on anonymous insiders; a senior DeepMind engineer reportedly rejected that account. On Harvey's legal benchmark Argon scored 19.6% against 25.42% for Muse Spark 1.2, and other commentators suspect tuning to preference data. Real-world coding quality remains disputed.
Tools and techniques
- Claude Code mods. Claude Code now supports TypeScript mods that change behavior, UI and features and ship as plugins; Anthropic used the mechanism for
/diffandAGENTS.mdsupport. If you build your own skills, start with our [Claude Code skills guide](/en/blog/claude-code-skills-guide) and the AI SKILLS skill catalog. - Sonnet 5.5 instead of Opus for routine work. Anthropic positions it for everyday tasks like fixing bugs; @edwinarbus says to avoid max effort, and Claude Code defaults to medium.
- A decision model on your laptop. llama.cpp added a
/v1/systemoneendpoint, and the model starts withllama serve -hf ggml-org/Kev-4B-GGUFper @ClementDelangue. It suits classification steps in workflows; ready-made ones are in our automation templates.
In brief
- FLUX 3 Image from Black Forest Labs generates up to 4K from up to ten references, with a 50% API discount through October 8, 2026; see the announcement.
- Tavus Griffin: the company says 48% of live participants mistook it for a human, versus under 3% for earlier systems.
- Meta Muse Home Link: 5,000 units free for subscribers while supplies last.
- SWE-sweep covers 100 repositories, 22 languages and about 4,000 real bugs; leading models solve under 5%.
- Pi 1.0 shipped with Pi Durable, making durable execution a harness primitive.
- Claude Haiku 5.5 will round out the family in the coming weeks, according to @mikeyk.