Claude Opus 5.5 Runs 40% Cheaper Than Opus 5 — AI Digest

Anthropic shipped Claude Opus 5.5 at Fable 5.1 quality and 40% lower running cost, and OpenAI answered the same day with GPT-6 Sol and Luna at $2/$10 and $0.10/$0.50 per million tokens. Xiaomi released MiMo-V2.6-Pro, 1.02T parameters under MIT.
Today's highlights
Claude Opus 5.5 claims Fable 5.1 quality at 40% lower running cost
Anthropic shipped Claude Opus 5.5 with the claim that it performs at the level of Claude Fable 5.1 "for most tasks" while costing 40% less to run than Opus 5. The launch post of 22 September 2026 calls it the first model in a new Claude 5.5 family, and Anthropic said it was available the same day in the Claude app and in Claude Code. Benchmark vendors recorded the token price dropping from $5/$25 to $4/$20 per million input and output tokens, with a 1M context window and 128K maximum output. For paid tiers the model became the default in Claude Code and the app, five-hour session limits in Claude Code went up by 20%, and the lower price stretches the same allowance roughly a quarter further. Anthropic also claimed Fable 5.1-class safeguards on cyber and bio risk, routing flagged requests to a different model. Sonnet 5.5 and Haiku 5.5 are promised within weeks.
Independent evaluations confirmed the claim in part. Cursor put the model first on CursorBench at 57.8%, Cognition first on FrontierCode 1.1 Extended at 65.3%, and Vals first on its RSI Index. On FrontierSWE it lands second at 62.3%, behind GPT-6 Astra at 65.5% but ahead of Fable 5.1 at 56.3% and Opus 5 at 52.0%. On table parsing it scored 93.9% on ParseBench, seven points above Opus 5, at 5.8 cents per page. The other half of the reaction was about prose: Anthropic said the model communicates more naturally and front-loads what matters, and one of its researchers put it bluntly — "we fixed the writing."
GPT-6 Sol and Luna: OpenAI answered the same day at $2/$10 and $0.10/$0.50
OpenAI released GPT-6 Sol and GPT-6 Luna hours after Anthropic's launch: faster, cheaper models that inherit GPT-6 Astra's advances. The announcement prices Sol at $2 per million input and $10 per million output tokens and Luna at $0.10 and $0.50 — each roughly half the price of its GPT-5.6 predecessor. Both rolled out to ChatGPT Work, Codex and the API, with Luna also reaching Free and Go users in the desktop app, though not yet in ordinary chat. OpenAI framed the comparison around cost per task rather than absolute capability: in its own examples, Sol at maximum effort beats Claude Opus 5 on AutomationBench at about 9% of the cost per task. Effort here is a declared reasoning budget: the higher it is set, the more tokens the model spends deliberating before it answers.
Partners confirmed the economics quickly. Cognition reported that Sol matches GPT-5.6 Sol at 61% lower cost per task and that Luna beats its predecessor at roughly a quarter of the price. Perplexity made Sol the default orchestrator at its light effort level. The infrastructure half of the release may matter more than the models: OpenAI announced discounts of up to 90% on cached input token reads and shipped a prompt caching dashboard with diagnostics for why reuse broke. For long-running agents that is money: cache invalidation when tools or settings change is what quietly consumes the budget of long runs.
Xiaomi MiMo-V2.6-Pro: 1.02T parameters under MIT and a $2.6M RL run
Xiaomi released MiMo-V2.6-Pro, an open-weights model with 1.02 trillion total and 42 billion active parameters, under the MIT license. The release covers not only weights and a technical report but the training environments and reinforcement-learning code as well. Artificial Analysis ranked it first among open-weights models on its Intelligence Index at 46, noting prices of $0.435 per million input and $0.87 per million output tokens. The MIT license was confirmed independently, which is unusual at this scale: trillion-parameter open weights normally arrive with conditions attached.
The sharper discussion was about training cost rather than scores. According to figures circulating with the release, the final RL run took 130 hours, 75 billion tokens and $2.6 million — an order of magnitude below what is usually assumed for this class. The team says it scales reinforcement learning on JAX and TPUs, where moving to a larger scale is "mostly a config change, not a code rewrite." If the numbers hold, the practical conclusion is that post-training is becoming a cheaper route to frontier-adjacent quality than pre-training, and open weights will keep closing the gap. Deployable open projects are collected in the AI SKILLS open-source section, and the wider landscape is covered in our guide to open-source LLMs.
Anthropic's system card reports multi-agent scaling to 100 parallel agents
Anthropic published system-card measurements of multi-agent scaling for the first time, covering up to 100 agents working in parallel on one task. An independent analyst flagged it as the first such disclosure in a lab system card, and a safety researcher pointed to section 8.12, where the results read like scaling laws applied to systems of agents rather than to a single model. The practical significance is that running agents in parallel stops being a hobbyist trick and becomes a measurable configuration with predictable returns.
The objection arrived immediately and on the merits. The author of ProgramBench argued that Anthropic scored a 166-task subset out of 200 and, in his reading, left out the hardest ones; he later separated the metrics — Anthropic's "average total test pass rate" against his team's "average task completion." This is the pattern of the season: the argument is not about whether the model works but about what exactly was counted. Claims like these have to be checked on your own tasks, which is precisely why the harness and the metric matter more than the line in the table.
Numbers and facts
- $5.98 against $5.86 — the cost of one Intelligence Index task for Opus 5.5 versus Opus 5 at maximum settings, by Artificial Analysis's accounting: the lower token price is almost entirely offset by higher token use.
- ~47% per quarter is how fast the cost of AI at a fixed level of performance has been falling since 2023, according to Epoch AI.
- 0.610 at $4.13 per task — Opus 5.5 on Perplexity's own WANDR evaluation, which the company says is 67.6% cheaper per task than Fable 5.1.
- 21.2% — the reduction in failed tool calls Perplexity reported in a live A/B test after post-training its agent on real user sessions.
- 80.9% on Terminal-Bench 2.1 — the claimed score of Step Code v0.1.0, a coding agent CLI released under the MIT license.
- 1.7% WER — StepAudio 3 ASR on the AA-WER index, effectively sharing the top spot among non-streaming speech-to-text systems.
- 0.73s to first token and 6.9 mWh per inference — Reka EdgeQ, a vision-language model running directly on Qualcomm's Hexagon NPU with the GPU left idle.
Our price table for the day (announced prices per million tokens, input/output):
| Model | Input | Output | Weights |
|---|---|---|---|
| Claude Opus 5.5 | $4 | $20 | closed |
| GPT-6 Sol | $2 | $10 | closed |
| GPT-6 Luna | $0.10 | $0.50 | closed |
| MiMo-V2.6-Pro | $0.435 | $0.87 | open, MIT |
Different views: did the frontier actually get cheaper?
In a single day the input token price fell at both leading vendors — from $5 to $4 at Anthropic, by about half against the previous generation at OpenAI — yet the cost of a solved task did not fall everywhere. Artificial Analysis's index accounting puts Opus 5.5 at $5.98 against Opus 5's $5.86: the model spends more tokens, especially on coding, and the discount goes into that increase.
For. Cost per task is measurably down at the partners: 61% lower for Cognition on Sol, 67.6% lower for Perplexity on WANDR, and a roughly ninefold gap in OpenAI's own AutomationBench examples. Where tasks are short and the cache is warm, the saving is real and shows up on the invoice.
Neutral. Cache-read discounts move the picture more than headline prices do: OpenAI took its discount to 90% and shipped cache diagnostics, and in Artificial Analysis's accounting caching was the main offset against rising token use. The conclusion is to price your own workload profile rather than the rate card.
Against. More effort is not better and is almost always more expensive: one published chart showed maximum effort costing 2.8x more than medium for a 3.2-point lower score on agentic coding. A practitioner summarised it in one line: stay on medium.
Tools and techniques
- Rewrite your prompts for the new generation. Anthropic published direct guidance for Opus 5.5: hand over the whole task rather than step-by-step instructions, define explicitly what "done" means, and drop "think carefully" — it no longer helps. Developers note that older prompting tricks may now clash with the training, so accumulated templates deserve re-testing instead of blind reuse. Ready-made prompts for specific jobs live in the AI SKILLS prompt library.
- Change the effort level mid-session without fear. On Claude Code 2.1.280 and later, switching effort does not break the prompt cache — so a long run can sit on medium and raise effort only for the hard step, without paying to rebuild the whole context.
- Compress the model, not the budget. Tim Dettmers opened a dynamic compression framework inside bitsandbytes2, aiming at 1.5–2.0 bits per weight with a "lazy" mode that finds the trade-off itself between memory, quality and speed at deployment time. It addresses the real obstacle to running trillion-parameter open weights locally: the growing KV cache.
In brief
- Qwen-Image-2.1 took first place among open models in both the image-editing and text-to-image arenas — per the leaderboard.
- Ming-Image-0.1-Design, a 6B open-weight image design family, launched with a claim of first place among open models on a UI/UX design leaderboard — announcement.
- Moondream released Parakeet Redux and Parakeet Ultra, local speech-to-text models covering 25 languages — announcement.
- fal published the stack behind H3 Max, which generates 5 seconds of video in 3 seconds — optimisation breakdown.
- DigitalOcean opened a public preview of managed agents supporting Claude Code, Codex and a choice of 75+ models — announcement.
- VS Code added an experimental Agent Merge mode in which the agent resolves review comments, failed checks and merge conflicts inside the pull request — announcement.
- NVIDIA DGX Spark: 128 GB of unified memory in a 1.2 kg box, with claimed local inference for models up to 200B parameters — hands-on.
- Amazon blocked Meta's Muse shopping agent — reported by The Rundown.
- OpenAI says its models solved 100 open mathematics problems — covered by Superhuman AI.
- Anthropic's new model showed a marked jump in 3D work: a company researcher describes a serious step up in 3D understanding, and an independent reviewer found code-generated scenes on par with GPT-6 Astra.