Claude Opus 5.5 Runs 40% Cheaper Than Opus 5 — AI Digest

Claude Opus 5.5 Runs 40% Cheaper Than Opus 5 — AI Digest

Anthropic shipped Claude Opus 5.5 at Fable 5.1 quality and 40% lower running cost, and OpenAI answered the same day with GPT-6 Sol and Luna at $2/$10 and $0.10/$0.50 per million tokens. Xiaomi released MiMo-V2.6-Pro, 1.02T parameters under MIT.

Today's highlights

Claude Opus 5.5 claims Fable 5.1 quality at 40% lower running cost

Anthropic shipped Claude Opus 5.5 with the claim that it performs at the level of Claude Fable 5.1 "for most tasks" while costing 40% less to run than Opus 5. The launch post of 22 September 2026 calls it the first model in a new Claude 5.5 family, and Anthropic said it was available the same day in the Claude app and in Claude Code. Benchmark vendors recorded the token price dropping from $5/$25 to $4/$20 per million input and output tokens, with a 1M context window and 128K maximum output. For paid tiers the model became the default in Claude Code and the app, five-hour session limits in Claude Code went up by 20%, and the lower price stretches the same allowance roughly a quarter further. Anthropic also claimed Fable 5.1-class safeguards on cyber and bio risk, routing flagged requests to a different model. Sonnet 5.5 and Haiku 5.5 are promised within weeks.

Independent evaluations confirmed the claim in part. Cursor put the model first on CursorBench at 57.8%, Cognition first on FrontierCode 1.1 Extended at 65.3%, and Vals first on its RSI Index. On FrontierSWE it lands second at 62.3%, behind GPT-6 Astra at 65.5% but ahead of Fable 5.1 at 56.3% and Opus 5 at 52.0%. On table parsing it scored 93.9% on ParseBench, seven points above Opus 5, at 5.8 cents per page. The other half of the reaction was about prose: Anthropic said the model communicates more naturally and front-loads what matters, and one of its researchers put it bluntly — "we fixed the writing."

GPT-6 Sol and Luna: OpenAI answered the same day at $2/$10 and $0.10/$0.50

OpenAI released GPT-6 Sol and GPT-6 Luna hours after Anthropic's launch: faster, cheaper models that inherit GPT-6 Astra's advances. The announcement prices Sol at $2 per million input and $10 per million output tokens and Luna at $0.10 and $0.50 — each roughly half the price of its GPT-5.6 predecessor. Both rolled out to ChatGPT Work, Codex and the API, with Luna also reaching Free and Go users in the desktop app, though not yet in ordinary chat. OpenAI framed the comparison around cost per task rather than absolute capability: in its own examples, Sol at maximum effort beats Claude Opus 5 on AutomationBench at about 9% of the cost per task. Effort here is a declared reasoning budget: the higher it is set, the more tokens the model spends deliberating before it answers.

Partners confirmed the economics quickly. Cognition reported that Sol matches GPT-5.6 Sol at 61% lower cost per task and that Luna beats its predecessor at roughly a quarter of the price. Perplexity made Sol the default orchestrator at its light effort level. The infrastructure half of the release may matter more than the models: OpenAI announced discounts of up to 90% on cached input token reads and shipped a prompt caching dashboard with diagnostics for why reuse broke. For long-running agents that is money: cache invalidation when tools or settings change is what quietly consumes the budget of long runs.

Xiaomi MiMo-V2.6-Pro: 1.02T parameters under MIT and a $2.6M RL run

Xiaomi released MiMo-V2.6-Pro, an open-weights model with 1.02 trillion total and 42 billion active parameters, under the MIT license. The release covers not only weights and a technical report but the training environments and reinforcement-learning code as well. Artificial Analysis ranked it first among open-weights models on its Intelligence Index at 46, noting prices of $0.435 per million input and $0.87 per million output tokens. The MIT license was confirmed independently, which is unusual at this scale: trillion-parameter open weights normally arrive with conditions attached.

The sharper discussion was about training cost rather than scores. According to figures circulating with the release, the final RL run took 130 hours, 75 billion tokens and $2.6 million — an order of magnitude below what is usually assumed for this class. The team says it scales reinforcement learning on JAX and TPUs, where moving to a larger scale is "mostly a config change, not a code rewrite." If the numbers hold, the practical conclusion is that post-training is becoming a cheaper route to frontier-adjacent quality than pre-training, and open weights will keep closing the gap. Deployable open projects are collected in the AI SKILLS open-source section, and the wider landscape is covered in our guide to open-source LLMs.

Anthropic's system card reports multi-agent scaling to 100 parallel agents

Anthropic published system-card measurements of multi-agent scaling for the first time, covering up to 100 agents working in parallel on one task. An independent analyst flagged it as the first such disclosure in a lab system card, and a safety researcher pointed to section 8.12, where the results read like scaling laws applied to systems of agents rather than to a single model. The practical significance is that running agents in parallel stops being a hobbyist trick and becomes a measurable configuration with predictable returns.

The objection arrived immediately and on the merits. The author of ProgramBench argued that Anthropic scored a 166-task subset out of 200 and, in his reading, left out the hardest ones; he later separated the metrics — Anthropic's "average total test pass rate" against his team's "average task completion." This is the pattern of the season: the argument is not about whether the model works but about what exactly was counted. Claims like these have to be checked on your own tasks, which is precisely why the harness and the metric matter more than the line in the table.

Numbers and facts

Our price table for the day (announced prices per million tokens, input/output):

ModelInputOutputWeights
Claude Opus 5.5$4$20closed
GPT-6 Sol$2$10closed
GPT-6 Luna$0.10$0.50closed
MiMo-V2.6-Pro$0.435$0.87open, MIT

Different views: did the frontier actually get cheaper?

In a single day the input token price fell at both leading vendors — from $5 to $4 at Anthropic, by about half against the previous generation at OpenAI — yet the cost of a solved task did not fall everywhere. Artificial Analysis's index accounting puts Opus 5.5 at $5.98 against Opus 5's $5.86: the model spends more tokens, especially on coding, and the discount goes into that increase.

For. Cost per task is measurably down at the partners: 61% lower for Cognition on Sol, 67.6% lower for Perplexity on WANDR, and a roughly ninefold gap in OpenAI's own AutomationBench examples. Where tasks are short and the cache is warm, the saving is real and shows up on the invoice.

Neutral. Cache-read discounts move the picture more than headline prices do: OpenAI took its discount to 90% and shipped cache diagnostics, and in Artificial Analysis's accounting caching was the main offset against rising token use. The conclusion is to price your own workload profile rather than the rate card.

Against. More effort is not better and is almost always more expensive: one published chart showed maximum effort costing 2.8x more than medium for a 3.2-point lower score on agentic coding. A practitioner summarised it in one line: stay on medium.

Tools and techniques

In brief