Opus 5.5 Leads Coding Agents at $13 a Task — AI Digest

Claude Opus 5.5 tops the Coding Agent Index at 66 points but costs $13.04 a task; GPT-6 Luna scores 37 for $0.068, and Meta unveils $1,299 VR glasses.
Top stories
Opus 5.5 leads the Coding Agent Index at 66, but a single task costs $13.04
Claude Opus 5.5 took first place on Artificial Analysis's Coding Agent Index with 66 points, up from 60 for Opus 5, while one index task now costs $13.04 to run. The Coding Agent Index is a composite score built from several coding-agent benchmarks. According to Artificial Analysis, Opus 5.5 scored 63.1% on Terminal-Bench 4.0, 68.4% on DeepSWE v1.1 and 66.4% on SWE-Atlas-QnA. Token prices fell to $4/$20 per million input/output tokens, with cache reads at $0.20. Per-task cost still went up, because Opus 5.5 burns 15.6M tokens per task and more than doubles its output tokens. Separately, the model set a record 2631 Elo on a writing benchmark, 307 points clear of the next model — though a max-effort run takes 17 minutes and $3.43 per script. Ready-made skills for coding agents are collected in our Claude Code skills guide.
The cheap end of the frontier: GPT-6 Luna scores 37 for $0.068 a task
At the low end of Artificial Analysis's price–performance curve sit GPT-6 Luna at 37 Intelligence Index points for $0.068 a task, MiMo-V2.6-Pro at 46 for $0.13 and GPT-6 Sol at 48 for $1.06. In the same Artificial Analysis run, Opus 5.5 reaches 58. The Pareto frontier is the set of models where you cannot get a higher score without paying more. Independent evaluator Vals puts GPT-6 Luna at $0.10/$0.50 per million tokens — roughly 100x cheaper than GPT-6 Astra per token — while landing within 8 points on the Vals Index. GPT-6 Luna has a 1M-token context window and up to 128k output tokens. The same shift shows up inside companies: per @Yuchenj_UW, Databricks engineers stopped reaching for closed models once their internal coding agents started routing work to open models. How to run open models yourself is covered in our open-source LLM guide.
Meta Connect 2026: $1,299 VR glasses, an FDA-cleared hearing aid and the WaveForms deal
At Connect 2026, Meta showed its first VR experience delivered in glasses rather than a headset, priced at $1,299, and turned its glasses into an FDA-cleared hearing aid. In Mark Zuckerberg's announcement, Meta VR Glasses are pitched as a private cinema, a multi-monitor workstation and a game console; the $1,299 price was reported by @kimmonismus. The glasses now work as an FDA-cleared hearing aid. Ray-Ban Meta Gen 3 brings longer battery life, upgraded microphones and new styles, including Aviators. Meta also acquired WaveForms AI, Alexis Conneau's speech and audio startup, whose work already surfaced at Connect. A new frontier model was only teased: Alexandr Wang said that "the most capable model we have ever trained" is coming soon.
Numbers and facts
- 100K+ H100-hours went into building OpenRSI-Index v0.1, which runs 60+ hour autonomous research trajectories on 1,000-GPU clusters.
- Four efficient architectures compared by @eliebakouch: DeepSeek V4.1 Flash and MiMo V3 use YOCO, while Qwen 3.8 Next Flash and GLM 5.3 Flash interleave sparse and linear attention 3:1.
- 1K–10K tokens is the context range where Jev became the top model on OpenRouter.
AI SKILLS summary table: cost per task vs Intelligence Index, September 2026
| Model | Intelligence Index | Cost per task |
|---|---|---|
| GPT-6 Luna | 37 | $0.068 |
| MiMo-V2.6-Pro | 46 | $0.13 |
| GPT-6 Sol | 48 | $1.06 |
| Claude Opus 5.5 | 58 | $5.98 |
Compiled by our editors from Artificial Analysis's run. Moving from GPT-6 Sol to Opus 5.5 buys 10 more points at 5.6x the cost per task.
Different views: does the world need a global AI regulator?
The AI regulation debate in September 2026 has split three ways: urgent global action, rejection of a world regulator, and "moderates" who refuse to choose between optimism and doom. The AI Daily Brief episode of 27 September 2026, "The Rise of the AI Moderates", describes a growing chorus rejecting the choice between unchecked optimism and inevitable doom.
For urgent action. Yoshua Bengio urged immediate action at the UN Security Council session on AI.
Neutral. The AI Daily Brief draws on Francis Fukuyama's essay on why he changed his mind about AI risk, the "AI as normal technology" view and a post by Jeffrey Katzenberg; the discussion covers AI risk, human creativity and who gets a say in shaping AI.
Against a global regulator. Kratsios rejected the idea of a global AI regulator.
Tools and techniques
- Try Jev as a reranker. The turbopuffer search database's native reranking now includes Jev, a model that returns a typed decision with a probability instead of reasoning text. RAG workflow templates are in the AI SKILLS automation templates catalog.
- Test interpretability tools on WorkspaceBench. Neel Nanda introduced the benchmark for evaluating such tools.
- Learn where GRPO came from. @cwolferesearch traces the reinforcement-learning lineage from VPG through REINFORCE and PPO to GRPO and its variants — a useful map before reading new RL papers.
In brief
- Gemini 4 is reportedly nearly finished training, per @kimmonismus.
- Redwood Research argues that latent "neuralese" reasoning in models would erode chain-of-thought oversight.
- Joby claims the first fully autonomous flight across the US, Superhuman AI reported on 27 September 2026.