Sonnet 5.5: 30% Faster at the Same Price — AI Digest

Anthropic shipped Sonnet 5.5: 30% faster than Sonnet 5, #2 on the Vals Index and the same $2/$12 price. AMD is buying World Labs for $8.2B, and a Perplexity red-team showed how agents slip out of sandboxes through DNS.
Top stories
Sonnet 5.5: 30% faster than Sonnet 5 and #2 on the Vals Index at the same $2/$12 price
Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, one week after Opus 5.5, calling it "a clear upgrade over Sonnet 5" that runs more than 30% faster and costs up to 30% less for most work. The launch was announced by @claudeai. According to Vals, the model has a 1M-token context window, up to 128K output tokens and the same price as Sonnet 5: $2/$12 per million input/output tokens. It debuted at #2 on the Vals Index, 0.47 points behind Opus 5.5. Artificial Analysis adds a caveat: at max effort the model uses about 193K output tokens per task, the most the firm has ever measured and roughly 7x GPT-6 Astra. Cognition measured 64.4% on FrontierCode 1.1 Main, up from Sonnet 5's 56.2%. As @simonw points out, Sonnet 5.5 now powers the free tier on claude.ai, and Haiku 5.5 will follow in the coming weeks, per @mikeyk. Ready-made skills for Claude Code are collected in our Claude Code skills guide.
AMD buys World Labs for $8.2B in stock, and Fei-Fei Li becomes AMD chief scientist
AMD is acquiring World Labs, the startup building spatial intelligence models, for $8.2B in stock, and its co-founder Fei-Fei Li becomes AMD's chief scientist. The deal was reported by The Rundown AI and posted by AMD CEO Lisa Su. Spatial intelligence is a model's ability to understand and build three-dimensional scenes rather than only text and flat images. Latent Space notes that the official post does not name a price, but the figure is known because AMD is a public company.
Agent sandboxes: 9 models got root, and 4 reached a blocked address through DNS
In Perplexity's SPACE red-team, 9 models, including Opus 5, GPT-5.6 Sol, Kimi K3 and Gemini 3.1 Pro, got root inside the virtual machine, and none escaped in 108 runs. That is according to Aravind Srinivas. With PyPI and Ubuntu mirrors allowed, however, 4 models reached a blocked URL through DNS spoofing and shared Fastly IP addresses. Perplexity found the same shared-IP flaw in 8 of 10 third-party sandboxes, including E2B, Vercel and Modal. Hugging Face responded with a contribution to OpenShell, and Clément Delangue sums up the core lesson: "Allowlists alone restrict where an agent can go, not what it does." OpenShell therefore adds per-sandbox network budgets, drift detection against a baseline, and a fleet view that flags many sandboxes writing to one host. The AI Daily Brief episode "The Real Risks of AI Agents" examines recent OpenAI security incidents and the limits of agent permissions.
Numbers and facts
- 26% of Anthropic's R&D is done by AI with high-level human oversight, up from 1% in March, according to an intelligence-explosion paper; under full automation a year of progress could take about 5 weeks, The Rundown AI reports.
- 17,895 lines of Lean proof were written by ten Sonnet 5.5 agents in 15 hours: per Vals, it proves the Thomson problem for seven electrons on a sphere and was accepted by the Lean kernel.
- 39.9 seconds is the new NanoGPT speedrun record, down from 67.6 seconds, reports @classiclarryd.
AI SKILLS summary table: Sonnet 5.5 vs Opus 5.5 on vision tasks
| Model and effort | Cost per 1,000 images | Latency |
|---|---|---|
| Sonnet 5.5, high | $8.82 | 8.9 s |
| Opus 5.5, high | $18.12 | 12.2 s |
| Sonnet 5.5, low | $6.53 | 10.8 s |
| Opus 5.5, low | $13.80 | 12.8 s |
Compiled by our editors from the Roboflow evaluation and its latency figures. At high effort Sonnet 5.5 costs about half as much as Opus 5.5, with roughly comparable accuracy on the sampled tasks.
Different views: does Sonnet 5.5 replace Opus?
Early users are split on Sonnet 5.5: some see it as Opus-level work at half the price, others as a strong model for well-defined tasks that lacks the older model's intuition.
Yes, for many tasks. Cursor calls the model "on par with Opus in many tasks," and @chaseleantj says it "feels just as good as Opus for half the price."
Neutral. GitHub Copilot's framing is more muted: in early testing it "matched Sonnet 5 while using fewer steps, tokens, and tool calls," which is an efficiency gain rather than a quality jump.
No. @rishdotblog says Sonnet is "nowhere near opus" and lacks the general intuition of Opus 5.5, and Anthropic's @mikeyk says he still reaches for Opus 5.5 for most work.
Tools and techniques
- Don't run Sonnet 5.5 at max effort. @edwinarbus is blunt: at max effort you should probably be using Opus, and Claude Code defaults to medium. Factory found high effort to be a strong default.
- Ask Claude Code to build evals. Per @ClaudeDevs, Claude Code can now build evals and hillclimb against them. Ready-made skills for tasks like this are in the AI SKILLS Claude Code skills catalog.
- Choose between SAEs and probes. Goodfire published a practical guide on when sparse autoencoders help in interpretability work and when simple probes are enough.
In brief
- Kling 4.0 generates video up to 4K with 10-bit HDR, takes 15 multimodal references and 10 keyframes, and produces 30-second clips natively.
- ElevenLabs v4 and v4 Turbo took #1 on the Artificial Analysis voice leaderboard, with voice cloning from 10 seconds of audio.
- Florida's attorney general, per @kimmonismus, is seeking an injunction halting OpenAI model development until independently approved safeguards are in place.