Sonnet 5.5 and Opus 5.5 Refuse 4.5% and 8.9% — AI Digest

Sonnet 5.5 and Opus 5.5 Refuse 4.5% and 8.9% — AI Digest

Artificial Analysis exposes Claude refusals: Sonnet 5.5 drops 4.5% of tasks, Opus 5.5 drops 8.9%. MAI-Transcribe-2 hits 2.5% WER at $0.54 per hour, and OpenAI fires three safety researchers.

Top stories

Claude refusals: Sonnet 5.5 drops 4.5% of coding tasks, Opus 5.5 drops 8.9%, and Opus 4.8 quietly takes over

Artificial Analysis now shows, inside its Coding Agent Index, when a model refuses a task and which model the provider swaps in: Claude Sonnet 5.5 refuses 4.5% of tasks and Claude Opus 5.5 refuses 8.9%. According to the Artificial Analysis audit, roughly 94% of Sonnet 5.5 refusals happen after work has already started, and the task is usually finished by a fallback to Opus 4.8. A fallback model is the model a service automatically switches to when the primary model refuses or fails. The analysts' conclusion: an agent score describes the provider-configured system, not uninterrupted work by the model named on the leaderboard. If you pick coding agents from rankings, part of the "Sonnet 5.5" row was earned by a different model.

Personal agents: how to choose between Dots, Muse, GrokBot and OpenClaw

On October 4, 2026, The AI Daily Brief published a guide to choosing a personal AI agent, comparing Dots, Muse, GrokBot and OpenClaw along four axes. A personal agent is an AI with access to your accounts and its own computer that runs in the background and reaches out to you first. The episode, How to Choose Your Personal AI Agent, weighs work versus personal use, model choice, ease of setup and data privacy. Ethan Mollick gives a concrete example in One Useful Thing: Muse noticed an airline credit of his was about to expire and, when asked, contacted American Airlines to request an extension. Sam Altman calls dot his favorite OpenAI product. Meta is answering with an ecosystem: Muse connectors for small businesses and open-source ESP32 firmware plus a Linux SDK for third-party hardware. Ready-made agent workflows live in AI SKILLS automation templates.

MAI-Transcribe-2-Streaming: 2.5% word error rate at $0.54 per hour

Microsoft's MAI-Transcribe-2-Streaming ranks first among 38 speech-to-text models on Artificial Analysis, with a 2.5% word error rate and a final transcript 0.13 seconds after speech ends. Streaming costs $0.54 per hour, or $9 per 1,000 minutes, per the Artificial Analysis evaluation. Mustafa Suleyman pitches the model as "55% faster and 60% cheaper than ElevenLabs", yet the same ranking lists ElevenLabs' Scribe v2 Realtime at $6.50 per 1,000 minutes, which is cheaper than Microsoft's streaming tier. The published posts do not reconcile the comparison bases, so test the price claim on your own volume. In the same week, Inworld agreed to acquire Ultravox, combining speech understanding with its own voice generation stack.

Context Language Models: editable context gives 11.4% more accuracy with 21.5% less compute

Meta researchers introduced Context Language Models, which edit their own context as a file through Bash, and report 11.4% higher accuracy on BrowseComp-Plus with 21.5% fewer FLOPs. A standard model only appends to its context, while a CLM rewrites it, and the context-management policy is learned in the weights rather than in an external harness (@RulinShao). Editing the middle of the context breaks prefix caching, so the authors added Suffix Cache Reuse, which cuts server compute by 35% versus standard SGLang. On a 24-hour multi-repository agent-swarm task, CLMs scored 65% higher with the same compute.

AI safety: OpenAI fires three safety researchers, agents probe 55 government sites

According to the WSJ, OpenAI fired three safety researchers for allegedly sharing confidential information with an external AI safety organization. OpenAI confirmed the departures and cites mishandling outside established procedures; the posts do not establish what was shared (report summary). The same week, Transluce reported aggressive agent activity against US government websites and a previously undisclosed, apparently unsuccessful hacking attempt against a Canadian government site; FT reporting says temporary inboxes and intermediary scanning services complicated tracing activity across 55 websites. Meanwhile NVIDIA shipped OpenShell, an open-source sandbox that limits agents at the runtime level instead of through prompt rules, with more than 100 companies joining the platform.

Numbers and facts

Our roundup of infrastructure numbers for the week of September 28 to October 4, 2026:

WhatNumberSource
MAI-Transcribe-2 streaming transcription$9 per 1,000 minutes, 2.5% WERArtificial Analysis
AMD EPYC 9006 Venice (Zen 6)about $700 to $14,904, up to 256 coresr/LocalLLaMA
Lambda GPU debt facility$1B+, investment-grade ratedLambda
Agent-written GPU kernels (Databricks)#1 on all 4 SOL-ExecBench tracks for ~$70K in tokens@Yuchenj_UW
Moondream Photon 2.6Qwen3.5 27B at 400+ tokens/s on B200@vikhyatk
Cloudflare AutoRouter model routerabout 30% lower spend in internal tests@ashleypeacock

Different views: was OpenAI right to fire its safety researchers?

According to the WSJ, OpenAI fired three safety researchers for sharing confidential information with an external AI safety organization; what was shared has not been disclosed, so the debate is about procedures rather than facts.

For. OpenAI, per the WSJ summary, frames the case as mishandling of confidential information outside established procedures, not as punishment for caring about safety.

Neutral. Josh Achiam called for details before drawing firm conclusions and argued that procedures should leave room for potential whistleblowing.

Against. John Schulman advocated greater research transparency, meaning that information about risks should be available beyond the lab's walls.

Tools and tips

In brief