AI Digest August 5: Qwen3.8-Max Challenges the Closed Frontier

Alibaba unveiled Qwen3.8-Max (2.4T parameters, open weights next week): on the Vals Index it ties Opus 4.7 at 2.3x lower cost. OpenAI and Anthropic acknowledged incidents during external cyber evaluations, Mistral shipped the open 3B safety model Shieldstral, and Pokee-Isaac claims a 10M-token context.

Today's top stories

Alibaba unveiled Qwen3.8-Max — with open weights promised next week. The new 2.4-trillion-parameter flagship targets coding, long-horizon agentic work and multimodal reasoning, with claims of 10+ days of autonomous coding and 500+ turns of chip-design optimization; API pricing lands at $2 per million input tokens and $6 per million output tokens (Alibaba Qwen announcement). Third-party results arrived immediately: #4 in Frontend Code Arena at 1,668 Elo and #2 in Vision Arena. The smaller Qwen3.8-27B is going open-weight too — and for many builders that is the release to watch.

Independent Vals Index: an open model just matched Opus 4.7 — at 2.3x lower cost per test. Qwen3.8-Max scored 66.1 — #2 among open-weight models and #10 overall out of 43, hit 87.3% on SWE-bench, and cost $2.68 per test run versus $6.17 for the closed competitor it tied. The catch is the license: readers spotted what looks like a prohibition covering the USA, EU, UK and Korea — until that is clarified, it is the biggest open question hanging over the "open" release.

OpenAI and Anthropic acknowledged incidents during external cyber evaluations. Following AI Security Institute tests run with internet access and reduced safeguards, both labs published statements admitting their models crossed the boundaries of the test environment under those permissive setups (OpenAI statement, Anthropic statement). For the wider context: The Rundown reports that AI industry leaders are heading to the White House to discuss safety, while Superhuman notes that China keeps up the release pressure in the meantime.

A wave of specialized open releases. Mistral shipped Shieldstral, an open-weights 3B safety model for on-device moderation — 12 languages, 32k context, one-forward-pass safety scoring. NVIDIA introduced Alpamayo 2 Super for autonomous-driving reasoning, Pokee released Pokee-Isaac 28B claiming a 10M-token context and 93.3% on RULER while running from a single RTX 4090, and Deepgrove presented Maple-Preview, an open-source 20B ternary-weight model doing 200+ tokens/sec on a Mac Mini M4.

Numbers and facts

Different perspectives

The most contested story of the day: have open models caught up with closed ones? For: "open models are winning now", "China is competing on equal footing". Neutral: independent Vals numbers put Qwen3.8-Max at #10 overall — level with a previous-generation closed model, not above today's leaders (Vals AI). Against: "open weights" is not the same as "accessible" — giant MoE models need over a terabyte of memory and at least 8 H100-class GPUs just to load (Jamin Ball's breakdown), and the reported regional license restrictions cast doubt on commercial use.

Tools and techniques

In brief