DeepSeek V4 Flash Vision and Ox Alpha — AI Digest

DeepSeek V4 Flash Vision and Ox Alpha — AI Digest

DeepSeek shipped V4-Flash-Vision-Exp at 83.9 on Terminal Bench 2.1 with images at Flash pricing. Stealth model Ox Alpha cleared 80%+ on DeepSWE tasks, and OpenAI cut GPT-5.6 Sol pricing by over 20% for three months.

Today's top stories

DeepSeek ships V4-Flash-Vision-Exp: 83.9 on Terminal Bench 2.1, images at Flash pricing

DeepSeek released DeepSeek-V4-Flash-Vision-Exp on 21 August 2026 — an image-capable version of its fast model that the company says preserves V4-Flash text performance while approaching Opus-4.8 on multimodal agent work (DeepSeek announcement). The published table shows 83.9 on Terminal Bench 2.1, 75.9 on Toolathlon-Verified and 64.3 on Chartography (table breakdown). Mixed text-and-image input runs through the API, where an image costs between 117 and 384 tokens and is billed at Flash rates (API details). DeepSeek also opened a Files API: upload an image once, reference it by id afterwards, and stop resending the payload on every request (Files API announcement).

Stealth model Ox Alpha posts 80%+ on DeepSWE tasks against Fable's 65%

An unnamed, unattributed model going by Ox Alpha scored above 80% on ten DeepSWE tasks, against 65% for Fable and 52% for GPT-5.6 Sol (measurement). A stealth model is one a lab puts into public hands anonymously, to collect unbiased feedback before it announces anything. Builders described it as tearing through their internal benchmarks, and one merged eight pull requests in a row on its approval (first reaction, end of day). Guesses converged on Zhipu's GLM family: Tim Dettmers read the fast output and weak partial prefill as signs of fewer active parameters (analysis), while other observers narrowed it to GLM-5.3/5.4 Vision (observation). Third-party clients had it the same day (Hermes Agent and OpenRouter, Cline).

OpenAI cuts GPT-5.6 Sol pricing by more than 20% for three months

OpenAI cut GPT-5.6 Sol pricing by over 20% across the API and credit-based products on 21 August 2026, for three months (OpenAI announcement, developer details). The discounts stack: on Devin, Sol now lands 76% below list through 3 October 2026 once promotions combine (Cognition's math), while the Code editor holds its own 50% off through 3 September (Code promotion). Separately OpenAI said Codex reached 20M active users and reset usage counters for every Codex and ChatGPT Work user (OpenAI statement). Teams also got hard spend ceilings: usage is now visible per API key, and monthly limits can be set per organization and per project (announcement).

NVIDIA AVO clears all 183 ARC-AGI-3 levels — on the public set only

NVIDIA's AVO agent reached 100% on ARC-AGI-3: all 183 levels across 25 environments, with no instructions, no stated rules and no stated goals (NVIDIA claim). François Chollet immediately clarified that this is the public demo and tutorial set rather than the full benchmark (Chollet's caveat). Community analysis landed on the same point: the run covered open environments, the private set was not tested, so generalization is not yet demonstrated (discussion). ARC-AGI-3 is a set of interactive environments where the agent has to infer the goal and the rules itself, which is exactly where the gap between memorized and understood shows up most clearly.

Nvidia to pay Poolside roughly $6B to license its model factory

Nvidia will reportedly pay startup Poolside around $6B for a non-exclusive license to its model-development system, plus roughly $1B invested at a $12B pre-money valuation, according to The Information (deal breakdown). The asset is Poolside's "model factory", the system behind its Laguna coding model; 109 employees who worked on Laguna received Nvidia offers. Reaction split: some expect a boost to Nvidia's open-model work, others call it an acqui-hire that ends Laguna as an independent line. A separate argument for open weights landed the same day from David Sacks: legal service Harvey reached state of the art in its domain on the open Kimi K3, at lower cost (post).

Numbers and facts

Different perspectives: does 100% on ARC-AGI-3 prove general reasoning?

The argument turns on one result: NVIDIA's AVO agent cleared all 183 levels across 25 ARC-AGI-3 environments with no instructions and no stated goals (NVIDIA claim). The question is what got measured — the ability to work out an unfamiliar environment, or a fit to one specific benchmark.

For. ARC-AGI-3 environments are interactive, and the agent has to derive the rules and the objective itself: clearing them back to back with no prior knowledge is a qualitatively different result from solving a set of static problems. Long autonomous scenarios are exactly where the research agenda is moving, from Google's environment reshaping to benchmarks like FACET (EnvHarness breakdown).

Neutral. François Chollet, who created ARC, does not dispute the result — he bounds it: what was cleared is the public demo and tutorial set, not the full benchmark (Chollet's caveat). That is not a rebuttal but a reminder that ARC-AGI keeps a private half precisely for claims like this one.

Against. Skeptics in the discussion argue that without a held-out run the result is indistinguishable from benchmark maxing, and propose a simple test — put the same agent into unrelated interactive games it never prepared for (discussion). The same skepticism showed up around open models that day: users found Qwen3.8-27B noticeably weaker on factual recall than version 3.6 once web search was switched off (breakdown).

Tools and techniques

In brief