DeepSeek V4.1-Flash: 40 Index Points at $0.30 — AI Digest

DeepSeek open-weighted V4.1-Flash under MIT: 40 index points at $0.30 per million input tokens, 763 billion parameters with only 8 billion active on input and 16 on output. It was run off an SSD at 300 tokens per second. OpenAI shipped a full-duplex voice model.
Today's highlights
DeepSeek V4.1-Flash: 40 on the intelligence index at $0.30 per million tokens
DeepSeek released V4.1-Flash under an MIT licence: 40 points on the Artificial Analysis Intelligence Index at $0.30 per million input tokens and $1.20 per million output — above its own former flagship V4 Pro and markedly cheaper than it. Cached input costs $0.006 per million, with an additional 50% off-peak discount. The model is stated at 763 billion total parameters, but only 8 billion are active for reading a request and 16 billion for generating the answer; the context window is one million tokens and the input accepts text and images. Vals placed it first among open-weight models on its index, ahead of Kimi K3, at $0.30 per test run — the cheapest in its top ten. Sebastian Raschka called the release a "big overhaul" and said it "should have been called DeepSeek V5". The Artificial Analysis breakdown, the Vals ranking, Raschka's assessment.
Architecture built for inference: input and output economics split apart
The headline novelty in V4.1-Flash is a causal encoder-decoder: the model spends 8 billion active parameters reading the request and 16 billion generating the answer, putting input and output cost on separate budgets. Everyone who read the report agrees it is all subordinated to compressing the KV cache. The KV cache is the model's memory of the text it has already read: it grows with context length, and on long documents it costs more than the computation itself. One analysis puts the cache at roughly 890 bytes per token in the regime the scores were measured in. Engineers note a local attention window paired with a sparse retrieval branch, and infer quantization-aware training for the cache — which would explain why the model holds quality with a four-bit KV cache. The technical read, the architecture read, on cache quantization.
A frontier model run off an SSD: 300 tokens per second on 32 GB of system RAM
Fraser Price reported running DeepSeek V4.1-Flash at full precision at over 300 tokens per second on four RTX Pro cards, keeping peak system memory below 32 GB — by offloading a 200-gigabyte structure to an NVMe drive. His first measurement was more modest: 200 tokens per second on four Max-Qs with 64 GB of RAM. Salvatore Sanfilippo independently ran the same model on a laptop with 128 GB and noted that SSD streaming turned out to be unexpectedly fast. The practical conclusion: for a top-tier open model, the demand for expensive video memory has stopped being the cut-off — the bottleneck moves to the drive, and drives are cheaper. A curated set of local-inference tools lives in the AI SKILLS Open Source section. The RTX Pro measurement, the first measurement, the laptop run.
Anthropic published its most detailed report on attempted misuse of Claude
Anthropic released a report on detected attempts to use Claude for cyberattacks, influence operations, surveillance, biology and weapons, and said it disrupted every operation described. It is the company's most detailed document of this kind. David Agranovich, formerly head of threat disruption at Meta, judged that the company deserves credit for this level of transparency, even if particular framings are worth debating. The publication landed alongside a separate discussion about the monitorability of reasoning: Redwood Research proposed disclosure norms for architectures that weaken the visible chain of thought, and Ryan Greenblatt argued that evidence and policies should be published before such architectures ship. Anthropic's report, Agranovich's assessment, the disclosure norms.
OpenAI opened a full-duplex voice model and hosting for agents
OpenAI shipped GPT-Live-1 into the API — a voice interface with full duplex, meaning it can listen while it is speaking, delegating reasoning and tool calls to a backend model. By the company's own measurements, paired with GPT-6 Astra it completes a task on the first attempt in 83.6% of cases on Tau3, scores 97.3% on the Artificial Analysis conversational-dynamics test, and starts replying with a 0.798-second onset latency. The same day brought public beta access to an Agents API with the Codex harness and OpenAI-hosted sandboxes — code, files and artifacts execute on the provider's side. Integrations appeared immediately at LiveKit, HeyGen and in Cognition's Devin Voice. The GPT-Live-1 launch, the measurements, the Agents API and sandboxes.
Numbers and facts
- 40 points at $0.30 / $1.20 per million tokens — DeepSeek V4.1-Flash on the Artificial Analysis Intelligence Index and its input and output pricing (breakdown).
- 763 billion total parameters, 8 billion active on input and 16 billion on output, a 1-million-token context, MIT licence (breakdown).
- 69% on AutomationBench-AA — level with GPT-6 Astra and above Grok 4.6 at 67%; on GDPval-AA v2 it gains 164 Elo, from 1468 to 1632, overtaking Kimi K3 at 1584 (breakdown).
- 89,000 tokens per task against GLM-5.3's 71,000 — V4.1-Flash is among the most verbose models measured, yet a task still costs $0.27 against GLM-5.3's $2.01 (breakdown).
- 300+ tokens per second with peak system memory under 32 GB — a full-precision run on four RTX Pro cards with SSD offload (measurement).
- 83.6% first-attempt completion and 0.798 seconds to start replying — GPT-Live-1 on Tau3 paired with GPT-6 Astra, and on the full-duplex test (measurements).
- Up to 70% cheaper at comparable quality — the claimed economics of Cognition's SWE-2; reinforcement learning was scaled to trillions of parameters (launch).
- 614 of 10,441 skills in the AI SKILLS catalog for Claude Code as of 12 September 2026 describe running models locally; 956 cover security and vulnerabilities.
- Up to 25% of trajectory tokens saved — the Elastic Horizon controller tunes an agent's horizon by the 90th percentile of successful trajectory lengths (write-up).
Different perspectives: is a cheap open model a breakthrough or an internal artifact?
The numbers are not disputed: 40 index points, first place among open models on the Vals board, $0.27 per task against $2.01 for the nearest competitor. The argument is whether those results carry over into real work.
For. Vals calls V4.1-Flash the new leading open-weight model on its board, Artificial Analysis leads with cost-adjusted intelligence, and Sebastian Raschka calls it a "big overhaul" that "should have been called V5" (Vals, Raschka).
Neutral. Some readers dissect rather than judge: the hybrid of local attention and sparse retrieval, the oddities in the first layers, the image tokenization. Aleksa Gordić notes that, going by its publications, DeepSeek looks like an "organic data" lab that apparently does not even use synthetic rephrasing (Gordić).
Against. The Teortaxes account objects on substance: DeepSeek's high internal evaluations sit alongside weak external robustness and gaps in particular skills, because the company "ships internal research artifacts and not products". The same account calls some results "very strange" — notably the AutomationBench first place next to a regression on CritPt (on products, on the strange results).
Tools and techniques
- An agent's horizon tunes itself. The Elastic Horizon controller watches the 90th percentile of successful trajectory lengths and moves the step limit accordingly: success rates rise while token spend falls by up to 25% (write-up).
- Long context gets read in parallel, not chunk by chunk. PARSER replaces sequential chunk reading with a set of frozen subagents under a lead agent trained with reinforcement learning: a reported +12 points at 896,000 tokens of context and up to 11x lower latency (write-up).
- A vocabulary for briefing a design agent. The author of Ben's Bites assembled Design Words — a set of ready phrasings for styles, layouts and components that you pick and paste into a prompt when you lack a designer's vocabulary of your own (the tool and the write-up). Ready-made agent scenarios live in the AI SKILLS automation templates catalog.
In brief
- OpenAI paused new $200 Pro sign-ups over capacity for Astra; existing users are unaffected and the API and other plans keep working (statement).
- Cursor shipped Projects — persistent work threads with a coordinator agent, shared memory and sync across devices (launch).
- ChatGPT Work gained a data agent: dashboards, answers and actions over connected corporate sources (launch).
- Training a weaker model on a stronger one's full trajectories hurts results by 4–30 points once the harness changes — only the failing turn in its own rollout should be rewritten (the Salesforce paper).
- At ByteDance, agents build and improve their own harnesses, but transfer is middling: only 34 of 64 changes moved held-out tasks in the right direction (HarnessDev write-up).
- Google Research's ToolGrad generates tool-call chains before the prompt is written and claims a near-100% pass rate for dataset construction (announcement).
- Hugging Face announced an Open Alignment team (Thomas Wolf).
- The Dioxus Labs team is joining Cognition to work on Devin's virtual machine, computer use and testing (announcement).