DeepSeek V4.1-Flash: 40 Index Points at $0.30 — AI Digest

DeepSeek V4.1-Flash: 40 Index Points at $0.30 — AI Digest

DeepSeek open-weighted V4.1-Flash under MIT: 40 index points at $0.30 per million input tokens, 763 billion parameters with only 8 billion active on input and 16 on output. It was run off an SSD at 300 tokens per second. OpenAI shipped a full-duplex voice model.

Today's highlights

DeepSeek V4.1-Flash: 40 on the intelligence index at $0.30 per million tokens

DeepSeek released V4.1-Flash under an MIT licence: 40 points on the Artificial Analysis Intelligence Index at $0.30 per million input tokens and $1.20 per million output — above its own former flagship V4 Pro and markedly cheaper than it. Cached input costs $0.006 per million, with an additional 50% off-peak discount. The model is stated at 763 billion total parameters, but only 8 billion are active for reading a request and 16 billion for generating the answer; the context window is one million tokens and the input accepts text and images. Vals placed it first among open-weight models on its index, ahead of Kimi K3, at $0.30 per test run — the cheapest in its top ten. Sebastian Raschka called the release a "big overhaul" and said it "should have been called DeepSeek V5". The Artificial Analysis breakdown, the Vals ranking, Raschka's assessment.

Architecture built for inference: input and output economics split apart

The headline novelty in V4.1-Flash is a causal encoder-decoder: the model spends 8 billion active parameters reading the request and 16 billion generating the answer, putting input and output cost on separate budgets. Everyone who read the report agrees it is all subordinated to compressing the KV cache. The KV cache is the model's memory of the text it has already read: it grows with context length, and on long documents it costs more than the computation itself. One analysis puts the cache at roughly 890 bytes per token in the regime the scores were measured in. Engineers note a local attention window paired with a sparse retrieval branch, and infer quantization-aware training for the cache — which would explain why the model holds quality with a four-bit KV cache. The technical read, the architecture read, on cache quantization.

A frontier model run off an SSD: 300 tokens per second on 32 GB of system RAM

Fraser Price reported running DeepSeek V4.1-Flash at full precision at over 300 tokens per second on four RTX Pro cards, keeping peak system memory below 32 GB — by offloading a 200-gigabyte structure to an NVMe drive. His first measurement was more modest: 200 tokens per second on four Max-Qs with 64 GB of RAM. Salvatore Sanfilippo independently ran the same model on a laptop with 128 GB and noted that SSD streaming turned out to be unexpectedly fast. The practical conclusion: for a top-tier open model, the demand for expensive video memory has stopped being the cut-off — the bottleneck moves to the drive, and drives are cheaper. A curated set of local-inference tools lives in the AI SKILLS Open Source section. The RTX Pro measurement, the first measurement, the laptop run.

Anthropic published its most detailed report on attempted misuse of Claude

Anthropic released a report on detected attempts to use Claude for cyberattacks, influence operations, surveillance, biology and weapons, and said it disrupted every operation described. It is the company's most detailed document of this kind. David Agranovich, formerly head of threat disruption at Meta, judged that the company deserves credit for this level of transparency, even if particular framings are worth debating. The publication landed alongside a separate discussion about the monitorability of reasoning: Redwood Research proposed disclosure norms for architectures that weaken the visible chain of thought, and Ryan Greenblatt argued that evidence and policies should be published before such architectures ship. Anthropic's report, Agranovich's assessment, the disclosure norms.

OpenAI opened a full-duplex voice model and hosting for agents

OpenAI shipped GPT-Live-1 into the API — a voice interface with full duplex, meaning it can listen while it is speaking, delegating reasoning and tool calls to a backend model. By the company's own measurements, paired with GPT-6 Astra it completes a task on the first attempt in 83.6% of cases on Tau3, scores 97.3% on the Artificial Analysis conversational-dynamics test, and starts replying with a 0.798-second onset latency. The same day brought public beta access to an Agents API with the Codex harness and OpenAI-hosted sandboxes — code, files and artifacts execute on the provider's side. Integrations appeared immediately at LiveKit, HeyGen and in Cognition's Devin Voice. The GPT-Live-1 launch, the measurements, the Agents API and sandboxes.

Numbers and facts

Different perspectives: is a cheap open model a breakthrough or an internal artifact?

The numbers are not disputed: 40 index points, first place among open models on the Vals board, $0.27 per task against $2.01 for the nearest competitor. The argument is whether those results carry over into real work.

For. Vals calls V4.1-Flash the new leading open-weight model on its board, Artificial Analysis leads with cost-adjusted intelligence, and Sebastian Raschka calls it a "big overhaul" that "should have been called V5" (Vals, Raschka).

Neutral. Some readers dissect rather than judge: the hybrid of local attention and sparse retrieval, the oddities in the first layers, the image tokenization. Aleksa Gordić notes that, going by its publications, DeepSeek looks like an "organic data" lab that apparently does not even use synthetic rephrasing (Gordić).

Against. The Teortaxes account objects on substance: DeepSeek's high internal evaluations sit alongside weak external robustness and gaps in particular skills, because the company "ships internal research artifacts and not products". The same account calls some results "very strange" — notably the AutomationBench first place next to a regression on CritPt (on products, on the strange results).

Tools and techniques

In brief