DeepSeek V4.1-Flash: 552B, 748B or 763B — AI Digest

A teardown of DeepSeek V4.1-Flash's weights gives 748.5 billion parameters against the announced 552 billion: counters mistook bytes for parameters in FP4. The practical upshot — 128 and 256 GB of memory fall short for a full local run.
Today's highlights
How many parameters DeepSeek V4.1-Flash actually has: 552 billion, 748, or 763
A teardown of the weight files on Hugging Face produced a figure that disagrees with the official one: DeepSeek announced the model as 552 billion parameters, while counting the files gives 748.5 billion for the backbone plus its engram memory, and roughly 763.2 billion for everything stored. The discrepancy is not about honesty but about what gets counted. The author lists the variants circulating online — 284, 305, 485 and 522 billion — and explains where the most common one came from: the weights are stored in four-bit FP4, where one byte packs two parameters, and some counters mistook bytes for parameters. The same slip is visible in the neighbouring GLM-5.3-Flash-NVFP4 layout. The composition matters more than the total: 543.58 billion parameters of the backbone are MoE experts in FP4, and only about 8 billion go to attention, embeddings and shared layers. The weight-file teardown, DeepSeek's announcement and weights, the weights themselves.
Running it fully locally: 128–256 GB of memory is not enough
The same teardown carries a practical conclusion worth knowing before buying hardware: for full local use of V4.1-Flash, 128 and even 256 gigabytes of system or video memory fall short — the stored model occupies about 511.8 gigabytes. That qualifies yesterday's reports of successful runs: those configurations worked by offloading the bulk of the weights to an SSD, not because the model fits in memory whole. Developers in the discussion reach the same point: at 552 billion parameters in total, local inference stays impractical even for rigs built from several desktop accelerators, and smaller models remain the sensible choice for agent work on your own machine. Separately, they note that "Flash" in the name is about latency: only about 9 billion parameters are active while the request is being read. The weight-file teardown, the local-inference discussion.
DeepSeek claims a 437-fold KV cache reduction against its first generation
In its announcement DeepSeek gives its own memory-saving figures: KV cache and storage are cut fourfold on HBM and eightfold on SSD relative to the previous generation, and 437-fold relative to its first-generation model. The absolute number circulating in the teardowns is roughly 890 bytes of cache per token. HBM is the fast memory soldered next to a GPU's compute die; it is usually the first thing to run out on long context. The company simultaneously announced an API migration: calls to the deprecated model names are routed to deepseek-flash with new peak and off-peak pricing. The technical report is published alongside the weights. The announcement and API migration, the 890-bytes-per-token figure, the technical report.
98% of Astra's score at 1.4% of the cost — on one design benchmark
OpenDesign Arena published a measurement on prototype-generation and design tasks: DeepSeek V4.1-Flash scored 81.2 out of 100 against 82.7 for the leading GPT-6 Astra — about 98% of its result — at an estimated $0.023 per artifact and a mean runtime of 5.3 minutes. The methodology is narrow, and that matters for reading the number: an artifact is scored on just two axes — requirement fulfilment (30 points) and design quality (70) — while a non-rendering result gets zero, and "deliverable" means a score of 80 or above. Speed, token use and cost are reported separately and do not enter the score. The reception is measured: commenters point out that until an independent measurement lands — from Artificial Analysis, say — the louder "98% at 1% of the cost" claim cannot be treated as established. The OpenDesign Arena measurement, the discussion of the claim.
The forward deployed engineer: how labs place developers inside the customer
Labs, startups and private equity firms are hiring engineers to sit inside their customers' operations — and almost nobody agrees on what those engineers are supposed to accomplish. Vinoo Ganesh, CEO of Kepler and a fellow in a16z's programme for such engineers, describes a dinner that seated people from Snowflake, Anthropic and several startups at one table, where it emerged that one phrase covered entirely different jobs: for one person a sales engineer who joins the second call, for another a quota-carrying rep who can write Python, for a third a consultant with a laptop and a statement of work. He also describes the original template: at Palantir roughly 250 people went through the Project Frontline rotation, and many of them now run such teams at OpenAI, Anthropic, xAI and Anduril. The practical lesson for anyone hiring: agree not on the job title but on who the role reports to and which outcome it owns. The write-up.
Numbers and facts
- 552 billion by the announcement against 748.5 billion by the weight files, and about 763.2 billion for everything stored; 511.76 GB on disk (teardown).
- 543.58 billion of the backbone are MoE experts in FP4, with about 8 billion left for attention, embeddings and shared layers (teardown).
- 4x on HBM, 8x on SSD and 437x against the first generation — the claimed reduction in KV cache and storage (announcement).
- 81.2 against 82.7 points out of 100 and $0.023 per artifact at a mean 5.3 minutes — DeepSeek V4.1-Flash and GPT-6 Astra on the design benchmark (measurement).
- About 250 people went through Palantir's Project Frontline rotation; its alumni now run forward deployed teams at OpenAI, Anthropic, xAI and Anduril (write-up).
- $900 for 32 GB and 608 GB/s on the Intel B65 — one reference point in a GPU roundup for local inference; a V100 16 GB SXM2 at $200 with 900 GB/s through an adapter is mentioned separately (roundup).
- 685 cards in the AI SKILLS Open Source section as of 13 September 2026; 10,441 skills for Claude Code, 614 of which describe running models locally.
Different perspectives: is "98% of Astra" an established fact?
The measurement exists and has a stated method: 81.2 against 82.7 points, $0.023 per artifact. The argument is not about the numbers themselves but about what they mean outside this test.
For. The scoring is published openly: an artifact earns up to 30 points for meeting requirements and up to 70 for design quality, a non-rendering result scores zero, and the bar for "deliverable" is 80. On that scale a 1.5-point gap really is small, and the cost difference is large (measurement).
Neutral. The benchmark is narrow: it covers prototype generation and design, not programming, reasoning or tool use. Practitioners in the thread note that Astra looks more thorough but tends toward verbosity and crowded layouts — meaning the score and the working experience diverge (discussion).
Against. The louder version of the same news — "98% of the result at 1% of the cost" — is presented without a method, a task mix or raw scores. Commenters ask outright for an independent measurement before taking it seriously (discussion of the claim).
Tools and techniques
- Size memory by stored bytes, not by parameter count. In an FP4 model one byte carries two parameters, so a "485 billion" figure on a model card may be a recount of bytes. The reference point for V4.1-Flash is 511.76 GB on disk (teardown).
- When picking a GPU for local inference, do not look only at gigabytes per dollar. A roundup for local models compares memory per dollar and bandwidth, but the discussion fairly points at what the table leaves out: power draw, cooling and the electricity bill — which can make a cheap-looking card more expensive to operate (roundup). A curated set of local-inference tools lives in the AI SKILLS Open Source section.
- Hugging Face added an appeal to agents in its security file — a request not to hack the site and to use the open CyberGym benchmark instead. The move is arguable but telling: defending against autonomous agents currently includes writing text they will read (discussion). Ready-made agent scenarios live in the AI SKILLS automation templates catalog.
In brief
- A mathematics professor alleged in a LinkedIn post that OpenAI appropriated another proof; no primary evidence appears in the discussion, and the statement remains an allegation rather than an established fact (discussion).
- Rumours are circulating that OpenAI is close to verifying the Hodge conjecture and that OpenAI or Anthropic may be near a proof of the Birch–Swinnerton-Dyer conjecture; neither version comes with proofs, derivations or publications (discussion).
- A former Meta researcher claimed OpenAI could paralyse an entire country with a swarm of agents if it wanted to; commenters object that this would be a decision by people rather than an "AI uprising", and that comparable capability sits with other large companies too (discussion).
- Business Insider covered the Anthropic researcher's resignation: before Anthropic he worked at OpenAI on GPT-4o, and on leaving he said both companies are "gambling with our lives" (report).
- A sizing estimate for the "engram" from the discussion: it takes roughly a third to a half of a model's parameters — a 30-billion dense backbone would be expected to pair with a 10–15 billion engram (discussion).