DeepSeek V4.1-Flash: 552B, 748B or 763B — AI Digest

DeepSeek V4.1-Flash: 552B, 748B or 763B — AI Digest

A teardown of DeepSeek V4.1-Flash's weights gives 748.5 billion parameters against the announced 552 billion: counters mistook bytes for parameters in FP4. The practical upshot — 128 and 256 GB of memory fall short for a full local run.

Today's highlights

How many parameters DeepSeek V4.1-Flash actually has: 552 billion, 748, or 763

A teardown of the weight files on Hugging Face produced a figure that disagrees with the official one: DeepSeek announced the model as 552 billion parameters, while counting the files gives 748.5 billion for the backbone plus its engram memory, and roughly 763.2 billion for everything stored. The discrepancy is not about honesty but about what gets counted. The author lists the variants circulating online — 284, 305, 485 and 522 billion — and explains where the most common one came from: the weights are stored in four-bit FP4, where one byte packs two parameters, and some counters mistook bytes for parameters. The same slip is visible in the neighbouring GLM-5.3-Flash-NVFP4 layout. The composition matters more than the total: 543.58 billion parameters of the backbone are MoE experts in FP4, and only about 8 billion go to attention, embeddings and shared layers. The weight-file teardown, DeepSeek's announcement and weights, the weights themselves.

Running it fully locally: 128–256 GB of memory is not enough

The same teardown carries a practical conclusion worth knowing before buying hardware: for full local use of V4.1-Flash, 128 and even 256 gigabytes of system or video memory fall short — the stored model occupies about 511.8 gigabytes. That qualifies yesterday's reports of successful runs: those configurations worked by offloading the bulk of the weights to an SSD, not because the model fits in memory whole. Developers in the discussion reach the same point: at 552 billion parameters in total, local inference stays impractical even for rigs built from several desktop accelerators, and smaller models remain the sensible choice for agent work on your own machine. Separately, they note that "Flash" in the name is about latency: only about 9 billion parameters are active while the request is being read. The weight-file teardown, the local-inference discussion.

DeepSeek claims a 437-fold KV cache reduction against its first generation

In its announcement DeepSeek gives its own memory-saving figures: KV cache and storage are cut fourfold on HBM and eightfold on SSD relative to the previous generation, and 437-fold relative to its first-generation model. The absolute number circulating in the teardowns is roughly 890 bytes of cache per token. HBM is the fast memory soldered next to a GPU's compute die; it is usually the first thing to run out on long context. The company simultaneously announced an API migration: calls to the deprecated model names are routed to deepseek-flash with new peak and off-peak pricing. The technical report is published alongside the weights. The announcement and API migration, the 890-bytes-per-token figure, the technical report.

98% of Astra's score at 1.4% of the cost — on one design benchmark

OpenDesign Arena published a measurement on prototype-generation and design tasks: DeepSeek V4.1-Flash scored 81.2 out of 100 against 82.7 for the leading GPT-6 Astra — about 98% of its result — at an estimated $0.023 per artifact and a mean runtime of 5.3 minutes. The methodology is narrow, and that matters for reading the number: an artifact is scored on just two axes — requirement fulfilment (30 points) and design quality (70) — while a non-rendering result gets zero, and "deliverable" means a score of 80 or above. Speed, token use and cost are reported separately and do not enter the score. The reception is measured: commenters point out that until an independent measurement lands — from Artificial Analysis, say — the louder "98% at 1% of the cost" claim cannot be treated as established. The OpenDesign Arena measurement, the discussion of the claim.

The forward deployed engineer: how labs place developers inside the customer

Labs, startups and private equity firms are hiring engineers to sit inside their customers' operations — and almost nobody agrees on what those engineers are supposed to accomplish. Vinoo Ganesh, CEO of Kepler and a fellow in a16z's programme for such engineers, describes a dinner that seated people from Snowflake, Anthropic and several startups at one table, where it emerged that one phrase covered entirely different jobs: for one person a sales engineer who joins the second call, for another a quota-carrying rep who can write Python, for a third a consultant with a laptop and a statement of work. He also describes the original template: at Palantir roughly 250 people went through the Project Frontline rotation, and many of them now run such teams at OpenAI, Anthropic, xAI and Anduril. The practical lesson for anyone hiring: agree not on the job title but on who the role reports to and which outcome it owns. The write-up.

Numbers and facts

Different perspectives: is "98% of Astra" an established fact?

The measurement exists and has a stated method: 81.2 against 82.7 points, $0.023 per artifact. The argument is not about the numbers themselves but about what they mean outside this test.

For. The scoring is published openly: an artifact earns up to 30 points for meeting requirements and up to 70 for design quality, a non-rendering result scores zero, and the bar for "deliverable" is 80. On that scale a 1.5-point gap really is small, and the cost difference is large (measurement).

Neutral. The benchmark is narrow: it covers prototype generation and design, not programming, reasoning or tool use. Practitioners in the thread note that Astra looks more thorough but tends toward verbosity and crowded layouts — meaning the score and the working experience diverge (discussion).

Against. The louder version of the same news — "98% of the result at 1% of the cost" — is presented without a method, a task mix or raw scores. Commenters ask outright for an independent measurement before taking it seriously (discussion of the claim).

Tools and techniques

In brief