Yandex Opens AliceAI-Foundation-80B Weights — AI Digest

Yandex published open weights for AliceAI-Foundation-80B-A3B, trained from scratch but not yet post-trained. The same day brought DeepSeek's 8-trillion-parameter plan, Alibaba's 5–10 trillion target and a slide confirming the Qwen4 line.
Today's highlights
Yandex releases AliceAI-Foundation-80B-A3B, a Russian model trained from scratch
Yandex published open weights for AliceAI-Foundation-80B-A3B-Base: 80 billion parameters with roughly 3 billion active per token. The "80B-A3B" notation matters — active parameters are the share of weights the model actually engages at each step, which is why a sparse model of that size runs faster than a dense one. The Hugging Face card describes it as trained from scratch rather than fine-tuned from someone else's checkpoint, and the LocalLLaMA thread adds two practical notes: there was no llama.cpp support at publication, and readers place the architecture close to the Qwen3-Next 80B-A3B layout, extended with Kimi Delta Attention.
The caveat is the word Base. This is a base model with no post-training: no instruction tuning, no RLHF, no chat alignment. Commenters on 23 September 2026 asked plainly that it not be compared with Qwen or DeepSeek as-is, since post-training is where the assistant behaviour that benchmarks measure actually comes from. A separate note urged running any third-party weights in a locked-down sandbox, because tokenizer and model code ship alongside the weights. For Russian-language work this is the first serious attempt in a while to build a foundation rather than a layer on top of someone else's; what it is worth will be visible only after post-training. Open projects that are ready to deploy are collected in the AI SKILLS open-source section.
The size race: DeepSeek is training 2 trillion parameters and planning 8
Within a day, plans from three Chinese labs surfaced, and all three are about multiplying size. According to a claim discussed in LocalLLaMA, DeepSeek is training a 2-trillion-parameter model and planning an 8-trillion one; the same thread lists current sizes — Flash at 552 billion and Pro at 1.6 trillion total with 49 billion active, a sparse arrangement. The primary source is a post on X with no training details attached. Alibaba, per a thread the same day, is preparing a model in the 5–10 trillion range and unveiled a chip of its own. And third: a conference slide confirms a Qwen4 line — Qwen4-27B alongside Qwen4-Max, Qwen4-Flash and Qwen4-Plus.
The practical reading in those threads was sober. A 5–10 trillion parameter model is out of reach for home hardware — "minutes per token," as one commenter put it — and its value will only reach local machines through distillation into compact descendants such as the Qwen4-27B now being discussed. A substantive counterargument to the race itself also appeared: as models grow they can memorise more of the training data without necessarily generalising better. How to pick an open model for a specific job is covered in our guide to open-source LLMs.
Biosecurity: genomic models have already assembled a working virus
Genomic language models trained on DNA were used to assemble complete bacteriophage genomes, which were then synthesised into functional viruses. A genomic language model predicts DNA sequence the way an ordinary one predicts text, except its alphabet has four letters. Eric Nguyen, co-founder of Radical Numerics, described this on the Latent Space podcast on 23 September 2026: he helped develop the Evo and Evo 2 models at Arc Institute, and a separate Arc and Stanford team used them to generate phage genomes that turned out to be viable.
Why this became possible only recently is a question of context length rather than model size. An average human gene runs to about 60,000 letters, the longest reaches 2.3 million, and the whole genome is around 3 billion; handling sequences of that length became practical roughly three years ago, well before frontier labs began advertising million-token context. Hence the argument over what to do about it: the same material carries Clement Delangue's position that defensive capability has to be open and has to keep pace with the offensive kind — open weights have already been used in the response to a real attack. The two domains Anthropic's filters flag as sensitive are exactly these: cyber-security and biology.
Google's ERA: a tree of experiments that produced at least ten papers
Google built a system that searches for a solution to any scientific problem that can be written down as a score, and it has yielded no fewer than ten publications. John Platt, author of two textbook machine-learning algorithms, explained ERA — Empirical Research Assistance — on the Latent Space podcast. The model keeps a running tree of past experiments as notebooks and, at each iteration, picks which one to mutate using the upper confidence bound rule, a close relative of Monte Carlo tree search. The choice is optimistic rather than greedy, so even the fifth-best notebook sometimes gets picked; about ten mutations are proposed at a time, and branch history is shared, so leaves learn from each other. Evolutionary algorithms date to the 1970s — this works because the model knows where to look: between Gemini 2.0 and 2.5 the approach went from not working to working well.
The most useful part of the account is a warning rather than a result. Google's contrail-detection competition was won by entrants who spotted a half-pixel error in the labels — whether a pixel's origin sits at its corner or its centre. That was worth $15,000 in prize money and moved the actual problem forward not at all: Goodhart's law in its purest form, where a measure that becomes a target stops being a measure. Platt's advice to anyone starting out is to fit a linear regression or a support vector machine first, and only then reach for anything else.
Numbers and facts
- 100 million parameters and under 10 hours on a single H100 is what the authors say it took to train from scratch — Supra2-IMG draws 256×256 images in roughly 2 seconds on a GPU and 20 on a CPU.
- Elo 1082 against 1052 is how the 6-billion-parameter Ming-Image-0.1-Design beats Ideogram 4.0 Quality on an interface-design leaderboard among open models (screenshot breakdown); the licence is MIT.
- 5.7 of 17 SWE-bench-Live tasks at a file size of 3.6 GB is the claimed result for SharpSpark, a quantised build of Spark-X2.5-4B aimed at small-VRAM machines; each solve takes about 74 minutes.
- 530 million parameters on a laptop with 8 GB of VRAM is the current size of mini-AGI, trained from scratch on a one-example-at-a-time stream; a commenter's own experiment found that layers inserted later contributed about 23 times less than the original ones.
- 18 MB and 2,000–4,000 tokens are the size and half-life of the KV-cache bank in phantom-kv, which changes model behaviour without editing weights; sceptics in the thread read it as an invisible system prompt by another name.
- From 2.2–2.5x down to 1.4–1.5x is how much better the Max 20× plan now is than Max 5×, by one measurement of Claude Code throughput: the former fell from 118 to 92 normalised weekly units while the latter holds near 65.
- 15.2 GB of VRAM was the peak when running Qwen-Image-2.1 in int8 on a 16 GB card, per a user report.
Our table of the week's announced sizes (total parameters / active per token):
| Model | Size | Status |
|---|---|---|
| DeepSeek Flash | 552B | released |
| DeepSeek Pro | 1.6T / 49B | released |
| DeepSeek (next) | 2T, 8T planned | training |
| Alibaba (next) | 5–10T | planned |
| Yandex AliceAI-Foundation | 80B / ~3B | open weights |
Different views: what is a claim of a hundred solved maths problems worth?
The day's argument circled a claim that an internal OpenAI model, after 24 days of training, solved the Navier–Stokes problem and more than a hundred open questions in mathematics. The original post carries no proofs, no papers and no model name — only text and a mention of an advisory group of mathematicians.
For. That the claim is being argued at all marks a shift in the question: the dispute is no longer whether a model can solve hard problems but whether people can check its output fast enough. Superhuman AI put the subject in its headline on 22 September 2026.
Neutral. The most substantive objection in the thread separates two abilities: solving a problem someone has already posed, and identifying a new one worth posing. The claim covers only the first, and it is the second that distinguishes a researcher from a solver.
Against. A second thread on the same subject is openly satirical: no benchmark, no proof artefacts, no independent validation. Until something verifiable appears, this is an announcement rather than a result.
Tools and techniques
- Ask for a probability, not for text. To turn an ordinary model into a classifier, it is enough to cap the output at one token and switch on log-probabilities: an answer of "1" at a log-probability of −0.00456 means about 99.5% confidence, while "0" at −5.395 means about 0.45%. A separate decision model for simple branches in a pipeline is often unnecessary — the right calling mode is.
- Check what the image licence actually permits. The Qwen team clarified separately that generated outputs are not "Licensed Materials" and that rights to the images stay with the user. The thread fairly points out that the licence text on Hugging Face still contains a clause barring commercial use of the materials — for client work, go by the binding text rather than the summary.
- A deleted chat is not deleted memory. Users demonstrated that details from erased conversations survive in the assistant's separate memory and resurface later. The workable mitigation from the same thread is to keep sensitive topics inside a project with isolated memory. Ready-made prompts for specific jobs live in the AI SKILLS prompt library.
In brief
- Readers of the Qwen4 thread are hoping for a variant whose architecture lowers VRAM requirements, and for Qwen4 Flash to fit inside 128 GB — the slide thread.
- Two agent skills shipped alongside the Ming-Image-0.1-Design models: interface design and turning an image into an editable presentation — announcement.
- Claude Opus 5.5 is billed as 40% cheaper and over 30% faster than Opus 5 at $4 per million input and $20 per million output tokens — the release card discussed.
- Five-hour limits on the Pro, Max and Team plans went up, and the limit reset can now be banked and spent when needed — the announcement thread.
- The Opus 5.5 benchmark table was published without methodology, sample sizes or error bars — the thread that looked at it — and by that table GPT-6 Astra still leads on business workflows and agentic scientific research.