Qwen3.8-27B on One RTX 4090: 115 tok/s — AI Digest

Qwen3.8-27B runs at 115 tokens per second on a single RTX 4090, the 320B GLM-5.3-Flash lands in llama.cpp, and Claude Sonnet 5.5 builds a 30-second video from code for $35.40.
Top stories
Qwen3.8-27B on a single RTX 4090: 115 tokens per second and 40 of 43 hidden DeepSWE tests
The open-weight Qwen3.8-27B, quantized to IQ4_XS, decodes at 115 tokens per second on one 24 GB RTX 4090 and passes 40 of 43 hidden tests on a DeepSWE task. The author of the r/LocalLLM run used the 14.25 GB Qwen3.8-27B-UD-IQ4_XS file from Unsloth in llama.cpp with a 196,608-token context, a q8_0 key-value cache and an MTP speculative draft; peak VRAM was 22,934 MiB. Quantization is the compression of model weights to fewer bits per parameter so the model fits into the memory of an ordinary GPU. The author is explicit about the limits: all 109 existing tests passed, yet the task as a whole counts as a fail, while frontier models reach an 85.3% full pass rate on the same task in the DeepSWE v1.1 data. This is one task out of 113, not parity with the frontier. For the bigger picture on running open weights, see our open-source LLM guide.
GLM-5.3-Flash lands in llama.cpp: a 320B text-and-vision model you can run locally
llama.cpp merged pull request #27773, adding support for GLM-5.3-Flash (GLM5-Next), a 320-billion-parameter hybrid model for text and images. According to the r/LocalLLaMA write-up, the implementation adds DSA indexing, hybrid indexed memory and a dedicated glm5v vision preprocessing path, and it reuses Kimi-K3 KDA layers and DeepSeek-style mixture-of-experts helpers. Validation on a random-weight model showed logits matching the Transformers reference across prefill and decode, with vision embeddings agreeing to about 1e-5. There is a compatibility catch: existing Unsloth quants use the architecture identifier glm5next instead of glm5-next, so mainline llama.cpp may refuse to load them without a metadata fix. Commenters complain that llama.cpp support for new architectures trails model releases by weeks or months. Ready-to-run local tools are collected in the AI SKILLS Open Source section.
Claude Sonnet 5.5 builds a 30-second video from code: $35.40 versus $50.48 on Opus 5.5
An r/ClaudeAI user had Claude Sonnet 5.5 produce a 30-second, 1080p, 60 fps video with no video model: canvas in headless Chrome drew the frames, ffmpeg encoded them and the audio was synthesized. By the author's own accounting, the job took about 2.4 hours, 779 tool calls, 739K output tokens and 101.6M cache-read tokens. At API prices that comes to $35.40 on Sonnet 5.5 versus $50.48 on Opus 5.5; the gap is less than 2x because cache reads cost $0.20 per million tokens on both models. The main pushback in the thread: applying Sonnet's token usage to Opus pricing is not a fair comparison, since Opus could use a different number of tokens and steps on the same task. A professional motion designer added that such a video is hard to edit frame by frame, which limits its value in a production pipeline.
AI safety: Opus 5.5 and Astra can rewrite over half a document without tripping the Pangram detector
Vals reports that Claude Opus 5.5 and GPT-6 Astra can rewrite more than 50% of a document without the Pangram detector flagging it. An AI-text detector is a classifier that estimates the probability that a text was written by a model rather than a person. The finding also undercuts research: estimates of how much web text is AI-generated often rely on exactly these labels. Separately, Artificial Analysis found that on the CyberGym-E2E-AA cyber benchmark some frontier models are safety-blocked on more than 85% of tasks. Goodfire introduced real-time biological-risk monitors and claims 3–5x greater adversarial robustness than established screening methods, with fewer refusals on dual-use requests. Apollo Research published principles for outside evaluators who receive employee-like access to frontier labs.
Numbers and facts
AI SKILLS summary table: running open models locally, early October 2026
| Model and build | Hardware | Result | Source |
|---|---|---|---|
| Qwen3.8-27B, IQ4_XS, 14.25 GB | 1x RTX 4090 24 GB | 115 tok/s, 22,934 MiB peak, 196,608 context | r/LocalLLM |
| Qwen3.8-Flash-Next, iq4_xs, MTP | DGX Spark | 28.36 → 43.88 tok/s (1.55x) | llama.cpp #29761 |
| GLM-5.3-Flash, 320B, text and vision | — | support merged, logits match Transformers | llama.cpp #27773 |
- +30% bandwidth and −15% power versus HBM4E is the claim for NVHBM, which moves the memory controller into a custom base die (@vikramskr); Micron says NVHBM will improve its margins, and the author disputes that.
- 1.6 ms p99 reads for Cloudflare KV Instant in private beta, but storage costs $100 per MB per month, reads $0.20 per million and writes $0.10 each (release roundup).
- Up to 10,000 tokens per second per user on models above 10T parameters is the target of the startup Volantis, which is betting on optics (@omarsar0).
- Up to 2.54x lossless speedups come from new DFlash draft models for Ornith-1.5 (@ornith_).
Different perspectives: should an AI lab argue with the Vatican about model consciousness?
According to the NYT, Anthropic's Chris Olah raised pulling out of the launch of the Pope's AI encyclical, whose text rejects machine consciousness; Christopher Hale shared the report, and the post drew 28.7K engagements. Olah's team reportedly lobbied the Pope's advisers to take model consciousness seriously, and in the end he attended.
For. The article opens with Olah's admission of uncertainty: "we don't know if A.I. models are conscious". If nobody knows, ruling the question out in an official document looks premature.
Neutral. Lucas Beyer pointed to a transcript wording change from "create" to "train", a sign that even the language around the encyclical is still being edited.
Against. Aidan Gomez called the campaign moral arrogance. The encyclical itself takes the same side: its text rejects machine consciousness.
Tools and techniques
- 360-degree orbit shots. The MiniMax-H3-360-Orbit LoRA generates an orbit from a first and last frame; the trick from the r/StableDiffusion thread is to forbid any subject motion in the prompt so that only the camera moves. The AI SKILLS prompt generator helps phrase prompts like this.
- Finding the culprit commit. When Cursor Rollouts catches a regression, it finds the offending pull request, opens an issue and offers a one-click cloud agent fix.
- Sandboxes for agents. Cloudflare Sandbox SDK 1.0 gives Durable Objects direct control over sandbox containers, a convenient base for agent workflows; ready-made automations live in AI SKILLS templates.
In brief
- Agents without search score just 2.9% on legal and 7.4% on finance tasks in the new Vals Web Search Index, versus 30–50% with search (details).
- AgentWorld: fewer than a third of multi-agent actions help complete the task, and coordination tasks reach only 12% success (@omarsar0).
- ProVer delivers relative gains of +9.91% on Qwen3.5-2B and +7.12% on Qwen3.5-4B over GRPO through precise credit assignment (@omarsar0).
- ScholarCatalyst, a research-taste benchmark, is labeled by 184 lead authors on 207 of their own projects (@yoonholeee).
- AI as a compiler: a model translates Triton directly to PTX with a verifier and speeds up FlashAttention on B200 by up to 1.37x (@Azaliamirh).
- Epoch AI estimates AI infrastructure could soon support hundreds of millions to billions of agents (@EpochAIResearch).
- Yoshua Bengio joined Canada's new National Council on AI (@Yoshua_Bengio).
- Google Research launched federated learning on trusted execution environments (TEE) with verifiable differential privacy (@GoogleResearch).
- The AI Daily Brief on October 5, 2026 explored why companies want AI they can own and how that demand could fuel an American open-weight resurgence — episode "Why Companies Want AI They Can Own".