Gemini 3.8 Flash Cyber Hits 86.2% on CyberGym — AI Digest

Google shipped a security-specialised model at Flash pricing, Meta pushed Muse Spark 1.3 with open weights promised, Alibaba's Wan 3.0 topped a video leaderboard, and two fresh papers questioned whether skill libraries help agents at all.
Today's main stories
Google: Gemini 3.8 Flash Cyber scores 86.2% on CyberGym without a premium price
Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-specialised model that scores 86.2% on CyberGym and 47.2% on CWE-Bench for patching. Sundar Pichai announced it, adding that an internal vulnerability-discovery benchmark puts the model above 70% success across twenty programming languages while keeping Flash-level speed and pricing. The pricing detail matters more than the headline number: security-tuned models have usually been sold as slow premium tiers, and here specialisation arrives without a surcharge. For engineering teams that changes where the model sits — automated vulnerability review stops being a separate expensive pass and becomes part of ordinary CI.
Meta: Muse Spark 1.3 targets agentic and coding work, open weights promised
Meta released Muse Spark 1.3, the strongest model in the Spark line for agentic and coding tasks, built for longer-horizon work and stricter instruction following. Shengjia Zhao introduced the model, and Alexandr Wang highlighted what it delivers per dollar. On r/LocalLLaMA the thread on incoming open weights focused on a reported MRCR result of 98.1% at 512K–1M context; if that holds up under independent testing it would be one of the strongest million-token retention numbers published so far. Superhuman AI put Meta's return to the frontier race in its 3 September 2026 headline.
Alibaba: Wan 3.0 takes first place in video editing with audio, from $0.05 per second
Wan 3.0 ranks first in Video Editing with Audio, second in Text-to-Video with Audio and fifth in Image-to-Video with Audio on the Artificial Analysis leaderboards. Artificial Analysis published the placements alongside the model's specification: it accepts text, images, video, audio, documents and web pages as references, supports native audio and generates up to 30 seconds at 1080p. Public preview pricing starts at $0.05 per second for 480p and rises to $0.20 per second for 1080p. For anyone running a production pipeline the interesting part is consolidation — generation, editing and audio in one model removes the stitching between three separate services. Shot briefs and prompts for video work live in the AI SKILLS prompt generator.
Qwen3.8-Max-0902 edges past Opus 5 Max by three points in web development
Qwen3.8-Max-0902 took first place on the Code Arena WebDev leaderboard with 1691, ahead of Claude Opus 5 Max at 1688 and Kimi K3 Max at 1674. The screenshot was dissected on r/LocalLLaMA; a three-point gap sits inside the noise band of arena scoring, so parity is the honest reading rather than victory. The discussion is more informative than the table: participants report that a locally run Qwen3.8-27B satisfies their coding needs better than a paid subscription, and the open question for the leader is whether Max weights will ever be released. Curated open models and local-run tooling sit in the AI SKILLS open source section.
Numbers and facts
- 86.2% — Gemini 3.8 Flash Cyber on CyberGym; 47.2% on CWE-Bench patching; above 70% on an internal vulnerability-discovery benchmark across twenty languages.
- 1691 vs 1688 vs 1674 — Qwen3.8-Max-0902, Claude Opus 5 Max and Kimi K3 Max on Code Arena WebDev.
- 30 seconds at 1080p and $0.05–$0.20 per second — Wan 3.0 length ceiling and public preview pricing.
- 2,207 held-out instances across four domains and six creator models — the scope of the HarnessDev evaluation.
- 17 models — coverage of the study on skill-retrieval effects.
- 10,266 Claude Code skills — active cards in the AI SKILLS catalogue as of 4 September 2026, which is exactly the library scale both papers argue about.
Different views: do skill libraries and self-built harnesses actually help?
Two fresh papers land on opposite sides, and both measure something that was previously taken on trust. A harness is the scaffolding around a model — tools, memory, call ordering and stopping rules — everything that turns a model into a working agent.
For. ByteDance Seed's HarnessDev has a model start from a weak but runnable seed, build its own harness, then improve it from downstream feedback. Elvis Saravia summarised the work: across four domains and 2,207 held-out instances, generated harnesses match or beat human-engineered systems in writing and ML experimentation.
Neutral. The same evaluation records that on code, search and research tasks self-built harnesses still trail mature human systems, and that gains are unstable and only partially transferable between models.
Against. A paper on the Retrieval-Invoked Actual-Use Effect, summarised by dair.ai, proposes an honest method: run the same task twice, with and without skills, and count only the cases where retrieval actually fired. Across 17 models on coding and maths it finds cases where a library lifts the aggregate score while hurting results on exactly the tasks that triggered it. For teams maintaining skill libraries that is a direct warning against trusting headline lift.
Tools and techniques
- Spark-X2.5 at 4B and 1.7B — the release thread on r/LocalLLaMA reports a native 1M-token context and roughly 20T training tokens, with llama.cpp support still waiting on pull request #27868. If the claims reproduce, this is the smallest published model carrying a million-token window.
- Photon 2.1 — Vikhyat Korrapati announced text-to-speech models and NVIDIA B200 support in the realtime multimodal inference engine.
- GLM-5.3 Fast on Baseten — the host announced availability with an emphasis on higher tokens per second for realtime deployments.
In brief
- Stanford is rebuilding The Modern Software Developer: Mihail Eric writes that 85% of the Fall 2025 material is being replaced with agent skills, context engineering, MCP portals and agentic code review.
- A second Stanford course, CS329Z Engineering AI Agents, covers building agents from scratch — announced by Diyi Yang and Michael Ryan.
- The open model Marin 535B-A23B is 13% through training, Percy Liang reported, with compute funded by the Jen-Hsun and Lori Huang Foundation and run on CoreWeave.
- Sebastian Raschka unpacked the Astra architecture rumours and pointed to Nanbeige 4.2-3B: a 22-layer stack reused twice behaves like a 44-layer model at the same memory footprint and roughly double the compute.
- The imagine service raised its limit to 14 references per video, covering images, voices and character references attached by tag in the prompt.
- Palmimo DevKit was introduced as a tabletop robot platform with open-source software and swappable AI brains, controllable from a few lines of Python.
- The AI Daily Brief episode of 3 September 2026 covered agentic loops for knowledge workers: how to define a verifiable finish line and stop a loop from running up costs.