Gemini 3.8 Flash Cyber Hits 86.2% on CyberGym — AI Digest

Gemini 3.8 Flash Cyber Hits 86.2% on CyberGym — AI Digest

Google shipped a security-specialised model at Flash pricing, Meta pushed Muse Spark 1.3 with open weights promised, Alibaba's Wan 3.0 topped a video leaderboard, and two fresh papers questioned whether skill libraries help agents at all.

Today's main stories

Google: Gemini 3.8 Flash Cyber scores 86.2% on CyberGym without a premium price

Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-specialised model that scores 86.2% on CyberGym and 47.2% on CWE-Bench for patching. Sundar Pichai announced it, adding that an internal vulnerability-discovery benchmark puts the model above 70% success across twenty programming languages while keeping Flash-level speed and pricing. The pricing detail matters more than the headline number: security-tuned models have usually been sold as slow premium tiers, and here specialisation arrives without a surcharge. For engineering teams that changes where the model sits — automated vulnerability review stops being a separate expensive pass and becomes part of ordinary CI.

Meta: Muse Spark 1.3 targets agentic and coding work, open weights promised

Meta released Muse Spark 1.3, the strongest model in the Spark line for agentic and coding tasks, built for longer-horizon work and stricter instruction following. Shengjia Zhao introduced the model, and Alexandr Wang highlighted what it delivers per dollar. On r/LocalLLaMA the thread on incoming open weights focused on a reported MRCR result of 98.1% at 512K–1M context; if that holds up under independent testing it would be one of the strongest million-token retention numbers published so far. Superhuman AI put Meta's return to the frontier race in its 3 September 2026 headline.

Alibaba: Wan 3.0 takes first place in video editing with audio, from $0.05 per second

Wan 3.0 ranks first in Video Editing with Audio, second in Text-to-Video with Audio and fifth in Image-to-Video with Audio on the Artificial Analysis leaderboards. Artificial Analysis published the placements alongside the model's specification: it accepts text, images, video, audio, documents and web pages as references, supports native audio and generates up to 30 seconds at 1080p. Public preview pricing starts at $0.05 per second for 480p and rises to $0.20 per second for 1080p. For anyone running a production pipeline the interesting part is consolidation — generation, editing and audio in one model removes the stitching between three separate services. Shot briefs and prompts for video work live in the AI SKILLS prompt generator.

Qwen3.8-Max-0902 edges past Opus 5 Max by three points in web development

Qwen3.8-Max-0902 took first place on the Code Arena WebDev leaderboard with 1691, ahead of Claude Opus 5 Max at 1688 and Kimi K3 Max at 1674. The screenshot was dissected on r/LocalLLaMA; a three-point gap sits inside the noise band of arena scoring, so parity is the honest reading rather than victory. The discussion is more informative than the table: participants report that a locally run Qwen3.8-27B satisfies their coding needs better than a paid subscription, and the open question for the leader is whether Max weights will ever be released. Curated open models and local-run tooling sit in the AI SKILLS open source section.

Numbers and facts

Different views: do skill libraries and self-built harnesses actually help?

Two fresh papers land on opposite sides, and both measure something that was previously taken on trust. A harness is the scaffolding around a model — tools, memory, call ordering and stopping rules — everything that turns a model into a working agent.

For. ByteDance Seed's HarnessDev has a model start from a weak but runnable seed, build its own harness, then improve it from downstream feedback. Elvis Saravia summarised the work: across four domains and 2,207 held-out instances, generated harnesses match or beat human-engineered systems in writing and ML experimentation.

Neutral. The same evaluation records that on code, search and research tasks self-built harnesses still trail mature human systems, and that gains are unstable and only partially transferable between models.

Against. A paper on the Retrieval-Invoked Actual-Use Effect, summarised by dair.ai, proposes an honest method: run the same task twice, with and without skills, and count only the cases where retrieval actually fired. Across 17 models on coding and maths it finds cases where a library lifts the aggregate score while hurting results on exactly the tasks that triggered it. For teams maintaining skill libraries that is a direct warning against trusting headline lift.

Tools and techniques

In brief