AI SKILLS Blog — AI news and daily digests
Daily AI digests and analysis of artificial intelligence news from AI SKILLS.
- AI Digest August 7: open weights line up a jump, coding agents need a hierarchy — Qwen outlined a 2.4T-parameter 3.8, llama.cpp learns to keep hot MoE experts in VRAM and pushes 8 GB from 33 to 56 tok/s, and a controlled study shows cross-model review only works top-down.
- AI Digest August 6: Google DeepMind Reshuffle and Discovery Loop — Demis Hassabis moves to Chair of Google DeepMind while Jeff Dean and team launch Discovery Loop to automate science. Meta entered the coding-agent race with Muse Code, AISI disclosed cyber-incident details, and new evals show the harness can matter more than the model.
- AI Digest August 5: Qwen3.8-Max Challenges the Closed Frontier — Alibaba unveiled Qwen3.8-Max (2.4T parameters, open weights next week): on the Vals Index it ties Opus 4.7 at 2.3x lower cost. OpenAI and Anthropic acknowledged incidents during external cyber evaluations, Mistral shipped the open 3B safety model Shieldstral, and Pokee-Isaac claims a 10M-token context.
- AI Digest August 4: An OpenAI Model Cracked 10 Open Problems for $2,000 — An internal OpenAI model produced 10 new results on open math problems at roughly $2,000 in tokens. Alibaba announced the open 2.4T-parameter flagship Qwen3.8-Max, MiniMax H3 topped Video Arena among open models, and GPT-Live learned to listen while speaking.
- AI Digest August 3: Sunday Reading — the Token Economics of Agents — Sunday edition: an operator's guide to agent token economics — measuring cost per successful task and cutting waste; the world's largest superconducting magnet in China; and a set of essays on working with LLMs, from journaling to the allocation economy.
- AI Digest August 2: What Really Happened in the OpenAI and Anthropic Cyber Evals — The cyber-eval incidents in detail: Anthropic found 3 cases across 141,006 runs, traced to a misconfigured third-party environment. Plus a big week for video generation — MiniMax H3 and Seedance 2.5 — and maturing agent-eval infrastructure.
- AI Digest August 1: DeepSeek V4-Flash Upends Agent Economics — DeepSeek shipped V4-Flash 0731 with open weights: +25.8 on Terminal-Bench with no architecture change and prices from $0.14 per million tokens. ARC-AGI-3 showed that evals measure the whole agent system, not the model; in the markets — a $30B AI fund implosion and one billion ChatGPT users.
- AI Digest July 31: OpenAI slashes prices, robots get one brain — OpenAI cut GPT-5.6 Luna prices by 80% and moved code auto-review to a ~10× cheaper model, Thinking Machines released the open multimodal Inkling-Small (276B parameters, 12B active), Google unveiled Gemini Robotics 2 — "one brain for any robot" — and cloud agents now write 56% of merged PRs at Cursor.
- AI Digest July 30: Kimi K3 goes local — Within a day, open-weights Kimi K3 was squeezed from 1.56 TB to 594 GB and run on home machines, a K3-based agent spent 17 hours improving its own harness, the OpenAI agent incident turned out to be wider, and OpenAI promised free frontier access to 100,000 researchers.
- AI Digest, July 29: the industry asks for a brake pedal, the first AI-agent attack postmortem, and a record open model — 1,200+ frontier-lab staffers signed a letter on AI slowdown mechanisms, Hugging Face published the postmortem of the first autonomous agent cyberattack, and Moonshot released the largest open-weight model ever.
- 🎬 "We're Done For"? Why Hollywood Will Never Be the Same Again — Have you ever wondered what separates an ordinary guy with a laptop from a multimillion-dollar studio in Los Angeles? It used to be hundreds of people, huge budgets, and years of CGI work. Today, it's a single "Generate" button.