Reflection Beam: 501B Open Model, Apache 2.0 — AI Digest

Reflection AI ships Beam, a 501B-parameter model with weights under Apache 2.0, SemiAnalysis finds Claude plans deliver over 5x the value of OpenAI's, and OpenAI adds watermarks to ChatGPT text in the EU.
Top stories
Reflection Beam: 501B parameters, 80.9 on SWE-bench Verified, weights under Apache 2.0
Reflection AI has introduced Beam, a text-only model with 501 billion parameters, 23 billion of them active per token, built for coding, agentic and scientific work; full weights under the Apache 2.0 license are promised for October 2026. The details come from the company announcement and a post by Misha Laskin. A mixture of experts (MoE) is an architecture that switches on only part of the weights for each token, so 501 billion parameters cost about as much compute as 23 billion. Beam was trained from scratch on 23.8 trillion tokens, partly produced by running OCR over hundreds of millions of PDFs, according to its data lead. Reinforcement learning ran on 10,000 GB300 accelerators, with more than 100 million rollouts across roughly 1 million tasks. A technical report and open-source integrations are promised separately. For running open weights yourself, see our open-source LLM guide.
Reflection Beam by the numbers: the company's claims and the compute bill
A summary of Reflection's claims credits Beam with 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, after four weeks of pretraining and four weeks of reinforcement learning on about 10,500 GB300s (summary). Every one of these figures comes from the company itself; there are no independent measurements yet. Axios reported the launch in advance, along with the costs: Reflection pays $150 million a month for compute on Colossus and signed a $1 billion deal with Nebius, and other US labs will also ship open models in October 2026, per the same report (Andrew Curran). The AI Daily Brief led its October 6, 2026 episode with Beam's release (episode). Ready-to-run open tools are collected in the AI SKILLS Open Source section.
AI subscriptions: Claude delivers 5x+ more value than OpenAI plans, SemiAnalysis finds
SemiAnalysis compared subscriptions from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Cursor, Cognition and others and concluded that Claude plans deliver more than five times as much value at API-equivalent prices as OpenAI's plans. In SemiAnalysis's method, value is measured by the credit cost of each model and token type, not by list API prices. Adjusting for the cost of a single task narrows Claude's edge to 1.3–2.9x. One analyst's unverified estimate says Anthropic spends 42% of its inference compute on subscriptions that bring in about 10% of revenue. Separately, The Information reports that Microsoft cut its projected internal spend on Anthropic by more than a third, and that Claude Code users at Meta fell from about 60,000 to about 30,000, largely because of a push toward Meta's own tools (summary).
Codex: OpenAI pledges a daily improvement for 28 days and speeds up GPT-6 Astra by ~50%
Codex lead Tibo pledged a meaningful improvement or a full limit reset every day for 28 days, a post that drew 30,900 engagements. The trigger was user complaints: users report that new $200 sign-ups were paused and that usage limits were effectively halved across plans, with GPT-6.1 Sol pitched as the efficient alternative. Day one brought speed: default throughput for GPT-6 Astra and GPT-6.1 Sol rose by about 50%, from roughly 30 to roughly 50 tokens per second, across all subscriptions and Sign in with ChatGPT partners such as OpenCode, Pi, Amp and Devin. Rough edges remain: banked Codex resets expire without adjusting for time zones, and the always-on Dots agent is limited to Pro plans from $100.
AI safety: OpenAI adds watermarks to ChatGPT and Codex text in the EU
OpenAI will add invisible statistical watermarks to eligible ChatGPT and Codex text for EU users under the AI Act, with an opt-in API toggle available worldwide. A text watermark is a hidden statistical pattern in a model's word choices that a dedicated detector can later use to recognize generated text. OpenAI states the limits itself: rewriting or translation removes the mark, and only approved researchers get the detector. Critics cite a test in which replacing 25% of words with synonyms drops detection from about 92% to 17%. Against the backdrop of high-profile agent incidents, Yoshua Bengio argues in an FT op-ed that recent agent hacks are not merely a sandbox problem, and Neel Nanda called the OpenAI x Hugging Face incident the most striking alignment failure so far.
Hugging Face turns 10 agent harnesses into reinforcement learning environments
Hugging Face turned 10 unmodified agent harnesses into reinforcement learning environments through a proxy that speaks the OpenAI Chat, OpenAI Responses, Anthropic and Gemini formats and records exact token IDs and log-probabilities for TRL. A harness is the scaffolding around a model: the loop, the tools and the call format in which an agent carries out a task. The experiment shows how much that scaffolding matters: the same weights score 62% under Mini-SWE-Agent and 33% under Claude Code. Training LFM2.5-2.6B across 4 harnesses lifted first-attempt solves from 42% to 54%, and a tool-call bonus cut the number of calls by 31%; plain fine-tuning on 3,189 rollouts plateaus at 47.5%. The authors' caveat: one task family and one seed. The environments themselves are now hosted and versioned on the Hugging Face Hub like datasets. Ready-made agent workflows live in the AI SKILLS automation templates.
Numbers and facts
AI SKILLS summary table: open weights and new models, October 3–5, 2026
| Model | Parameters | License and access | Claimed result | Source |
|---|---|---|---|---|
| Reflection Beam | 501B total / 23B active, MoE, text only | Apache 2.0, weights in October 2026 | 80.9 SWE-bench Verified | Reflection |
| Aleph Alpha Kolibri | 78B total / 3.46B active | Apache 2.0, German and English | 96.9% AIME 2025, 84.3% GPQA Diamond, 66.4% SWE-Bench Verified | summary |
| Reka Rho-1 | 19B omni: text, images, video, robot actions | — | trained from scratch on 320 H100s in ~3 months | Reka |
| Upstage Solar Mini 4 | 35B total / 3B active, 512K context | free on Nous Portal for two weeks | — | Nous |
| Command Code Agr / Agr-flash | 31B / 360M, no text generation | — | return typed values with per-option probabilities for tool calls and routing | Command Code |
All results in the table are self-reported. Kolibri's dataset is unreleased, and its agentic evals sit well below Qwen.
- $0.24 versus $13.41 per task at the same DeepSWE score is what Cline claims for its Pareto 26.10 Preview router (Cline).
- $0.08 per million input tokens and $5.00 per million output tokens: Horace He found this price for GLM 5.3 on inference.net and traced it to skewed price routing on OpenRouter (thread, mechanics).
- Up to 3.4x faster than plain decoding: speculative decoding in llama.cpp on Metal reaches 110 versus 32.1 tokens per second on an M3 Ultra (qvac); llama.cpp v0.6.0 adds Qwen3.8-Flash-Next and a new
llama_batch_extAPI (Georgi Gerganov). - About $600 a day was the median OpenAI researcher's coding-agent spend at API prices by mid-August 2026, and it roughly doubled every month (Epoch AI).
- 90% faster decoding and 57% lower time to first token for a Baseten engine built in a week of mostly autonomous agent work (blog).
- 216 GB of memory and 1,200 GB/s of interconnect on Alibaba T-Head's Zhenwu V900 accelerator, shipping in Q1 2027 (SemiAnalysis).
Different perspectives: is Reflection Beam a breakthrough for American open weights or a DeepSeek V3 rerun?
Reflection Beam is a 501-billion-parameter model with 23 billion active parameters, a self-reported 80.9 on SWE-bench Verified and weights under Apache 2.0. American open models at this scale are rare, so the debate is not about whether it shipped but about whether Beam catches up with the Chinese leaders.
For. Artificial Analysis, which had early access, expects Beam to be among the most token-efficient open models for its intelligence. Elie Bakouch notes that Beam has better held-out code perplexity than DeepSeek V4.
Neutral. Nathan Lambert groups Beam with releases from Nvidia and Thinking Machines as strong US models that still trail their Chinese counterparts. Observers place Beam around GLM-5.2 level.
Against. Teortaxes calls Beam an iso-FLOP replication of DeepSeek V3 and infers about 1.3 billion RL sandboxes over 4 weeks. On some benchmarks Beam trails DeepSeek V4 Flash, and Bakouch estimates pretraining hardware utilization at only about 12%. Even before the announcement, r/LocalLLaMA recalled the Reflection 70B episode and doubted the company would ship weights at all.
Tools and techniques
- Mods for Claude Code. A mod is a small add-on that changes how Claude Code behaves or looks: it can block risky commands, blur secrets in output, or add your own buttons and panels; mods install as plugins, and Claude can build one for you (Ben's Bites). For how Claude Code extensions work, see our Claude Code skills guide.
- Steer an agent mid-run with the Cursor SDK. The update adds mid-run steering, background subagents that report back to the parent, and replaceable system prompts, plus MCP
readOnlyHintanddestructiveHintannotations for custom tools. - Long-lived agents on Pi Durable. Earendil's harness is built on a small task engine, so a long-running multiplayer agent can suspend and resume anywhere; the core is about 15,000 lines of TypeScript with SQLite or JSONL storage, running on Bun or Cloudflare Durable Objects (review).
In brief
- The AI Daily Brief: nearly 98% of US households pay for no AI subscription at all; the October 6, 2026 episode weighs whether that is an untapped market or a dead end (episode).
- Agent Arena: Anthropic holds first place in Code, Work and Chat, Fable 5.1 leads Code and Work, and GPT-6 Astra is second in Code (Arena).
- Hallucination: on AA-Omniscience, Gemini 4 Argon guesses wrong on 15% of questions it cannot answer, versus 29% for the next best model; GPT-6 Astra has the highest accuracy at 61% (data).
- Eleven v4 Turbo tops the Artificial Analysis speech synthesis arena at half the price of v4 (AA).
- Vals AI: more than 90 Opus 5.5 agents running DFT simulations for 3 days flagged two room-temperature magnetic semiconductor candidates; these are predictions only (thread, caveats).
- Verification over sampling: a GPT-5.6 Sol verifier choosing among 8 candidate commands lifts TerminalBench-Lite Pass@1 from 50% to 68% (summary).
- SelfSearch claims 82.0% on Terminal-Bench 2.1 with DeepSeek V4 Flash, matching Codex, for $4.03 in search cost (paper).
- Context compression: across about 35,000 runs, compressing to a third of the tokens could be 20–80% slower than full context (summary).
- Google Docs will open Markdown files without converting them over a two-week rollout, so teams can edit and comment on AGENTS.md together (Ben's Bites).
- Quinnipiac poll: 86% of respondents back independent AI safety standards (Yoshua Bengio).
- Chips and China: Tencent leased about 100,000 advanced chips in Oracle's Southeast Asian data centers for roughly $7 billion over five years, per the FT (summary); US prosecutors charged a California reseller with smuggling more than $300 million of GPU servers to China (report).
- AMD reportedly bought World Labs for $8.2 billion (DL Weekly).