Reflection Beam: 501B Open Model, Apache 2.0 — AI Digest

Reflection Beam: 501B Open Model, Apache 2.0 — AI Digest

Reflection AI ships Beam, a 501B-parameter model with weights under Apache 2.0, SemiAnalysis finds Claude plans deliver over 5x the value of OpenAI's, and OpenAI adds watermarks to ChatGPT text in the EU.

Top stories

Reflection Beam: 501B parameters, 80.9 on SWE-bench Verified, weights under Apache 2.0

Reflection AI has introduced Beam, a text-only model with 501 billion parameters, 23 billion of them active per token, built for coding, agentic and scientific work; full weights under the Apache 2.0 license are promised for October 2026. The details come from the company announcement and a post by Misha Laskin. A mixture of experts (MoE) is an architecture that switches on only part of the weights for each token, so 501 billion parameters cost about as much compute as 23 billion. Beam was trained from scratch on 23.8 trillion tokens, partly produced by running OCR over hundreds of millions of PDFs, according to its data lead. Reinforcement learning ran on 10,000 GB300 accelerators, with more than 100 million rollouts across roughly 1 million tasks. A technical report and open-source integrations are promised separately. For running open weights yourself, see our open-source LLM guide.

Reflection Beam by the numbers: the company's claims and the compute bill

A summary of Reflection's claims credits Beam with 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, after four weeks of pretraining and four weeks of reinforcement learning on about 10,500 GB300s (summary). Every one of these figures comes from the company itself; there are no independent measurements yet. Axios reported the launch in advance, along with the costs: Reflection pays $150 million a month for compute on Colossus and signed a $1 billion deal with Nebius, and other US labs will also ship open models in October 2026, per the same report (Andrew Curran). The AI Daily Brief led its October 6, 2026 episode with Beam's release (episode). Ready-to-run open tools are collected in the AI SKILLS Open Source section.

AI subscriptions: Claude delivers 5x+ more value than OpenAI plans, SemiAnalysis finds

SemiAnalysis compared subscriptions from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Cursor, Cognition and others and concluded that Claude plans deliver more than five times as much value at API-equivalent prices as OpenAI's plans. In SemiAnalysis's method, value is measured by the credit cost of each model and token type, not by list API prices. Adjusting for the cost of a single task narrows Claude's edge to 1.3–2.9x. One analyst's unverified estimate says Anthropic spends 42% of its inference compute on subscriptions that bring in about 10% of revenue. Separately, The Information reports that Microsoft cut its projected internal spend on Anthropic by more than a third, and that Claude Code users at Meta fell from about 60,000 to about 30,000, largely because of a push toward Meta's own tools (summary).

Codex: OpenAI pledges a daily improvement for 28 days and speeds up GPT-6 Astra by ~50%

Codex lead Tibo pledged a meaningful improvement or a full limit reset every day for 28 days, a post that drew 30,900 engagements. The trigger was user complaints: users report that new $200 sign-ups were paused and that usage limits were effectively halved across plans, with GPT-6.1 Sol pitched as the efficient alternative. Day one brought speed: default throughput for GPT-6 Astra and GPT-6.1 Sol rose by about 50%, from roughly 30 to roughly 50 tokens per second, across all subscriptions and Sign in with ChatGPT partners such as OpenCode, Pi, Amp and Devin. Rough edges remain: banked Codex resets expire without adjusting for time zones, and the always-on Dots agent is limited to Pro plans from $100.

AI safety: OpenAI adds watermarks to ChatGPT and Codex text in the EU

OpenAI will add invisible statistical watermarks to eligible ChatGPT and Codex text for EU users under the AI Act, with an opt-in API toggle available worldwide. A text watermark is a hidden statistical pattern in a model's word choices that a dedicated detector can later use to recognize generated text. OpenAI states the limits itself: rewriting or translation removes the mark, and only approved researchers get the detector. Critics cite a test in which replacing 25% of words with synonyms drops detection from about 92% to 17%. Against the backdrop of high-profile agent incidents, Yoshua Bengio argues in an FT op-ed that recent agent hacks are not merely a sandbox problem, and Neel Nanda called the OpenAI x Hugging Face incident the most striking alignment failure so far.

Hugging Face turns 10 agent harnesses into reinforcement learning environments

Hugging Face turned 10 unmodified agent harnesses into reinforcement learning environments through a proxy that speaks the OpenAI Chat, OpenAI Responses, Anthropic and Gemini formats and records exact token IDs and log-probabilities for TRL. A harness is the scaffolding around a model: the loop, the tools and the call format in which an agent carries out a task. The experiment shows how much that scaffolding matters: the same weights score 62% under Mini-SWE-Agent and 33% under Claude Code. Training LFM2.5-2.6B across 4 harnesses lifted first-attempt solves from 42% to 54%, and a tool-call bonus cut the number of calls by 31%; plain fine-tuning on 3,189 rollouts plateaus at 47.5%. The authors' caveat: one task family and one seed. The environments themselves are now hosted and versioned on the Hugging Face Hub like datasets. Ready-made agent workflows live in the AI SKILLS automation templates.

Numbers and facts

AI SKILLS summary table: open weights and new models, October 3–5, 2026

ModelParametersLicense and accessClaimed resultSource
Reflection Beam501B total / 23B active, MoE, text onlyApache 2.0, weights in October 202680.9 SWE-bench VerifiedReflection
Aleph Alpha Kolibri78B total / 3.46B activeApache 2.0, German and English96.9% AIME 2025, 84.3% GPQA Diamond, 66.4% SWE-Bench Verifiedsummary
Reka Rho-119B omni: text, images, video, robot actions—trained from scratch on 320 H100s in ~3 monthsReka
Upstage Solar Mini 435B total / 3B active, 512K contextfree on Nous Portal for two weeks—Nous
Command Code Agr / Agr-flash31B / 360M, no text generation—return typed values with per-option probabilities for tool calls and routingCommand Code

All results in the table are self-reported. Kolibri's dataset is unreleased, and its agentic evals sit well below Qwen.

Different perspectives: is Reflection Beam a breakthrough for American open weights or a DeepSeek V3 rerun?

Reflection Beam is a 501-billion-parameter model with 23 billion active parameters, a self-reported 80.9 on SWE-bench Verified and weights under Apache 2.0. American open models at this scale are rare, so the debate is not about whether it shipped but about whether Beam catches up with the Chinese leaders.

For. Artificial Analysis, which had early access, expects Beam to be among the most token-efficient open models for its intelligence. Elie Bakouch notes that Beam has better held-out code perplexity than DeepSeek V4.

Neutral. Nathan Lambert groups Beam with releases from Nvidia and Thinking Machines as strong US models that still trail their Chinese counterparts. Observers place Beam around GLM-5.2 level.

Against. Teortaxes calls Beam an iso-FLOP replication of DeepSeek V3 and infers about 1.3 billion RL sandboxes over 4 weeks. On some benchmarks Beam trails DeepSeek V4 Flash, and Bakouch estimates pretraining hardware utilization at only about 12%. Even before the announcement, r/LocalLLaMA recalled the Reflection 70B episode and doubted the company would ship weights at all.

Tools and techniques

In brief