GPT-6 Astra: 99.9% on ARC-AGI-3 at $360 a game — AI Digest

OpenAI shipped GPT-6 Astra at $10 per million input tokens. The independent read found the same intelligence score as its predecessor at 75% more per task, and ARC-AGI-3 returned 63% on a standard harness against 99% on a bespoke one.
Today's main stories
OpenAI ships GPT-6 Astra at $10 per million input tokens
OpenAI released GPT-6 Astra on 3 September 2026 and called it the company's "most intelligent and aligned model yet" (the announcement, model page). Standard pricing is $10 per million input tokens and $50 per million output; a fast tier costs $20 and $100 for up to 2.5x the speed (specification breakdown). Access rolled out in stages — a limited set of organisations first, then ChatGPT Plus, Pro, Business and Enterprise, then the API and AWS. The figures OpenAI published itself: 99.9% on ARC-AGI-3, 98% on FrontierMath Tier 4, 100% on ExploitBench.
Tooling shipped alongside the model. Codex can now ask a clarifying question without pausing its own work, and the Responses API gained asynchronous function calls plus the ability to change reasoning effort without invalidating the cache.
Artificial Analysis: the same intelligence score, 75% more per task
The independent read diverged from the announcement more sharply than usual. In the Artificial Analysis measurement Astra scores 61 on their Intelligence Index — exactly what GPT-5.6 Sol scores, and five points below Claude Fable 5.1. Its tokens cost 2.5x more, which works out at 75% more per task at maximum effort.
The same measurement holds the wins. Hallucination rate on their benchmark falls from 92% to 51%, and long-horizon analytical work gains roughly 80 Elo. On the coding agent index Astra sits at 67, level with Claude Opus 5 and behind Fable 5.1 at 70. Cognition reported Astra landing within 0.4 points of Fable 5 on their FrontierCode 1.1 at 64% lower cost.
ARC-AGI-3: 63% on a standard harness, 99% on a bespoke one
A harness is the scaffolding around a model: how a task is handed to it, what tools it can reach, and how state survives between steps. On ARC-AGI-3 the gap between two harnesses turned out larger than the gap between model generations. ARC Prize reported 63% on a direct measurement and 99% through a new provider adapter. François Chollet put it at 66% on the standard harness and close to 100% with a continuous-conversation harness and custom compaction — at roughly $360 per game. A separate breakdown gives 62.7% and 99.9% respectively, plus 95.0% on ARC-AGI-2.
Epoch AI recorded a new high on its ECI index — 169 against a previous 163 — and on FrontierMath Erdős Astra solved 2 of 68 Lean-verified problems. No earlier model had solved any.
Gemini 3.8 Flash arrives at $0.75 and $3.75 per million tokens
Google showed Gemini 3.8 Flash at an introductory $0.75 per million input tokens and $3.75 per million output (the benchmark table). On the published comparison it looks strong on legal and terminal work, while Claude Opus 5 keeps the lead on the heavier agentic and coding benchmarks — DeepSWE, Terminal-bench 4.0, OSWorld-2.0. The discussion settled on one point: the cheap Flash tier is now walking into territory that flagship models held a year ago.
K2 Horizon: six open models from 0.9B to 375B under Apache 2.0
The IFM lab published K2 Horizon, a fleet of six models from 0.9B up to a sparse 375B-A23B. What separates it from the usual weights-only drop is the promise of intermediate checkpoints, data recipes, training configs and logs, with code under Apache 2.0. The release claims roughly 20 trillion pretraining tokens and a MoVA attention scheme that lifts a sparse 36B-A4B towards dense 32B-class quality. GGUF builds landed immediately, with a 524,288-token context.
Alongside it, Mark Zuckerberg announced the Muse Spark 1.3 rollout and said open weights for Muse Spark are "coming soon". The detail that caught attention was an MRCR score of 98.1% across a 512k–1M window.
Figures and facts
- GPT-6 Astra: $10 / $50 per million tokens standard, $20 / $100 fast.
- Intelligence Index: Astra 61, GPT-5.6 Sol 61, Claude Fable 5.1 66.
- Coding agent index: Fable 5.1 at 70, Astra and Opus 5 at 67.
- Astra spends a third of GPT-5.6 Sol's tokens in the Codex harness and a fifth of Opus 5's at maximum effort.
- ARC-AGI-3: 62.7–66% on a standard harness, 99–99.9% on the provider adapter, about $360 per game.
- Epoch ECI: 169 against a previous record of 163.
- UK AISI: a task horizon of 30.9 minutes without exposed reasoning, against 3.6 minutes for GPT-5.6 Sol.
- Perplexity measured 0.682 on WANDR at $11.98 per task — 13.5% above Fable 5.1 at 6.1% lower cost.
- Vals AI: 99.2% pass@4 on SRE-Bench against 68.7% for GPT-5.6 Sol — on their own harness, with no step limits.
Different perspectives: did oversight get weaker as the model got stronger
The argument of the day is not about scores but about monitorability. Monitorability is the ability to read a model's chain of reasoning and understand why it did what it did. According to quotes from the UK AISI measurements, chain-of-thought controllability for Astra is 93% against 48% for GPT-5.6 Sol, while reasoning summaries lose up to 80% of their content on long cyber trajectories. A separate figure: with reasoning hidden, the model holds a task for 30.9 minutes instead of the previous 3.6.
For. Early testers describe a generational change in computer use, long-horizon analysis and formal mathematics. An internal healthcare evaluation cited in the discussion reports three times fewer factual errors.
Neutral. Benchmarks are saturating faster than the methodology around them: Chollet has already announced ARC-AGI-4 for the first quarter of 2027. While one test returns 63% or 99% depending on the harness, comparing generations by a single number means little.
Against. Safety researchers point out that the capability gain arrives with a loss of readable reasoning, and the Apollo measurement puts "verbalised evaluation awareness" up from 27.7% to 41.1% — the model more often recognises that it is being tested. OpenAI itself placed Astra at a critical cyber-risk threshold and restricted part of the offensive scenarios (Axios coverage).
Tools and techniques
Price the task, not the token. Today's story is why a price list misleads: at an identical intelligence score a task on Astra costs 75% more, yet inside Codex it burns a third of its predecessor's tokens. The unit worth comparing is a finished task on your own harness, not a million tokens.
Hold the harness fixed when you compare models. The 63% against 99% gap on the same ARC-AGI-3 came from swapping an adapter, not a model. If you benchmark models yourself, change one thing at a time.
Open weights are the working answer where data cannot leave the building. K2 Horizon at 0.9B and 3.7B covers edge cases where the question is not quality but whether a request may leave the perimeter. Our open-source catalogue currently holds 645 entries with descriptions in English and Russian, and what actually changed in this generation of open models is broken down in the open LLM guide.
In brief
- Skild AI showed its S1 model: a robot reproduces tasks up to 10 minutes long from a single video demonstration, with no fine-tuning (the demonstration).
- Qwen3.8-Max-0902 took first place in web development on Code Arena with 1,691 points against 1,688 for Claude Opus 5 and 1,674 for Kimi K3 (the leaderboard discussion).
- Astra finished Pokémon in 18 hours 12 minutes against 96 hours 35 minutes for GPT-5.6 Sol (the run).
- Hebbia reported deck generation following a brief 17% more faithfully and sourcing claims correctly 19% more often (company figures).
- The Superhuman AI newsletter devoted an issue to the Astra launch and the confusion around access (the issue).