Data Gave 12x, Models 3.7x in LLM Progress — AI Digest

Data Gave 12x, Models 3.7x in LLM Progress — AI Digest

A 2019–2025 measurement: data delivered 12.0x of compute-efficiency gain against 3.7x from architectures. GPT-6 Astra one week in scores 67 on the Coding Agent Index against Fable 5.1's 70, and costs 75% more per task.

Today's highlights

Data delivered 3.24x more progress than model architectures

A study published on 8 September 2026 measured how much of the progress in language models came from data versus architecture between 2019 and 2025: data improvements delivered 3.24x more compute-efficiency gain than model improvements. In absolute terms that is 12.0x from data against 3.7x from models at a 1e19 FLOPs training budget. Pretraining is the first stage of training, where a model learns from a large text corpus before any task-specific tuning. The researchers took one model recipe and one public corpus per year, trained every combination, and scored the results on OLMES, an aggregate of ten relatively easy benchmarks. The two contributions turned out to be almost independent: 88% of the variance in OLMES scores is explained by simply adding the data effect and the model effect. For scale, the 2019 OpenWebText corpus held roughly 9 billion tokens; by 2025 open corpora such as UltraFineWeb were orders of magnitude larger and filtered by a trained classifier. Full write-up with charts.

GPT-6 Astra one week in: "a show horse, not a workhorse"

A week after launch, the hands-on verdict on GPT-6 Astra diverges from its benchmark record: the model sets records while users describe it as uneven. The author of Ben's Bites burned through 4 billion tokens over a weekend and, by his own account, finished nothing. Kieran Klaassen's summary is "it is a show horse, not a workhorse"; developer Theo writes "I both love and absolutely detest this model." Astra still wins ARC-AGI-3, posts the top score on Zapier's AutomationBench and costs the same as Claude Fable 5.1. Its computer use draws particular attention: Astra draws a portrait in Canva the way a person would with a mouse, and is fast enough to play piano. Inside Codex the model now asks questions without blocking: it continues everything that does not depend on your answer, then folds the answer in without losing the original task. A week of hands-on use.

Astra saves tokens yet costs more per task

An independent measurement by Artificial Analysis found Astra uses 70% fewer tokens than GPT-5.6 Sol, yet its pricing makes it 75% more expensive per task at maximum effort. Astra scores 67 on the Coding Agent Index, level with Claude Opus 5 and Fable 5, while Fable 5.1 leads at 70. In the Codex harness Astra spends one third of GPT-5.6 Sol's tokens and one fifth of Claude Opus 5's at maximum effort, and reaches the same score for less than half the cost of Claude Fable 5. On the Intelligence Index Astra scores 61, matching GPT-5.6 Sol and trailing Claude Fable 5.1 by 5 points. Its hallucination rate on that benchmark falls from 92% to 51% at maximum effort while accuracy rises 4 points. Regressions exist too: roughly 80 Elo down on GDPval-AA v2 and 2–3 points on τ³-Banking, SciCode and AA-LCR. Full measurement thread.

OpenAI says it hit its "automated research intern" goal

OpenAI announced it has reached its own "automated research intern" milestone: its agents now take on research tasks that would occupy a skilled person for several days. Humans still set the direction and judge the output. The company's next stated target is a fully automated AI researcher by March 2028. OpenAI is separately describing the internal advantage it gets from its own research agents in a Rundown piece dated 8 September 2026. The claim matters because it moves the argument about agent usefulness from demos toward something measurable: a named class of task and a named horizon. It cannot be verified from outside yet — no independent measurement of multi-day agent research work is publicly available. Source of the claim.

Anthropic is testing plugins that change how Claude Code behaves

Anthropic is testing a mechanism that lets Claude Code write plugins which change the interface, log actions, and restrict what agents are allowed to do. The feature has not shipped; the team is collecting feedback. The third capability matters more than the first two: constraining an agent's permissions inside the tool itself answers what companies keep asking for — a predictable agent rather than a maximally free one. For anyone already building extensions around Claude Code, the foundation shifts: today behaviour is configured through skill files, tomorrow through plugin code as well. How skills are structured is covered in our guide to Skills for Claude Code. Source.

Numbers and facts

Different perspectives: did data or architecture drive AI progress?

The measurement favours data: 12.0x of gain against 3.7x from architectures across 2019–2025. Yet the same researchers warn this is the wrong way to value architectural work.

For "data wins." The numbers are direct: at a 1e19 FLOPs budget, data delivered 3.24x more gain than models, and the two effects barely overlap. The naive reading is that the pretraining era of 2019–2024 was mostly good data engineering — extraction, curation, filtering (study).

Neutral. The experiment runs at small scale — up to 1e19 FLOPs — and covers pretraining only. A large share of today's gains comes from reinforcement learning and post-training, which this setup does not touch at all.

Against. The researchers argue against their own naive reading: the main contribution of architectural work was not compute efficiency but making large amounts of compute usable in the first place, as parameter counts, context lengths, run durations and cluster sizes all scaled up.

Tools and techniques

Plugins instead of config files: connecting an agent through an ordinary browser login. A review of Grok Bot describes plugin installation as signing in to a website: you find the plugin in a catalogue, a browser window opens, you log in, and the agent is connected — no MCP server JSON to edit, no API keys to paste. The author connected Freshdesk this way and built a bot that checks for new support tickets every fifteen minutes. OpenClaw 2.0, released this week, narrows the gap: its Quick Start reuses an existing Claude Code or Codex login, and its browser app moves setup and plugin management into a graphical interface. The underlying difference stands: OpenClaw gives you a gateway you run yourself, while Grok Bot supplies and operates the machine as part of the product (five days of hands-on use).

Subagents for data collection with integrity checks. The author of Ben's Bites assembled 107 million rows of public UK council spending data by asking an agent to find the sources: it spun up three subagents in separate threads, built a catalogue of 31 official sources, and logged every download with a checksum so changed source data could later be told apart from the original pull. The result was 8 councils, 1.8 million rows and £9 billion of spending in one clean CSV (build write-up). Comparable pipelines for routine data collection and processing sit in the automation template catalogue.

In brief