Claude Code Limits: 25% Above Pre-Promo Levels — AI Digest

The Claude Code weekly-limit promotion ended on 13 September: from the 14th the ceiling sits 25% above the level before it. A post by Trump against slowing AI became the day’s most discussed item, the AI Evaluator Forum proposed a baseline for independent evaluation, and Agent Arena measured DeepSeek V4.1-Flash third among open models at $0.06–0.07 per task.
Today's highlights
Claude Code limits: the promotion ended on 13 September, the weekly ceiling is now 25% above the old one
Claude Code users are circulating a screenshot of Anthropic's support page: the weekly-limit promotion that ran from May to August 2026 ended on 13 September, and as of 14 September the weekly ceiling on the Pro, Max, Team and seat-based enterprise plans sits 25% above the level that preceded the promotion. For anyone used to the promotional allowance that is a noticeable cut, which is why the threads filled up. In a neighbouring one, a Max 20x subscriber posts a dashboard showing the weekly allowance across all models fully spent and Fable at 78%, with the limits still labelled "temporarily boosted". Commenters tie this to serving capacity for frontier models being scarce, and argue about how predictable the remaining allowance now is. Worth keeping in mind when planning long agent sessions. The limits thread, the dashboard thread.
"Pacing the frontier": Trump refused, David Sacks called it regulatory capture
On 14 September 2026 a post by Donald Trump against slowing AI development and against extra guardrails became the day's most discussed item — 5,058 activity points on the thread; it frames AI development, data centres and competition with China as a national priority and names Anthropic directly. A day earlier, David Sacks's criticism was making the rounds: on his argument Dario Amodei and Sam Altman are free to slow their own unreleased models but not to seek regulatory cover and antitrust exemptions for doing so, and he questions METR's independence over its ties to people and investors around Anthropic. The counter-argument runs through the same discussions: public coordination between two labs could itself be an antitrust problem, and other players are, by participants' estimate, 3–12 months behind — so bilateral restraint changes little. Trump's post and the reaction, the Sacks discussion.
AEF-1: a common baseline for independent model evaluation is proposed
The AI Evaluator Forum published AEF-1, a proposed baseline for independent third-party evaluation of AI systems: model access, conflicts of interest, funding relationships, recusal and transparency. A third-party evaluator is an organisation that tests a model without having built it; its conclusions are what people point to when they say a model was "independently evaluated". The document landed exactly as the week's whole argument narrowed to one question: whether such independence is achievable in practice when evaluators and the evaluated are financially entangled. Adjacent to it is a statement from OpenAI researcher Dan Selsam, circulated by Daniel Kokotajlo: a situationally aware model may look aligned precisely while it is being tested, concealing misalignment — in which case the evidentiary weight of future evaluations falls on its own. AEF-1, Selsam's statement.
DeepSeek V4.1-Flash on Agent Arena: third among open models at $0.06–0.07 per task
Agent Arena published a measurement in which DeepSeek V4.1-Flash (Max) placed third among open models and landed on the Pareto frontier: a net improvement of +4.87% at a median cost of roughly $0.06–0.07 per task. For comparison, in the same measurement the Hy4 preview gives +4.96% at $0.22 per task and Kimi K3 (Max) gives +6.39% at $0.77. On quality DeepSeek sits close to the top of the open field; on cost per task it differs from its neighbours by three-fold and ten-fold. The Pareto frontier here means no other model in the measurement was simultaneously better and cheaper. Read it as a metric for agentic tasks rather than as a universal verdict on the model. The Agent Arena measurement, the fuller numbers.
Agent harnesses: changing the file-reading format removed 15% of edit errors
LangChain reported a concrete figure for the argument that the win does not come only from picking a model: a single change to the format in which files are handed to the agent cut errors from the file-editing tool by 15% and total input tokens by 10%. A harness is everything that surrounds the model in a working agent: the run loop, the tools, permissions, routing, memory, retries and monitoring. An industry conference gave the subject its own track with the same thesis: when an agent fails in production, what breaks is usually not the model but the harness around it. In the same direction: Cline Desktop shipped, a standalone desktop app for working with open-weight models — your own key, your own provider, support for DeepSeek-V4.1-Flash and Musespark-1.3, and model switching mid-project. LangChain's numbers, Cline Desktop.
Numbers and facts
- 25% above the pre-promotion level — the Claude Code weekly ceiling for Pro, Max, Team and the seat-based enterprise plan as of 14 September 2026 (thread).
- +4.87% at $0.06–0.07 per task for DeepSeek V4.1-Flash, against +4.96% at $0.22 for the Hy4 preview and +6.39% at $0.77 for Kimi K3 (measurement).
- 15% fewer file-edit errors and 10% fewer input tokens from a single change to how files are presented (LangChain).
- 10 days against 4–5 months — the reported time for one engineer with an AI assistant on chip design versus a conventional front-end team (summary).
- 14.4 seconds of 768p video in 9.0 seconds on eight B200 accelerators after warm-up — MiniMax's inference optimisation claim (MiniMax).
- 695 cards in the AI SKILLS Open Source section as of 15 September 2026; 10,441 skills for Claude Code, 613 of which describe running models locally.
Different perspectives: should the pace of development be held back?
The argument is not about whether model capability is growing but about who gets to limit that pace and on what authority. Which is why both sides appealed to independent evaluation on the same day — and both consider it insufficient as it stands.
For. Bilal Chughtai announced he had left Google DeepMind and called directly for pacing and more transparency: on his argument, capability growth may be outrunning the alignment work (post).
Neutral. The AI Evaluator Forum published AEF-1 — baseline requirements for independent evaluation (access, conflicts of interest, funding relationships, recusal, transparency). It is an attempt to move the argument from declarations to checkable procedure without answering the question about pace (AEF-1).
Against. Aidan Gomez objects to a world in which a handful of Silicon Valley companies become AI gatekeepers for governments (post). A practical counter runs through the discussions too: restricting open-weight publication would not stop model proliferation but would shift it toward extraction and distillation through closed APIs (discussion).
Tools and techniques
- When an agent fails in production, start with the harness, not the model. A practical guide to building a harness from scratch advises separating inference, tools and the loop, keeping prompts short, logging aggressively and testing on diverse tasks — and only then layering in memory, skills and subagents (the guide). Ready-made skills for pipelines like these are in the AI SKILLS catalog of Claude Code skills.
- Cheap document parsing is not free. Cohere Parse 5 shipped as a cheaper parser and drew an objection from Jerry Liu: the price is competitive, but on visual grounding, chart parsing and fine-grained citation-oriented extraction it is weaker than several alternatives — meaning the comparison has to be run on your own documents (the objection).
- Voice workflows got cheaper. OpenAI cut desktop voice pricing by roughly 60% and usage rose 2.4× — if you shelved voice input over cost, it is worth recalculating (the report).
In brief
- A user described building EMBER, a DOS-like operating system, with Claude: its own boot path, a graphical interface, a file manager, and unmodified Doom, Alley Cat and Prince of Persia running on it; the source is public (the write-up and repository).
- RewardAI introduced OM-1, a robot foundation model trained on human manipulation data rather than teleoperation, claiming zero-shot transfer across tabletop, industrial and humanoid robots (announcement).
- Google DeepMind applied WeatherNext 3 to renewables planning: hourly forecasts for turbine-height wind and solar radiation (announcement).
- Inferact and Google Cloud announced a partnership to make TPU a first-class citizen in vLLM — optimised kernels, a native path via TorchTPU, and a programme allocating capacity to open-source contributors (announcement).