xAI, OpenAI and Anthropic cosign AEF-1 — AI Digest

xAI, OpenAI and Anthropic have cosigned AEF-1, the first shared standard for independent AI audits. The pacing fight produced a DeepMind resignation and cartel accusations, while DeepSeek-V4.1-Flash reached the open-model top three at $0.06–0.07 per task.
Today's main stories
AEF-1: xAI, OpenAI and Anthropic all cosign one standard for outside audits
The AI Evaluator Forum published AEF-1, a baseline standard for independent third-party evaluation of AI systems, and xAI, OpenAI and Anthropic all signed on (AI Evaluator Forum, Latent Space breakdown). The document covers the five points where "independence" usually breaks down: what model access an evaluator gets, how conflicts of interest are disclosed, who funds the work, when an evaluator must recuse itself, and what gets published. Third-party evaluation means testing a model's capabilities and risks through an outside organisation rather than the lab that built it, under access and disclosure rules agreed in advance. Until now every lab negotiated its own private terms with evaluators, which is exactly why the argument over whether outside audits can be trusted kept ending in "we can't see the contract." Three competing labs agreeing on one text is the first shared denominator that argument has had all year.
The pacing split: a DeepMind resignation on one side, cartel accusations on the other
The public fight over whether labs should slow capability growth hardened into two camps with incompatible logic on 14–15 September 2026. For a pause: researcher Bilal Chughtai announced he had left Google DeepMind, arguing progress is outrunning alignment and calling for both pacing and more transparency (post). Against it: Aidan Gomez of Cohere, whose objection is about power rather than safety — he does not want a world where a handful of Silicon Valley companies become the AI gatekeepers for governments (post); Cohere itself added that public extinction-risk discourse drifts into science fiction (post). The open-weight community opened a third front: an r/singularity thread drew 2,173 responses around the claim that "safe pacing" is a Western frontier-lab cartel wearing a safety badge (thread).
DeepSeek-V4.1-Flash: third among open models at $0.06–0.07 per task
Agent Arena placed DeepSeek-V4.1-Flash third among open models and on the Pareto frontier: +4.87% net improvement at a median cost of $0.06–0.07 per task (measurement, fuller data). In the same run, Hy4 preview scores +4.96% at $0.22 and Kimi K3 (Max) reaches +6.39% at $0.77. So for a difference measured in tenths of a percent, Hy4 charges three times more and the leader charges eleven times more. The caveat came from practitioners: on r/LocalLLaMA commenters noted the model also ranks high on a hallucination benchmark, and that winning one private test is not broad superiority (thread). At its 10 September launch Artificial Analysis scored it 40 on its intelligence index at $0.30 per 1M input and $1.20 per 1M output tokens (measurement).
The harness saves more than a model upgrade does
LangChain changed the file-reading format inside its agent and got 15% fewer edit_file errors and 10% fewer input tokens — with no model change at all (report). That is precisely the thesis the AI Engineer World's Fair turned into its own track: when an agent fails in production the culprit is usually not the model but everything around it — the harness, permissions, tool routing, memory, retries, kill switches and monitoring (track announcement). A harness is the layer around the model that runs the agent loop: tool calls, memory, retries and stopping conditions. Omar Shorbagy published a practical build-from-scratch guide: separate inference, tools and loop; keep prompts minimal; log aggressively; test on diverse tasks; only then add memory, skills and subagents (guide, follow-up on cost). Ready-made harness patterns for n8n and Make are in the AI SKILLS automation templates.
Numbers and facts
- AEF-1 covers five areas: model access, conflicts of interest, funding, evaluator recusal and transparency of results (source).
- DeepSeek-V4.1-Flash: +4.87% at $0.06–0.07 per task versus +6.39% at $0.77 for Kimi K3 (Max) (Agent Arena).
- OpenAI cut desktop voice pricing by roughly 60% and usage rose 2.4× (summary).
- MiniMax H3: 14.4 seconds of 768p video generated in 9.0 seconds (MiniMax claim).
- Swift-Qwen3.8-27B from UkisAI: −58.3% thinking tokens and a 1.95× speed-up while keeping xhigh accuracy (breakdown).
- Cognichip ACI Enterprise: one engineer finished in 10 days work the company compares to 4–5 months for a traditional front-end chip design team (summary).
- Our own Open Source catalogue holds 696 active entries as of the morning of 16 September 2026, 23 of them added in the last seven days — the flow of open tools is not slowing in the exact week the industry argues about slowing it down.
Different views: who benefits from pacing the frontier?
The argument is not about whether autonomous agents are dangerous. It is about who gets to set the speed. Three positions voiced on 14–15 September 2026.
For. Bilal Chughtai left Google DeepMind saying capability is outrunning alignment, and asked for both pacing and transparency (post). Daniel Kokotajlo published Dan Selsam's statement, which goes further and frames the moment as situationally acute (post).
Neutral. AEF-1 offers a third route: stop arguing about speed and make what already happens verifiable — evaluator access, funding and disclosure (AI Evaluator Forum). The harness-engineering track at AI Engineer World's Fair says the same thing in engineering terms: control comes from permissions, retries and kill switches, not declarations (announcement).
Against. Aidan Gomez objects to a world where a few companies become AI gatekeepers for states (post). The r/singularity thread pointed at a practical hole: banning open-weight releases would not stop model proliferation, it would push replication toward extraction and distillation of closed systems through their APIs (thread).
Tools and techniques
- Cline Desktop — a standalone app for open-weight models with your own key and mid-project model switching; DeepSeek-V4.1-Flash and Musespark-1.3 are among the supported models (announcement).
- Automatic model tiers in GitHub Copilot — efficiency, balance and intelligence instead of manual switching (announcement), plus an
/askmode while the agent is still working (details). - Deep Research inside Gemini Live — start research by voice, then discuss the finished report in chat (Google announcement). Prompts for that kind of workflow live in the AI SKILLS prompt library.
In brief
- Arcee launched Forge with Bolt: opted-in Bolt Pro users get 50× more usage across open-weight models in exchange for anonymised development-session data, with weights promised for release afterwards (announcement).
- Inferact and Google Cloud are making TPU a first-class citizen in vLLM: production serving, optimised kernels and a native PyTorch path via TorchTPU (announcement).
- RewardAI introduced OM-1, a robot foundation model trained on human manipulation data rather than teleoperation, claiming zero-shot transfer across tabletop, industrial and humanoid robots (announcement).
- Google DeepMind applied WeatherNext 3 to renewables planning, with hourly forecasts for turbine-height wind and solar radiation (announcement).
- NVIDIA listed the RTX PRO 5500 Blackwell with 84 GB of ECC GDDR7 and 1,398 GB/s of memory bandwidth (discussion).
- Good Start Labs is turning games into training material for models — an interview with founder Alex Duffy on why games turned out to be a convenient training environment (Latent Space).