xAI, OpenAI and Anthropic cosign AEF-1 — AI Digest

xAI, OpenAI and Anthropic cosign AEF-1 — AI Digest

xAI, OpenAI and Anthropic have cosigned AEF-1, the first shared standard for independent AI audits. The pacing fight produced a DeepMind resignation and cartel accusations, while DeepSeek-V4.1-Flash reached the open-model top three at $0.06–0.07 per task.

Today's main stories

AEF-1: xAI, OpenAI and Anthropic all cosign one standard for outside audits

The AI Evaluator Forum published AEF-1, a baseline standard for independent third-party evaluation of AI systems, and xAI, OpenAI and Anthropic all signed on (AI Evaluator Forum, Latent Space breakdown). The document covers the five points where "independence" usually breaks down: what model access an evaluator gets, how conflicts of interest are disclosed, who funds the work, when an evaluator must recuse itself, and what gets published. Third-party evaluation means testing a model's capabilities and risks through an outside organisation rather than the lab that built it, under access and disclosure rules agreed in advance. Until now every lab negotiated its own private terms with evaluators, which is exactly why the argument over whether outside audits can be trusted kept ending in "we can't see the contract." Three competing labs agreeing on one text is the first shared denominator that argument has had all year.

The pacing split: a DeepMind resignation on one side, cartel accusations on the other

The public fight over whether labs should slow capability growth hardened into two camps with incompatible logic on 14–15 September 2026. For a pause: researcher Bilal Chughtai announced he had left Google DeepMind, arguing progress is outrunning alignment and calling for both pacing and more transparency (post). Against it: Aidan Gomez of Cohere, whose objection is about power rather than safety — he does not want a world where a handful of Silicon Valley companies become the AI gatekeepers for governments (post); Cohere itself added that public extinction-risk discourse drifts into science fiction (post). The open-weight community opened a third front: an r/singularity thread drew 2,173 responses around the claim that "safe pacing" is a Western frontier-lab cartel wearing a safety badge (thread).

DeepSeek-V4.1-Flash: third among open models at $0.06–0.07 per task

Agent Arena placed DeepSeek-V4.1-Flash third among open models and on the Pareto frontier: +4.87% net improvement at a median cost of $0.06–0.07 per task (measurement, fuller data). In the same run, Hy4 preview scores +4.96% at $0.22 and Kimi K3 (Max) reaches +6.39% at $0.77. So for a difference measured in tenths of a percent, Hy4 charges three times more and the leader charges eleven times more. The caveat came from practitioners: on r/LocalLLaMA commenters noted the model also ranks high on a hallucination benchmark, and that winning one private test is not broad superiority (thread). At its 10 September launch Artificial Analysis scored it 40 on its intelligence index at $0.30 per 1M input and $1.20 per 1M output tokens (measurement).

The harness saves more than a model upgrade does

LangChain changed the file-reading format inside its agent and got 15% fewer edit_file errors and 10% fewer input tokens — with no model change at all (report). That is precisely the thesis the AI Engineer World's Fair turned into its own track: when an agent fails in production the culprit is usually not the model but everything around it — the harness, permissions, tool routing, memory, retries, kill switches and monitoring (track announcement). A harness is the layer around the model that runs the agent loop: tool calls, memory, retries and stopping conditions. Omar Shorbagy published a practical build-from-scratch guide: separate inference, tools and loop; keep prompts minimal; log aggressively; test on diverse tasks; only then add memory, skills and subagents (guide, follow-up on cost). Ready-made harness patterns for n8n and Make are in the AI SKILLS automation templates.

Numbers and facts

Different views: who benefits from pacing the frontier?

The argument is not about whether autonomous agents are dangerous. It is about who gets to set the speed. Three positions voiced on 14–15 September 2026.

For. Bilal Chughtai left Google DeepMind saying capability is outrunning alignment, and asked for both pacing and transparency (post). Daniel Kokotajlo published Dan Selsam's statement, which goes further and frames the moment as situationally acute (post).

Neutral. AEF-1 offers a third route: stop arguing about speed and make what already happens verifiable — evaluator access, funding and disclosure (AI Evaluator Forum). The harness-engineering track at AI Engineer World's Fair says the same thing in engineering terms: control comes from permissions, retries and kill switches, not declarations (announcement).

Against. Aidan Gomez objects to a world where a few companies become AI gatekeepers for states (post). The r/singularity thread pointed at a practical hole: banning open-weight releases would not stop model proliferation, it would push replication toward extraction and distillation of closed systems through their APIs (thread).

Tools and techniques

In brief