Open Jev in two days: 90% against 93% — AI Digest

Open Jev in two days: 90% against 93% — AI Digest

Jev got open reproductions within two days: Bespoke Nimble on Qwen3.5-9B scores 90% against 93% for the original. Claude Code now reads AGENTS.md, and Anthropic and Accenture are putting at least a billion dollars into independent evaluation.

Today's main stories

Open clones of Jev shipped two days after the announcement

Jev, the decision model shown on 15 September 2026, got open reproductions within two days. Bespoke Nimble is an "open Jev" built as a LoRA fine-tune on top of Qwen3.5-9B using synthetic contrastive data curation and constrained decoding: on its curated eval the base Qwen rose from 66% to 90%, against 93% for Jev itself, at 100ms on an H100 and runnable locally (Bespoke announcement). At the other end of the range sits Kev-0.5B, built on Qwen2.5-0.5B and small enough to run on a MacBook Pro (announcement). In parallel Jev turned up as an eval model inside Braintrust with a claimed 400x lower scoring cost (Ankur Goyal). The speed of reproduction says the value was never in the weights — the framing "a discriminative model instead of a generative one" replicates on open bases almost immediately.

Claude Code now reads AGENTS.md, turning a convention into a standard

Claude Code version 2.1.277 now looks for an AGENTS.md file when a project has no CLAUDE.md, with a config-level toggle (announcement). AGENTS.md is an instruction file for a coding agent at the root of a repository: what the project is, how to build it, which rules to follow. The significance is industrial rather than technical: a format that some tools supported and others ignored has been acknowledged as shared. The practical payoff is immediate — no more shim files that exist only to point one format at the other (Simon Willison). For teams running several agents side by side it means one instruction rather than three copies drifting apart.

Anthropic and Accenture will put at least a billion dollars into independent evaluation

Anthropic announced a partnership with Accenture on independent evaluation of frontier AI, with both sides expecting to invest at least $1B over five years to build the capacity for it (Anthropic announcement). This continues the outside-oversight argument running since 14 September, and it is the first commitment at this scale. The reaction ranged from reserved to hostile: critics question whether a consulting firm is the right vehicle for model red-teaming and safeguard assessment, while Transluce argues the real issue is not the size of the cheque but the conditions of independence and meaningful access (Transluce). Regulatory pressure is building alongside it: California's governor signed an executive order convening an expert panel to recommend stronger AI safety laws, including possible kill switches and embedded outside monitors (The Rundown).

Teams are splitting models by role: frontier to plan, cheap to execute

Practitioners keep describing the same stack shape: a strong model plans, a cheap one executes. One published stack runs GPT 5.6 Sol XHigh for planning, GLM 5.3 Flash for implementation and DeepSeek V4.1 Flash for everything else (Ahmad Osman). The economics of that move are measurable: an internal knowledge-base pipeline walked Opus → Sonnet → GLM 5.2 → GLM 5.3 Flash and cut spend by roughly two orders of magnitude since spring (Kyle Russell). The price of the top tier shows up in concrete numbers too: on GPT-6 Astra a high-effort setting lifted detection from 82.1% to 83.6% mAP@50 while roughly doubling per-image cost from $0.050 to $0.101 and raising latency from 11s to 32s (Roboflow measurement). One and a half points for double the price is exactly the trade that pushes teams to split roles.

Numbers and facts

Different views: is Jev a new primitive or a rediscovery of classifiers

The argument is not about the numbers but about whether any of this is new, and it splits people noticeably by how long they have been in machine learning. Three positions were voiced on 18 September 2026.

It is an architectural shift. Han Xiao argues Jev can pull tool calling, routing and MCP-style decisions back from small generative models toward discriminative ones (position). The same thought extended further: a judgment layer with near-zero marginal cost running on the device itself, for notifications, UI adaptation and sensor-driven decisions (continuation).

It depends where you came from. Mikhail Parakhin noted the split: people who entered the field after ChatGPT treated Jev as a revelation, while those doing machine learning before it were more puzzled by the excitement (observation).

The demos are about speed, not quality. The sharpest objection: almost every demonstration leans on latency, and there is still no standard benchmark for this class of model, so there is nothing to compare against (objection).

Tools and techniques

In brief