Open Jev in two days: 90% against 93% — AI Digest

Jev got open reproductions within two days: Bespoke Nimble on Qwen3.5-9B scores 90% against 93% for the original. Claude Code now reads AGENTS.md, and Anthropic and Accenture are putting at least a billion dollars into independent evaluation.
Today's main stories
Open clones of Jev shipped two days after the announcement
Jev, the decision model shown on 15 September 2026, got open reproductions within two days. Bespoke Nimble is an "open Jev" built as a LoRA fine-tune on top of Qwen3.5-9B using synthetic contrastive data curation and constrained decoding: on its curated eval the base Qwen rose from 66% to 90%, against 93% for Jev itself, at 100ms on an H100 and runnable locally (Bespoke announcement). At the other end of the range sits Kev-0.5B, built on Qwen2.5-0.5B and small enough to run on a MacBook Pro (announcement). In parallel Jev turned up as an eval model inside Braintrust with a claimed 400x lower scoring cost (Ankur Goyal). The speed of reproduction says the value was never in the weights — the framing "a discriminative model instead of a generative one" replicates on open bases almost immediately.
Claude Code now reads AGENTS.md, turning a convention into a standard
Claude Code version 2.1.277 now looks for an AGENTS.md file when a project has no CLAUDE.md, with a config-level toggle (announcement). AGENTS.md is an instruction file for a coding agent at the root of a repository: what the project is, how to build it, which rules to follow. The significance is industrial rather than technical: a format that some tools supported and others ignored has been acknowledged as shared. The practical payoff is immediate — no more shim files that exist only to point one format at the other (Simon Willison). For teams running several agents side by side it means one instruction rather than three copies drifting apart.
Anthropic and Accenture will put at least a billion dollars into independent evaluation
Anthropic announced a partnership with Accenture on independent evaluation of frontier AI, with both sides expecting to invest at least $1B over five years to build the capacity for it (Anthropic announcement). This continues the outside-oversight argument running since 14 September, and it is the first commitment at this scale. The reaction ranged from reserved to hostile: critics question whether a consulting firm is the right vehicle for model red-teaming and safeguard assessment, while Transluce argues the real issue is not the size of the cheque but the conditions of independence and meaningful access (Transluce). Regulatory pressure is building alongside it: California's governor signed an executive order convening an expert panel to recommend stronger AI safety laws, including possible kill switches and embedded outside monitors (The Rundown).
Teams are splitting models by role: frontier to plan, cheap to execute
Practitioners keep describing the same stack shape: a strong model plans, a cheap one executes. One published stack runs GPT 5.6 Sol XHigh for planning, GLM 5.3 Flash for implementation and DeepSeek V4.1 Flash for everything else (Ahmad Osman). The economics of that move are measurable: an internal knowledge-base pipeline walked Opus → Sonnet → GLM 5.2 → GLM 5.3 Flash and cut spend by roughly two orders of magnitude since spring (Kyle Russell). The price of the top tier shows up in concrete numbers too: on GPT-6 Astra a high-effort setting lifted detection from 82.1% to 83.6% mAP@50 while roughly doubling per-image cost from $0.050 to $0.101 and raising latency from 11s to 32s (Roboflow measurement). One and a half points for double the price is exactly the trade that pushes teams to split roles.
Numbers and facts
- Bespoke Nimble: 66% → 90% against 93% for Jev, 100ms on an H100 (announcement).
- Astra, high-effort mode: 82.1% → 83.6% mAP@50 with cost $0.050 → $0.101 and latency 11s → 32s (Roboflow).
- CUA-Bench: every frontier model scores below 20% on real-time keyboard and mouse control across six games (Vals AI).
- Turbo-dLLM: 2.48x faster diffusion-LLM training at 512K context and 7.59x at 1M on 8× H100 (announcement).
- ABC-130K: 3,500 hours of teleoperation, 130K+ episodes, 195 tasks, collected on an $8K bimanual rig, sim-to-real correlation r = 0.91 (details).
- Our own data: as of 19 September 2026, 100 of the 10,441 skills in the AI SKILLS catalogue mention
CLAUDE.mdand only 52 mentionAGENTS.md— the older format still leads two to one. Those figures come from our database and appear in no primary source (Claude Code skills catalogue).
Different views: is Jev a new primitive or a rediscovery of classifiers
The argument is not about the numbers but about whether any of this is new, and it splits people noticeably by how long they have been in machine learning. Three positions were voiced on 18 September 2026.
It is an architectural shift. Han Xiao argues Jev can pull tool calling, routing and MCP-style decisions back from small generative models toward discriminative ones (position). The same thought extended further: a judgment layer with near-zero marginal cost running on the device itself, for notifications, UI adaptation and sensor-driven decisions (continuation).
It depends where you came from. Mikhail Parakhin noted the split: people who entered the field after ChatGPT treated Jev as a revelation, while those doing machine learning before it were more puzzled by the excitement (observation).
The demos are about speed, not quality. The sharpest objection: almost every demonstration leans on latency, and there is still no standard benchmark for this class of model, so there is nothing to compare against (objection).
Tools and techniques
- Cut your agent's tool set down to four. The "Harness Tax" analysis shows a simple set — read, write, edit, bash — reaching the Pareto frontier on benchmark performance while trimming unnecessary spend (analysis). A companion paper on harness design for coding agents makes the same point: benchmark outcomes are increasingly shaped by harness structure, context setup, turn budgets and tool affordances rather than the base model (paper).
- Give the decision model a browser. The most convincing use of Jev so far is browser control: a LangChain and Jev pairing held up on tasks like the Wikipedia game and structured routines (breakdown), and Cline shipped a plugin that hands Jev a browser inside the editor (Cline). Ready-made automation scenarios for that shape of workflow sit in the AI SKILLS automation templates.
- Test your agent on computer use, not only on text. CUA-Bench measures real-time keyboard and mouse work across six games, three of them held back from training, and every frontier model is still under 20% (announcement). The first family of small computer-use models is open as well — CUA-S1-FORMS (announcement).
In brief
- On the RSI-Exam leaderboard GPT-6 Astra holds first place at 0.5126 and Claude Fable 5.1 enters second at 0.4813, with no model yet reaching the calibrated reference (update).
- Epoch AI reported another open FrontierMath problem solved in an interactive session with GPT-6 Astra (Epoch AI); the SAIR foundation launched Open Math Model, open models and tooling for mathematics (announcement).
- Grok Voice Transcribe 2.0: 2.7% word error rate on streaming final transcripts against 3.9% for its predecessor, 0.49s after end of speech, pricing unchanged at $0.20 per streaming hour (Artificial Analysis).
- fal launched H3 Max Lip Sync with an 11-second median generation time and first place on both speed and quality in its own evals (announcement); the model was built by pushing diffusion RL into a verifiable lip-sync task (details).
- ProgramAsWeights: you describe a function in English, compile it once, and then run a small neural program locally on the CPU with Wi-Fi off; code and models are public (announcement).
- DeepSeek V4.1 Flash improved noticeably in quality after its context was extended to 1M tokens — evidence for the argument that agents are above all context hungry (observation).