Open Jev Clone Hits 90% vs Jev's 93% — AI Digest

Jev, the decision model from TypeSafe AI, was cloned in the open within two days: a Qwen3.5-9B fine-tune scores 90% against the original's 93%. Claude Code v2.1.277 now reads AGENTS.md, and the Harness Tax analysis says four tools are enough.
Today's highlights
Jev: an open clone on Qwen3.5-9B scores 90% against the original's 93%
Jev, the decision model from TypeSafe AI, was reproduced in the open within two days, and the best clone trails the original by three percentage points. A decision model is not a text generator but a classifier with a calibrated probability: it answers by picking from a predefined schema rather than by writing prose. Bespoke Nimble is a LoRA fine-tune of Qwen3.5-9B trained on synthetically curated contrastive data; on its own curated eval the base Qwen improved from 66% to 90% against Jev's 93%, at a reported 100 ms on an H100 and usable locally. At the other end of the size range, Kev-0.5B builds on Qwen2.5-0.5B and runs on a MacBook Pro. The conclusion the discussion reached on 17 September 2026: this is not a chatbot story but a fast "System 1" sitting next to a large model — routing, citation selection, escalation. Braintrust already offers Jev as an eval model at roughly 400x lower scoring cost than prior setups, and calibrated probability matters wherever a decision has to be justified rather than guessed. On 21 September 2026 the TypeSafe AI founder laid out the thinking in a Latent Space interview: models for production, not for god.
Claude Code v2.1.277 reads AGENTS.md when CLAUDE.md is absent
Anthropic added AGENTS.md support to Claude Code, adopting an instruction-file convention that until now belonged to competing tools. Version v2.1.277 checks for AGENTS.md when a project has no CLAUDE.md, with the behaviour exposed as a config toggle. That treats AGENTS.md as an emerging cross-tool standard rather than someone else's house rule, and Simon Willison immediately named the payoff: a whole class of shim files — files that existed only to point one format at another — stops being necessary. For teams running several coding agents against one repository, it removes the chore of keeping two rule sets in sync. Our own catalogue held 10,453 Claude Code skills on 22 September 2026; how skills and the SKILL.md file are structured is covered in our guide to Claude Code skills.
Harness Tax: read, write, edit and bash reach the Pareto frontier
A measurement suggests four plain tools give a coding agent close to peak benchmark quality at lower spend. A harness is the scaffolding around an agent: its tool set, context layout, turn budget and stopping rules. The Harness Tax analysis argues that the minimal set — read, write, edit, bash — reaches the Pareto frontier on benchmarks, while anything beyond it tends to add cost more reliably than capability. Alongside it came the paper An Empirical Study of Harness Design for Coding Agents, which finds benchmark outcomes increasingly shaped by harness structure rather than the base model. The practical consequence is already visible in production stacks: frontier for planning, cheap for execution — GPT 5.6 Sol XHigh to plan, GLM 5.3 Flash to implement, DeepSeek V4.1 Flash for the rest.
Robotics: ABC-130K ships 3,500 hours of teleoperation from an $8K rig
Open manipulation data grew to 130,000 episodes collected on hardware costing eight thousand dollars. ABC-130K is described as the largest open teleoperation dataset to date: 3,500 hours, more than 130,000 episodes, 195 tasks, a bimanual setup costing $8K, with open hardware, training code, simulator and evaluation. The baseline science reports a sim-to-real correlation of r = 0.91 on task progress, meaning a policy can be screened in simulation instead of on the robot. The release answers the standard objection to robotics datasets — that the rig itself is the barrier to entry. For scale: the full ABC release includes more than 400 hours of simulation data across 24 tasks and 5,850 labelled policy-evaluation episodes.
Numbers and facts
- 2.7% WER for Grok Voice Transcribe 2.0 on streaming final transcripts, 0.49 s after end of speech, priced at $0.20 per hour streaming and $0.10 per hour non-streaming — Artificial Analysis measurement, up from 3.9% on its predecessor.
- Below 20% is where every frontier model lands on CUA-Bench, a real-time keyboard-and-mouse test across six games, three of them kept private.
- 82.1% → 83.6% mAP@50 detection gain for GPT-6 Astra on its high-effort setting in Roboflow's testing, while per-image cost went from $0.050 to $0.101 and latency from 11 to 32 seconds (details).
- 2.48x at 512K and 7.59x at 1M context — training speedups for diffusion language models on eight H100s in the Turbo-dLLM library.
- 105,493 downloads in six days for Swift Qwen 3.8 27B, the top fine-tune and ninth on Hugging Face Trending, up from 24,000 on day two.
- Under 6 GB is the size of Ternary Bonsai 2 (27B), a ternary-weight derivative of Qwen3.8-27B — nine times smaller than the original, with a claimed 98% of capability retained; commenters lined up to test that 98% claim.
- At least $1B over five years is what Anthropic and Accenture say they will invest in independent evaluation of frontier AI.
Different views: what is an autonomous agent demo actually worth?
The argument centres on Vals AI's claim that GPT-6 Astra completed Factorio: Space Age in about two wall-clock days and 165+ in-game hours. The dispute is not about the game but about whether such runs count as evidence of capability.
For. A long-horizon game with production chains resembles real work more closely than short benchmarks: the agent has to hold a plan across hundreds of steps and recover from its own mistakes.
Neutral. Individual integrations are measurable without games at all: browser use was called the strongest application for decision models so far, and a plugin for Cline hands such a model a browser inside the editor. That can be checked against your own tasks in an evening.
Against. The comments asked plain questions: 165 in-game hours in two days implies accelerated time or excluded pauses, and no replay was published for auditing. The same scepticism follows neighbouring demos — a virtual fusion reactor lab built in four hours drew criticism for shipping no equations or methodology: polished, but nothing to verify.
Tools and techniques
- Put a decision model in front of the large one. Open Jev clones are already testable locally: Laya at 421M parameters (a ModernBERT-large encoder plus a scoring head) and openjev can be wired into a router that decides which model gets the request. Deployable open-source builds are collected in the AI SKILLS open-source section.
- Cut the harness, then measure. Before adding another tool to an agent, test the Harness Tax hypothesis on your own task set: keep read, write, edit and bash, then compare quality and spend against your current configuration.
- Price the step, not the subscription. The "frontier to plan, cheap to execute" stack cut one practitioner's spend by roughly two orders of magnitude since spring: Opus → Sonnet → GLM 5.2 → GLM 5.3 Flash. Ready-made automation chains that take such models as nodes are in automation templates.
In brief
- California Governor Gavin Newsom signed an executive order convening an expert panel on stronger AI safety laws, including possible kill switches and embedded outside monitors — reported by The Rundown.
- On the RSI-Exam leaderboard GPT-6-astra holds first place at 0.5126 and Fable 5.1 entered second at 0.4813 — update from the test's author.
- Fal released H3 Max Lip Sync, claiming first place on both speed and quality in its own evals at a median generation time of 11 seconds — announcement.
- ProgramAsWeights compiles a function specified in English into a small neural program that runs locally on CPU with Wi-Fi switched off — author's description.
- Trycua open-sourced CUA-S1-FORMS, the first in a family of small "System One" computer-use models — announcement.
- Figure shipped Helix 2.5, another update to its robot control model — covered by Superhuman AI.
- Ethan Mollick's essay "The Overhang" describes the gap between what models can already do and what organisations have managed to adopt — the text.