Open Jev Clone Hits 90% vs Jev's 93% — AI Digest

Open Jev Clone Hits 90% vs Jev's 93% — AI Digest

Jev, the decision model from TypeSafe AI, was cloned in the open within two days: a Qwen3.5-9B fine-tune scores 90% against the original's 93%. Claude Code v2.1.277 now reads AGENTS.md, and the Harness Tax analysis says four tools are enough.

Today's highlights

Jev: an open clone on Qwen3.5-9B scores 90% against the original's 93%

Jev, the decision model from TypeSafe AI, was reproduced in the open within two days, and the best clone trails the original by three percentage points. A decision model is not a text generator but a classifier with a calibrated probability: it answers by picking from a predefined schema rather than by writing prose. Bespoke Nimble is a LoRA fine-tune of Qwen3.5-9B trained on synthetically curated contrastive data; on its own curated eval the base Qwen improved from 66% to 90% against Jev's 93%, at a reported 100 ms on an H100 and usable locally. At the other end of the size range, Kev-0.5B builds on Qwen2.5-0.5B and runs on a MacBook Pro. The conclusion the discussion reached on 17 September 2026: this is not a chatbot story but a fast "System 1" sitting next to a large model — routing, citation selection, escalation. Braintrust already offers Jev as an eval model at roughly 400x lower scoring cost than prior setups, and calibrated probability matters wherever a decision has to be justified rather than guessed. On 21 September 2026 the TypeSafe AI founder laid out the thinking in a Latent Space interview: models for production, not for god.

Claude Code v2.1.277 reads AGENTS.md when CLAUDE.md is absent

Anthropic added AGENTS.md support to Claude Code, adopting an instruction-file convention that until now belonged to competing tools. Version v2.1.277 checks for AGENTS.md when a project has no CLAUDE.md, with the behaviour exposed as a config toggle. That treats AGENTS.md as an emerging cross-tool standard rather than someone else's house rule, and Simon Willison immediately named the payoff: a whole class of shim files — files that existed only to point one format at another — stops being necessary. For teams running several coding agents against one repository, it removes the chore of keeping two rule sets in sync. Our own catalogue held 10,453 Claude Code skills on 22 September 2026; how skills and the SKILL.md file are structured is covered in our guide to Claude Code skills.

Harness Tax: read, write, edit and bash reach the Pareto frontier

A measurement suggests four plain tools give a coding agent close to peak benchmark quality at lower spend. A harness is the scaffolding around an agent: its tool set, context layout, turn budget and stopping rules. The Harness Tax analysis argues that the minimal set — read, write, edit, bash — reaches the Pareto frontier on benchmarks, while anything beyond it tends to add cost more reliably than capability. Alongside it came the paper An Empirical Study of Harness Design for Coding Agents, which finds benchmark outcomes increasingly shaped by harness structure rather than the base model. The practical consequence is already visible in production stacks: frontier for planning, cheap for execution — GPT 5.6 Sol XHigh to plan, GLM 5.3 Flash to implement, DeepSeek V4.1 Flash for the rest.

Robotics: ABC-130K ships 3,500 hours of teleoperation from an $8K rig

Open manipulation data grew to 130,000 episodes collected on hardware costing eight thousand dollars. ABC-130K is described as the largest open teleoperation dataset to date: 3,500 hours, more than 130,000 episodes, 195 tasks, a bimanual setup costing $8K, with open hardware, training code, simulator and evaluation. The baseline science reports a sim-to-real correlation of r = 0.91 on task progress, meaning a policy can be screened in simulation instead of on the robot. The release answers the standard objection to robotics datasets — that the rig itself is the barrier to entry. For scale: the full ABC release includes more than 400 hours of simulation data across 24 tasks and 5,850 labelled policy-evaluation episodes.

Numbers and facts

Different views: what is an autonomous agent demo actually worth?

The argument centres on Vals AI's claim that GPT-6 Astra completed Factorio: Space Age in about two wall-clock days and 165+ in-game hours. The dispute is not about the game but about whether such runs count as evidence of capability.

For. A long-horizon game with production chains resembles real work more closely than short benchmarks: the agent has to hold a plan across hundreds of steps and recover from its own mistakes.

Neutral. Individual integrations are measurable without games at all: browser use was called the strongest application for decision models so far, and a plugin for Cline hands such a model a browser inside the editor. That can be checked against your own tasks in an evening.

Against. The comments asked plain questions: 165 in-game hours in two days implies accelerated time or excluded pauses, and no replay was published for auditing. The same scepticism follows neighbouring demos — a virtual fusion reactor lab built in four hours drew criticism for shipping no equations or methodology: polished, but nothing to verify.

Tools and techniques

In brief