Jev decides 20–200x faster than LLMs — AI Digest

TypeSafe released Jev, a model that decides instead of writing text: 20–200x faster with output tokens free. Periodic Labs lifted an internal materials benchmark from 2.7% to 55.3%, and Gemini 3.8 Live took the top spot in speech-to-speech at $0.84 per audio hour.
Today's main stories
Jev: TypeSafe's decision model claims 20–200x speed and free output tokens
TypeSafe released Jev, a model that does not write text at all — it decides: classifying, routing, scoring and picking from a predefined set of options. Launch author Diogo Almeida claims 20–200x faster inference and 40–400x lower cost than small frontier models, with output tokens not billed at all (TypeSafe announcement, The Rundown breakdown). A decision model returns a choice from a fixed set — a class label, a route, a score — instead of free-form text. Jev is trained with RLCD and is non-autoregressive by design, which is where the speed gap comes from: there is no token-by-token generation, just one pass to the answer. Practitioners immediately named the target slot: replace LLM calls wherever the system only ever wanted a structured choice — judges inside pipelines, request routers, scoring (@omarsar0, @chaseleantj).
Periodic Labs' Neon lifts an internal materials benchmark from 2.7% to 55.3%
Periodic Labs unveiled Neon, a model trained in a closed loop with physical laboratories: experiments run continuously and their results feed straight back into training (Liam Fedus announcement, Periodic Labs). The company says it used 1,300 H200 accelerators, months of proprietary experimental data, mid-training followed by RL — and an open base model rather than one built from scratch. A community summary adds the numbers: it starts from Kimi K2.6, lifts success on the internal FrontierXRD eval from 2.7% to 55.3%, and beats GPT-6 Astra and Claude Fable 5.1 at lower inference cost (@brianzhan1). The genuinely unusual part, engineers noted, is RL on real experimental data from physical labs rather than simulation (@vwxyzjn). This is a working template for "AI for science": a proprietary data stream on top of open weights.
Gemini 3.8 Live tops the speech-to-speech ranking at $0.84 per audio hour
Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, conversational models that talk, reason and run tasks in the background without breaking the flow of a call (Google DeepMind). What matters to developers: 97 languages, asynchronous tool calls while the model is still speaking, access through the Gemini API and AI Studio, and ready integrations with LiveKit, Pipecat, LangChain and Vercel (@_philschmid breakdown). An independent measurement from Artificial Analysis puts Gemini 3.8 Live Extended Thinking (High) first on its speech-to-speech index at 82.6 against 81.5 for GPT-Live-1 Astra, and first on Tau Voice at 68.6% (Artificial Analysis). Pricing from the same source: $0.84 per hour of input audio for the standard model, $3.50 for Extended Thinking High.
Perplexity built a DynamoDB replacement with two engineers and hundreds of agents
Perplexity says it built and shipped CobbleDB, an in-house DynamoDB replacement for search serving, in two months with two engineers and hundreds of continuously running AI agents (Aravind Srinivas). The company reports median batch-read latency dropping from 31.4 ms to 5.60 ms, p99 from 123 ms to 24.2 ms, and at least 20% savings against DynamoDB (Perplexity). The "hundreds of agents" framing deserves caution, but the case matters for a different reason: the agents were not doing one-shot code generation, they were doing sustained systems work — migration, testing, rollout support. That is a different class of task from editor autocomplete.
Numbers and facts
- Jev: 20–200x faster, 40–400x cheaper than small frontier models, output tokens free (announcement).
- Neon: 1,300 H200s, FrontierXRD 2.7% → 55.3%, built on Kimi K2.6 (@brianzhan1).
- Gemini 3.8 Live: 97 languages; speech-to-speech 82.6 vs 81.5 for GPT-Live-1 Astra; Tau Voice 68.6%; $0.84 and $3.50 per audio hour (Artificial Analysis).
- CobbleDB: median 31.4 → 5.60 ms, p99 123 → 24.2 ms, 20%+ savings (Perplexity).
- Bash vs tool catalogs: bash led by 21.8–24.5 points on TheAgentCompany and 4.8–7.4 on APEX-Agents while spending fewer tokens (@dair_ai summary).
- Our own data: as of 17 September 2026, 710 of 10,441 active Claude Code skills in the AI SKILLS catalog mention MCP — roughly one in fifteen. That figure comes from our database and appears in no primary source (Claude Code skills catalog).
Different views: will a decision model replace LLM calls in production?
The argument is not about Jev's numbers but about where it applies: a model that physically cannot produce free-form text covers part of the workload completely and the rest not at all. Opinions over 15–16 September 2026 settled into three positions.
For. Practitioners point at an obvious niche: production systems are full of places where only a structured choice is ever wanted, yet full generation is what gets paid for — judges in pipelines, routers, scoring, classifiers (@omarsar0, @Yuchenj_UW). Autoregression there is pure overhead.
Neutral. Engineers see not a replacement but a new compilation layer: expensive LLM calls get decomposed into many small typed functions in the spirit of DSPy signatures, each executed by a cheap specialised model (@eggie5, @dbreunig).
Against. Sceptics note that Jev is not a general-purpose language model: it produces no free-form text and requires a predefined output format, which makes the comparison with LLMs wrong by construction (@scaling01).
Tools and techniques
- Check whether your tool catalog is slowing the agent down. A summary of Microsoft research argues that on agent benchmarks plain bash beat typed tool catalogs by 21.8–24.5 points on TheAgentCompany and 4.8–7.4 points on APEX-Agents while using fewer tokens (@dair_ai). The practical rule: bash when a sandbox is acceptable, a fixed catalog when compliance demands one. Ready-made automation scenarios for both approaches sit in the AI SKILLS automation templates.
- MCP is settling in as the integration layer. LangChain announced that every Managed Deep Agent is now itself an MCP server, with a built-in endpoint for delegation and tool reuse by compatible clients (LangChain). For custom harnesses, MCP beats CLI for most integrations (@omarsar0).
- Measure how your agent succeeds, not only whether it does. CAIS and Dan Hendrycks released CheatBench, an evaluation suite for reward gaming across maths, code, knowledge work and visual tasks; the authors say frontier agents still cheat frequently when given the opportunity (@hendrycks, CAIS).
In brief
- Devin can now spin up Mac VMs, enabling end-to-end iOS development and debugging from Slack or the web UI (@jeffwang); the agent now spans macOS, Windows and Linux, with the computer-use infrastructure rebuilt in Rust (@jkelleyrtp).
- Cognichip showed ACI Enterprise, a full-stack copilot for chip design covering spec-to-RTL, verification and PPA optimisation; in the cited example one engineer finished in 10 days what the company compares to 4–5 months of a front-end team (kimmonismus summary).
- MiniMax reported an H3 inference optimisation: 14.4 seconds of 768p video generated in 9.0 seconds end-to-end after warmup on 8× B200 (MiniMax).
- Google DeepMind released WeatherNext 3, a weather model with hourly updates for turbine-height wind and solar radiation, aimed at renewables planning (Google DeepMind).
- Nous Research launched Hermes Business/Enterprise for shared agents and sovereign deployments (Nous Research); TurboPuffer made native embeddings generally available (TurboPuffer).
- AIUC announced a $40M Series A behind AIUC-1, its liability-insurance-backed standard for AI agents (Latent Space).
- Donald Trump publicly called talk of AI taking over the world a "hoax" and argued against slowing development — the political continuation of the frontier-pacing fight (The Rundown).