Stripe Buys OpenRouter, 10T Tokens a Day — AI Digest

Stripe has bought OpenRouter, the model router for 10M developers handling 10T tokens a day. UkisAI Swift reasons 63.4% less at the same accuracy, and in biotech AI made analysis cheap but not experiments.
Top stories
Stripe buys OpenRouter: 10 trillion tokens a day and a bet on stopping agentic fraud
Stripe has acquired OpenRouter, the neutral model-routing layer used by more than 10 million developers and handling more than 10 trillion tokens a day. OpenRouter co-founder Alex Atallah and AMP's Anjney Midha walked through the deal in a Latent Space episode on 25 September 2026. OpenRouter bet from day one that no single model would win everything, and grew from what venture investors dismissed as "just a wrapper" into the distribution layer that labs rely on to reach developers. The strategic logic of the deal is fraud. Token fraud is abuse of paid model access: stolen keys, hijacked accounts and resold tokens. Midha argues that inference gateways are increasingly becoming targets, and that the next wave of attacks will come not only from people but from autonomous agents going after valuable token flows. Stripe's fraud infrastructure is the answer to exactly that.
UkisAI Swift: 63.4% fewer reasoning tokens at the same accuracy
UkisAI has released Swift, a family of Qwen-based models trained not to overthink: Swift Flash Next uses 63.4% fewer thinking tokens and runs 1.8x faster, with only a −0.2% score change at xhigh. Overthinking is when a model keeps reasoning after it already has the answer, paying for it in time and tokens. According to the release post on r/LocalLLaMA, those tokens were penalised in training and accuracy was recovered with GSPO and on-policy distillation. Swift1.5 27B cuts reasoning by 58.5% while scoring +0.35% higher, averaged over 5 seeds on GPQA, AIME26, LiveCodeBench, ERQA and Terminal Bench 2.1. The models ship as GGUF, NVFP4, MLX and W4A16, with a 9B variant planned. Separately, Xiaomi revealed the core of the upcoming MiMo-V3 — HySparse2, a sparse-attention design that shrinks the KV cache and prefill cost. How to run open models like these yourself is covered in our open-source LLM guide.
Foundries vs navigators: AI made thinking cheap in science, not experiments
Adrian Sanborn of Endura Therapeutics argues in a Latent Space guest post on 24 September 2026 that AI's main gap in science is simple: thinking got cheap, doing did not. Knowledge work around an experiment has sped up dramatically, yet every hypothesis is still checked by a physical experiment that takes days or weeks. Biotech, in his view, has adapted in two ways. Foundries — Xaira, NewLimit, Insitro, Lila, Periodic Labs and others — cut the cost of doing, through sequencing, multiplexing, high-throughput microscopy and physical automation. Navigators spend the surplus of thinking on decisions: which tools to build and which questions deserve an experiment. Sanborn's navigation numbers: analysis takes an hour instead of a week, software once licensed for six figures becomes a one-day build, and disease programmes are chosen from 500 candidates instead of five. Navigation runs fastest in early-stage startups, which have no legacy processes or contracts to shed.
Ben Tossell built an 87-device site in a morning: 40 messages, 7 subagents
Ben's Bites author Ben Tossell built an interactive site of 87 devices from 1976 to 2026 in a single morning — started in Codex, finished in Factory, with every image generated by GPT-6 Astra. In his write-up on 25 September 2026 the whole build took 40 messages and 7 subagents, and the site has already collected 1,455 votes. A subagent is a separate agent that works in its own session and reports back only the result; Tossell uses them when there is a batch of repetitive work. The habit he leans on most is asking the agent for three versions at once, because "I often don't know what I want until I can see it". In the previous Ben's Bites issue, co-author Keshav put a price on this kind of work with GPT-6 Sol: a four-hour task building a marketing course used 22% of the weekly limit on a $100 Codex plan. Ready-made skills for long, subagent-heavy tasks are in our Claude Code skills guide.
Numbers and facts
- 82.5% win rate — GPT-6 Astra in DOOM agent matches; Sol is the fastest there, and Luna gives the most wins per dollar.
- 3rd try — the attempt on which GPT-6 Astra reportedly beat NetHack, per Ethan Mollick.
- Up to 6.3× — speed-up for Qwen-Image-2.1 with Pruna few-step LoRAs at 5–8 steps.
- About 870 ms — response latency of Muse Realtime Avatar, Meta's model that animates the Muse assistant in sync with its voice.
- $17 — what it takes, per @DataChaz, to train your own Jev-style model in minutes.
- 0.8 points — the gap between the best open model, AutoJev-27B, and Jev on Decision Index v0.2.
AI SKILLS summary table: fast decision models, 24 September 2026
| Model | Where it runs | Latency |
|---|---|---|
| Tev1 0.8B | locally in Ollama | about 50 ms end to end |
| GLiNER2.5-Decide (Fastino) | GPU | 38–47 ms |
| GLiNER2.5-Decide (Fastino) | CPU | 167 ms |
| Jev (TypeSafe) as a judge | API | 152 ms median |
Compiled by our editors from the primary sources: Tev1 0.8B, GLiNER2.5-Decide and Jev-as-a-Judge. Decision models return a finished decision — a label or a structure — rather than prose, so latency matters more than reasoning quality. The authors measured on different hardware, so the rows show orders of magnitude, not a head-to-head.
Different views: is the AI safety alarm overblown?
Nvidia CEO Jensen Huang told Ezra Klein in an interview released on 25 September 2026 that predictions of imminent AI doom are overblown — the same week researchers showed that agents often ignore a monitor telling them to stop. The argument is not about whether risks exist, but about who closes them and how.
Against the alarm. Huang thinks AI alarmism has gone too far and that predictions of imminent doom are exaggerated.
Neutral. Hugging Face's Clem Delangue argues that open models counter the asymmetry between attack and defence capabilities. Meanwhile Anthropic has resumed billing for requests blocked by its safeguards, citing a false-positive rate below 0.1%.
For caution. Research highlighted by @maksym_andr shows agents often keep working when a monitor tells them to stop. A monitor is a separate system that watches an agent's actions and can halt it; if the agent ignores it, oversight stops working.
Tools and techniques
- Try Swift instead of stock Qwen on a home server. One r/LocalLLaMA user has run the 27B version for weeks as a homelab and sysadmin assistant; the quantised file is Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. On smaller machines, wait for the 9B version. More open models for local use are in the AI SKILLS open-source catalogue.
- Ask your agent for three versions at once. Tossell's habit works for any task where the result is easier to judge by eye than to describe: an interface, a piece of copy, a page layout. Hand repetitive work — finding exact release dates, generating images — to subagents.
- Turn on self-repair in Zapier. Zaps, per Ben's Bites, can now fix themselves when they break (beta) and get cheaper the more they run. Ready-made n8n and Make workflows are in the automation templates catalogue.
In brief
- Amazon is investing $100M in robot manufacturing, Superhuman AI reported on 26 September 2026.
- Meta's Muse Spark 1.3 is available on Google Cloud and Oracle, and Spark 1.4 has appeared on OpenCode.
- Perplexity's local agents now run on AMD Ryzen AI Max under the Portable Computer project.
- Google Research announced a multi-agent framework for long-form video that stays consistent over time.
- LangChain shipped Engine v2 with red-teaming and validated fixes.
- Latent Space is merging its AINews newsletter with its Discord: the newsletter has more than 200K subscribers, and Supabase Select takes place in San Francisco on 2 October 2026.
- GPT-6 Astra tops MentalHealthBench, a benchmark for mental-health conversations, per Ben's Bites.