Opus 5.5 Builds a Video From Code for $4 — AI Digest

Claude Opus 5.5 is turning into a video tool: users render full clips entirely from code for as little as $4. TypeSafe is raising $1B+ at a $10B+ valuation for its Jev judge model, and Gemini 3.8 Flash hits 89.2% on ARC-AGI-2.
Top stories
Claude Opus 5.5 builds entire videos out of code: a $4 clip and an 8-hour island
Within two days of release, Claude Opus 5.5 has turned into a video-making tool: users are assembling clips whose script, graphics, animation and voiceover are all written in code, with not a single frame coming from a video model. The author of a post on r/ClaudeAI produced a 30-60 second clip in 1 hour 20 minutes: about $20 on Opus and $3.21 on external APIs via OpenRouter. Another user replicated the same prompt for their own project and came in at roughly $4 over 1.5-2 hours. Dan Greenheck spent 8 hours building the interactive TideWater island, complete with birds, fish and a boat — burning through $1,874 in tokens, 59% of his weekly Max limit. The day's most popular post was a Claude-made video about Western civilization (30.6K engagements). These examples have already sparked debate over whether diffusion is even needed for this kind of video.
TypeSafe is raising $1B at a $10B valuation: Jev judge is 277x cheaper than GPT-6
Startup TypeSafe is raising more than $1B at a valuation above $10B, according to journalist Stephanie Palazzolo — just a week after a $200M round. Jev is a "System 1" model: instead of writing out reasoning, it returns a typed decision with a probability attached. In the Jev-as-a-Judge paper, a Jev-based judge costs $0.044 per 1,000 evaluations at a median latency of 152ms — roughly 277 times cheaper than GPT-6. On RewardBench and HaluEval, Jev trails GPT-6 by no more than 3 points; on JudgeBench, by 14.5. A cascade that hands off uncertain cases to GPT-6 Astra keeps 99% accuracy at 57% of the cost. Ramp reports it matched GPT-5.6 Luna's reranking accuracy with a tail latency of 300ms — 10 times lower — and at a third of the cost. The AI Daily Brief broke down six categories of real-world Jev use cases, from ad-campaign analysis to inbox triage.
Gemini 3.8 Flash: 89.2% on ARC-AGI-2 at $0.40 per task
Google's Gemini 3.8 Flash scored 89.2% on ARC-AGI-2 at a cost of $0.40 per task, and 98.5% on the first version of the test. The picture is more modest on ARC-AGI-3: 10.4% with a standard harness and 35% with Google's own harness. A harness is the scaffolding wrapped around a model — tools, prompts and the agent's execution loop; the 3.4x difference shows how much of the result on novel tasks comes from the harness itself. On the Artificial Analysis index, Gemini 3.8 Flash scored 41 points at a speed of 291 tokens per second with a 1 million token context, and in Cline the model is currently available for free.
Perplexity Photon: p99 latency drops from 800ms to 65ms, engine written by hundreds of agents
Perplexity has released Photon, a Rust-based search and ranking engine that a small team and hundreds of agents built for roughly $300K worth of tokens. According to the company, internal p99 latency dropped from ~800ms to ~65ms on 20% fewer machines, while the amount of data per document grew 2.5x. The Fast Search API runs on top of the engine: 160ms median, 230ms at the 95th percentile, and 68% lower cost per task. Nous Research made it free in Hermes Agent, and at Shopify, according to Mikhail Parakhin, the API has become the primary search backend. Perplexity's report is unusual in naming both the cost of the agents' work and the measured production gains.
Numbers and facts
- 88.4% — Claude Opus 5.5 topped SimpleBench; on vision, @skalskip92 ranks it above Fable 5 and GPT-6 Sol but below GPT-6 Astra, at roughly 60% cheaper than Fable 5.1.
- 62% on xhigh vs. 59% on max — Opus 5.5 on Terminal-Bench-Science; the best model outside OpenAI and Anthropic, Qwen3.8 Max, scored 12%.
- 46 points and $0.13 per task vs. $1.99 — Xiaomi MiMo-V2.6-Pro on the AA index, one point below GPT-5.6 Sol; Xiaomi released the RL code and training environments.
- From 23.3% to 44.3% — Harness-Zero bakes the harness into the model itself: without any scaffolding, it solves more than the base model does with one (41.7%).
- Down to 46.7% — that's how much DeepMind's XYEval test drags down agent performance with a single confidently false user hint.
- 3.13x — the decoding speedup for LFM2.5-VL-3B using Liquid AI's DSpark draft model on an M5 Max; speculative decoding is when a small model proposes tokens and a large one just verifies them in a batch.
- 469 tokens per second — GLM-5.3 on 8x AMD MI355X in vLLM and TileRT, for a single user.
AI SKILLS summary table: cost per solved task, September 24, 2026
| Model | What was measured | Price |
|---|---|---|
| Jev (TypeSafe) | acting as judge, 1000 evaluations | $0.044 |
| Xiaomi MiMo-V2.6-Pro | AA index task | $0.13 |
| Gemini 3.8 Flash | ARC-AGI-2 task | $0.40 |
| GPT-6 Luna [Max] | 1M tokens, blended price | ~$0.40 |
| Grok 4.7 | Agent Arena task | $1.14 |
The table was compiled by the editors from the primary sources above; the units differ, so the rows can't be compared directly — they illustrate the order of magnitude of prices. Luna Max entered Code Arena WebDev in 24th place (1,593 points), and Grok 4.7 took 16th place in Agent Arena.
Different views: Is Jev a new class of model or a well-packaged classifier?
Jev really is cheap and fast, but the argument is over how to classify it: on the Banking77 task, a simple combination of BGE-small and logistic regression, per an independent benchmark, scored 93.3% against Jev's 83.2%. The answer determines what Jev should be compared to — LLMs or classical classifiers.
For. Investors are valuing TypeSafe at $10B, and Jev proved 140 Software Foundations theorems for under $1 — roughly 130 times cheaper than Astra. Jev is the top model on OpenRouter for the 1-10K token context range.
Neutral. Proponents of CLM offer an open alternative: CLM is about 9 times faster than Jev as a verifier on long tasks, but falls behind on "broad" zero-shot tasks: 95.2% vs. 99.2% on BFCL v4.
Against. The author of a breakdown on r/LocalLLaMA argues Jev is standard classification over a fixed set of labels, and that the "0% hallucinations" claim only guarantees the format of the answer, not its correctness. In his view, Jev should be compared to embedding models and cross-encoders, not to LLM-based JSON generation.
Tools and techniques
- Don't set Opus 5.5's reasoning effort to max. Reasoning effort is a setting for how many tokens a model spends thinking before it answers. Results on max are lower than on xhigh, because max forces a minimum reasoning budget — Theo advises avoiding it. Alongside the Opus 5.5 release, Claude Code's five-hour limits grew by 20%, Ben's Bites reports; ready-made skills for long-running tasks are covered in the Claude Code skills guide.
- Load GGUF files straight into Transformers. Hugging Face Transformers can now open quantized GGUF files via
from_pretrained(..., gguf_file=...): on an M2 Max, Qwen3.8-27B in UD-Q4_K_M delivers 15.9 tokens per second versus 13.4 for llama.cpp. For how to pick and run open-source models locally, see the open-source LLM guide. - Set
balanceexplicitly in Weaviate 1.39. The new version makes MMR result diversity stable, but the default value of 0.0 means pure diversity with no regard for relevance. RAG and agent-scenario templates are available in the AI SKILLS automation templates catalog.
In brief
- Google is sending four TPUs into orbit aboard a Planet prototype satellite on SpaceX's Transporter-18 mission — the Suncatcher project; launch is set for next week.
- At its Interrupt conference, LangChain released Managed Deep Agents 0.8 with user and agent memory, plus LangSmith Fine-Tuning, which turns traces into fine-tuning datasets.
- Quail is an open-source AI-SQL engine that processes more than 1 billion input tokens per minute on a single H100.
- Odyssey Agora-2 simulates up to 20 people and agents in a single world in real time.
- Meta gave agents dedicated memory agents to fight "context rot": Sonnet 4.5 improved from 37.6% to 45.9%.
- SmolDataEnvs offers more than 5,000 verifiable RL environments for data analysis, for models up to 10B on a single GPU.
- A NeurIPS paper shows that suppressing a single MLP neuron removes safety refusals in 7 models ranging from 1.7B to 70B parameters.
- Sakana AI appointed Jurgen Schmidhuber chief scientific advisor for its self-improving systems lab; C5R built an AI-run lab and the SciUniverse benchmark in 12 weeks.
- OpenRouter, per a Latent Space episode, has become a neutral model-routing layer for more than 10 million developers.