Opus 5.5 Builds a Video From Code for $4 — AI Digest

Opus 5.5 Builds a Video From Code for $4 — AI Digest

Claude Opus 5.5 is turning into a video tool: users render full clips entirely from code for as little as $4. TypeSafe is raising $1B+ at a $10B+ valuation for its Jev judge model, and Gemini 3.8 Flash hits 89.2% on ARC-AGI-2.

Top stories

Claude Opus 5.5 builds entire videos out of code: a $4 clip and an 8-hour island

Within two days of release, Claude Opus 5.5 has turned into a video-making tool: users are assembling clips whose script, graphics, animation and voiceover are all written in code, with not a single frame coming from a video model. The author of a post on r/ClaudeAI produced a 30-60 second clip in 1 hour 20 minutes: about $20 on Opus and $3.21 on external APIs via OpenRouter. Another user replicated the same prompt for their own project and came in at roughly $4 over 1.5-2 hours. Dan Greenheck spent 8 hours building the interactive TideWater island, complete with birds, fish and a boat — burning through $1,874 in tokens, 59% of his weekly Max limit. The day's most popular post was a Claude-made video about Western civilization (30.6K engagements). These examples have already sparked debate over whether diffusion is even needed for this kind of video.

TypeSafe is raising $1B at a $10B valuation: Jev judge is 277x cheaper than GPT-6

Startup TypeSafe is raising more than $1B at a valuation above $10B, according to journalist Stephanie Palazzolo — just a week after a $200M round. Jev is a "System 1" model: instead of writing out reasoning, it returns a typed decision with a probability attached. In the Jev-as-a-Judge paper, a Jev-based judge costs $0.044 per 1,000 evaluations at a median latency of 152ms — roughly 277 times cheaper than GPT-6. On RewardBench and HaluEval, Jev trails GPT-6 by no more than 3 points; on JudgeBench, by 14.5. A cascade that hands off uncertain cases to GPT-6 Astra keeps 99% accuracy at 57% of the cost. Ramp reports it matched GPT-5.6 Luna's reranking accuracy with a tail latency of 300ms — 10 times lower — and at a third of the cost. The AI Daily Brief broke down six categories of real-world Jev use cases, from ad-campaign analysis to inbox triage.

Gemini 3.8 Flash: 89.2% on ARC-AGI-2 at $0.40 per task

Google's Gemini 3.8 Flash scored 89.2% on ARC-AGI-2 at a cost of $0.40 per task, and 98.5% on the first version of the test. The picture is more modest on ARC-AGI-3: 10.4% with a standard harness and 35% with Google's own harness. A harness is the scaffolding wrapped around a model — tools, prompts and the agent's execution loop; the 3.4x difference shows how much of the result on novel tasks comes from the harness itself. On the Artificial Analysis index, Gemini 3.8 Flash scored 41 points at a speed of 291 tokens per second with a 1 million token context, and in Cline the model is currently available for free.

Perplexity Photon: p99 latency drops from 800ms to 65ms, engine written by hundreds of agents

Perplexity has released Photon, a Rust-based search and ranking engine that a small team and hundreds of agents built for roughly $300K worth of tokens. According to the company, internal p99 latency dropped from ~800ms to ~65ms on 20% fewer machines, while the amount of data per document grew 2.5x. The Fast Search API runs on top of the engine: 160ms median, 230ms at the 95th percentile, and 68% lower cost per task. Nous Research made it free in Hermes Agent, and at Shopify, according to Mikhail Parakhin, the API has become the primary search backend. Perplexity's report is unusual in naming both the cost of the agents' work and the measured production gains.

Numbers and facts

AI SKILLS summary table: cost per solved task, September 24, 2026

ModelWhat was measuredPrice
Jev (TypeSafe)acting as judge, 1000 evaluations$0.044
Xiaomi MiMo-V2.6-ProAA index task$0.13
Gemini 3.8 FlashARC-AGI-2 task$0.40
GPT-6 Luna [Max]1M tokens, blended price~$0.40
Grok 4.7Agent Arena task$1.14

The table was compiled by the editors from the primary sources above; the units differ, so the rows can't be compared directly — they illustrate the order of magnitude of prices. Luna Max entered Code Arena WebDev in 24th place (1,593 points), and Grok 4.7 took 16th place in Agent Arena.

Different views: Is Jev a new class of model or a well-packaged classifier?

Jev really is cheap and fast, but the argument is over how to classify it: on the Banking77 task, a simple combination of BGE-small and logistic regression, per an independent benchmark, scored 93.3% against Jev's 83.2%. The answer determines what Jev should be compared to — LLMs or classical classifiers.

For. Investors are valuing TypeSafe at $10B, and Jev proved 140 Software Foundations theorems for under $1 — roughly 130 times cheaper than Astra. Jev is the top model on OpenRouter for the 1-10K token context range.

Neutral. Proponents of CLM offer an open alternative: CLM is about 9 times faster than Jev as a verifier on long tasks, but falls behind on "broad" zero-shot tasks: 95.2% vs. 99.2% on BFCL v4.

Against. The author of a breakdown on r/LocalLLaMA argues Jev is standard classification over a fixed set of labels, and that the "0% hallucinations" claim only guarantees the format of the answer, not its correctness. In his view, Jev should be compared to embedding models and cross-encoders, not to LLM-based JSON generation.

Tools and techniques

In brief