GPT-6.1 Sol at $0.56 per Task on Agent Arena — AI Digest

GPT-6.1 Sol entered Agent Arena at #5 for $0.56 per task, while Anthropic took the top three spots. Hugging Face showed the same weights scoring 62% vs 33% across harnesses.
Today's top stories
GPT-6.1 Sol costs $0.56 per task on Agent Arena, while Anthropic holds the top three spots
OpenAI's GPT-6.1 Sol entered the Agent Arena leaderboard at #5 with a median cost of $0.56 per task, while Anthropic models took the top three places. According to Arena's data, Sol in Max mode is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher, and 81% cheaper than GPT-6 Astra while trailing it by just 1.04 points. Claude Sonnet 5.5 in Max mode debuted at #3 and ranked first in the Chat category, but it costs $2.74 per task versus $1.58 for Opus 5.5 at #2 — which, in Arena's view, keeps Sonnet 5.5 off the cost-performance frontier. GPT-6.1 Sol is priced at $2/$10 per million input and output tokens, against $10/$50 for Astra, DL Weekly reports. In WebDev Arena, Sol briefly reached #3 before Sonnet 5.5 pushed it to #4, and Gemini 4 Argon in High mode took #1 in Text Arena, according to Arena's weekly recap.
Hugging Face: the same weights score 62% in one harness and 33% in another
Hugging Face showed that a model's score depends heavily on its harness: identical weights reached 62% in one setup and 33% in another. A harness is the software wrapped around a model to make it an agent: the call loop, the tool set, the prompt format and the rules for managing context. Hugging Face released a method for reinforcement learning across several harnesses at once: a proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves. After this training, LFM2.5-2.6B rose from 42% to 54% on average across four harnesses and made 31% fewer tool calls. For comparison, plain supervised fine-tuning on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%. Hugging Face open-sourced the trainer, the data and all seven trained models; for the bigger picture on open weights, see our open-source LLM guide.
Pi 1.0 ships MCP by default as DeepSeek, Anthropic and T3 Code roll out their own harnesses
Earendil released Pi 1.0, the stable version of its minimal agent harness, with native MCP support built into Codemode. Per Earendil's announcement, the release also adds deferred tool loading, cache warming for Anthropic models and mid-conversation system messages, while the experimental Pi Durable runs on Cloudflare Durable Objects via agents SDK v0.26.0. DeepSeek Harness gained desktop builds for macOS and Windows. Claude Code introduced mods — plugins with middleware-style hooks; our Claude Code skills guide explains how Claude Code extensions fit together. The T3 Code orchestrator passed 400,000 users, and its rewrite — 823 commits across 1,912 files — has now merged. Elvis Saravia frames this wave as a shift toward malleable harnesses. Ready-made agent workflows for n8n and Make live in the AI SKILLS automation templates catalog.
Decision models land in llama.cpp, and Cloudflare releases the open Clef
llama.cpp added a /v1/systemone endpoint for running decision models locally. A decision model is a model that, instead of writing free text, picks exactly one option from a given list: an action, a class or a continuation. Georgi Gerganov announced the endpoint, and Clément Delangue showed a one-command launch with llama serve -hf ggml-org/Kev-4B-GGUF. Cloudflare released Clef, an open-weights decision model post-trained from Qwen3.8-27B, plus a lighter clef-flash built on Qwen3.5-9B; Clef models are already on Ollama. Perplexity claims pplx-decider-v1-27b averages 85.7% across 11 benchmarks. Skeptics question the novelty: Merve Noyan calls decision models a rebrand of zero-shot classifiers.
Meta publishes six math papers produced with Muse Spark in a plain chat
Meta published six papers on open math problems produced with Muse Spark 1.1 and 1.2 through the ordinary meta.ai chat, with no custom scaffold. AI at Meta announced the release, and Alexandr Wang shared the list. Each paper labels which passages were drafted primarily by humans and which by AI, and a second group of mathematicians reviewed the work. In parallel, Google presented Cogentic, a Gemini-based multi-agent system that produced new results on five open theory problems: every draft must pass two adversarial verifiers, and the agents share a ledger of verified lemmas. Most problems took Cogentic about 100 model calls; the hardest took about 1,000.
Numbers and facts
AI SKILLS summary table: cost per task on Agent Arena as of October 2, 2026. Per-task costs for GPT-6 Sol and GPT-6 Astra are our own calculation from the percentages Arena published; all other figures come from Arena and DL Weekly.
| Model | Agent Arena rank | Cost per task | Price per 1M tokens (in/out) |
|---|---|---|---|
| Claude Opus 5.5 | 2 | $1.58 | — |
| Claude Sonnet 5.5 [Max] | 3 | $2.74 | — |
| GPT-6.1 Sol [Max] | 5 | $0.56 | $2/$10 |
| GPT-6 Sol | — | ≈$0.92 (calculated) | — |
| GPT-6 Astra | — | ≈$2.95 (calculated) | $10/$50 |
- $2.6–5.3 trillion a year would have to go into AI infrastructure even at 20% utilization, while lab revenue reaches roughly $1 trillion by the end of 2027, Epoch AI estimates.
- 62.8% is how much accuracy seven open models lose as context grows from 4K to 128K tokens, per NVIDIA's Long-Transduction study.
- Up to 48% less peak context with Microsoft's training-free FOCUS method, which also raises task success by up to 8.9 points (dair.ai).
- From 63.7% to 71.5% on ProgramBench for GPT-5.5 with a dedicated controller from Meta Superintelligence Labs, using the same workers and budget; Codex scores 58.0% (dair.ai).
- From 61.88% to 73.33% AIME 2024 pass@1 with Apple's LoopCD, which also halves recurrent loops (Aran Komatsuzaki).
- 9% more efficient than MXFP4 — NVFP4 on B200, with higher accuracy, as measured by Stas Bekman.
- From 576 to 352 bytes per MLA latent cache row after Prime Intellect moved it to NVFP4, fitting about 50% more cached tokens than FP8 (Prime Intellect).
- About $5.7 trillion — Nvidia's record market value after a $150 billion buyback increase, per Bloomberg (summary).
Different perspectives: has Claude Opus 5.5 gotten worse since launch?
In early October 2026, Reddit users complained en masse about a quality drop in Claude Opus 5.5, and the sentiment tracker modelsentiment.com recorded the model's score falling from 71–73 to 55 out of 100. Yet none of the complainants offered a reproducible measurement, and independent leaderboards this same week keep Opus 5.5 at #2 on Agent Arena.
- For. The author of the "Opus 5.5 nerfing" thread says that 5–6 days after launch the model handles complex C++, 3D and physics work noticeably worse, and advises saving launch-day prompts and outputs for comparison. A Claude Code Enterprise user writes that usage jumped from 70% to 90% of the limit within an hour, with wordier, duplicated code.
- Neutral. The modelsentiment.com tracker shows Opus 5.5 dropping from 71–73 (September 25–28) to 55, but the commenter citing it notes that it measures user opinion, not model behavior.
- Against. Commenters in both threads offer neither control runs nor a measurement history. Reading 324 thinking summaries, Design Arena found that Opus 5.5 commits early in about 4 of 5 cases, while GPT-6 Astra hedges roughly 20 times as often.
Tools and techniques
- Use an iPhone as a second GPU for a local model. backburner, a llama.cpp fork, offloads part of Qwen 3.8 27B from a MacBook to an iPhone 17 Pro Max over USB-C and, per the author's measurements, speeds up prefill by 35% at 8K context and 44% at 16K. More local-inference tools are in the AI SKILLS Open Source section.
- Turn on MTP decoding for Qwen3.8-Flash Next. llama.cpp PR #29761 adds the
--spec-type draft-mtpflag: on DGX Spark decode speed rose from 28.36 to 43.88 tokens per second, a 1.55× gain. Mind the size: the IQ4_NL quant weighs about 102 GB. - Install the "You should know" plugin in Claude Code. It spins off a side agent that flags important output the user might otherwise miss.
In brief
- Tavus unveiled Griffin, a live video interaction model that 44% of participants mistook for a real person, versus about 3% for other systems; more from The Rundown and Superhuman.
- Google DeepMind introduced SynthID Bio, watermarks for proteins, The Rundown reports.
- StepFun's Step 5 Preview ranks #7 among open-weight models on Vals at $2.54 per task, with a 1M-token context window.
- Qwen3.8-27B-Humanlike-Chat 2.0 lifted IFBench from 37.3 to 43.7 and When2Call from 48 to 58 over its base model, but slipped on MMLU-Pro from 78.5 to 72.5.
- webAI's 3.66B TwIL-LM3-Pro roughly matches Qwen3-8B on formal logic; the Q4 quant is 2.09 GiB and the license is non-commercial.
- Reka released RIDM, an inverse dynamics model under Apache 2.0 that extracts motor and camera actions from real video.
- Andrej Karpathy asked a model "land or water?" for 16,200 coordinate pairs and got a recognizable world map.
- OpenAI's Agents API claims 99.97% turn reliability and 20% faster tool calls.
- According to The Batch, open-weight GLM-5.3 nearly matched Claude Mythos at exploiting vulnerabilities — 12% vs 14%.
- Nathan Lambert and Tom Zick launched Trillium Labs, a non-profit for open post-training recipes, backed by Halcyon Futures and Schmidt Sciences.
- A group run by a former White House deputy chief of staff plans to spend at least $100 million on a campaign framing AI warnings as a coordinated effort.
- Meta made 5,000 Muse Home Link smart-home bridges, free for subscribers while supplies last.