GPT-6.1 Sol at $0.56 per Task on Agent Arena — AI Digest

GPT-6.1 Sol at $0.56 per Task on Agent Arena — AI Digest

GPT-6.1 Sol entered Agent Arena at #5 for $0.56 per task, while Anthropic took the top three spots. Hugging Face showed the same weights scoring 62% vs 33% across harnesses.

Today's top stories

GPT-6.1 Sol costs $0.56 per task on Agent Arena, while Anthropic holds the top three spots

OpenAI's GPT-6.1 Sol entered the Agent Arena leaderboard at #5 with a median cost of $0.56 per task, while Anthropic models took the top three places. According to Arena's data, Sol in Max mode is 39% cheaper than GPT-6 Sol while scoring 1.52 points higher, and 81% cheaper than GPT-6 Astra while trailing it by just 1.04 points. Claude Sonnet 5.5 in Max mode debuted at #3 and ranked first in the Chat category, but it costs $2.74 per task versus $1.58 for Opus 5.5 at #2 — which, in Arena's view, keeps Sonnet 5.5 off the cost-performance frontier. GPT-6.1 Sol is priced at $2/$10 per million input and output tokens, against $10/$50 for Astra, DL Weekly reports. In WebDev Arena, Sol briefly reached #3 before Sonnet 5.5 pushed it to #4, and Gemini 4 Argon in High mode took #1 in Text Arena, according to Arena's weekly recap.

Hugging Face: the same weights score 62% in one harness and 33% in another

Hugging Face showed that a model's score depends heavily on its harness: identical weights reached 62% in one setup and 33% in another. A harness is the software wrapped around a model to make it an agent: the call loop, the tool set, the prompt format and the rules for managing context. Hugging Face released a method for reinforcement learning across several harnesses at once: a proxy speaks the OpenAI, Anthropic and Gemini API formats and records sampled token IDs and logprobs for training, with no changes to the harnesses themselves. After this training, LFM2.5-2.6B rose from 42% to 54% on average across four harnesses and made 31% fewer tool calls. For comparison, plain supervised fine-tuning on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%. Hugging Face open-sourced the trainer, the data and all seven trained models; for the bigger picture on open weights, see our open-source LLM guide.

Pi 1.0 ships MCP by default as DeepSeek, Anthropic and T3 Code roll out their own harnesses

Earendil released Pi 1.0, the stable version of its minimal agent harness, with native MCP support built into Codemode. Per Earendil's announcement, the release also adds deferred tool loading, cache warming for Anthropic models and mid-conversation system messages, while the experimental Pi Durable runs on Cloudflare Durable Objects via agents SDK v0.26.0. DeepSeek Harness gained desktop builds for macOS and Windows. Claude Code introduced mods — plugins with middleware-style hooks; our Claude Code skills guide explains how Claude Code extensions fit together. The T3 Code orchestrator passed 400,000 users, and its rewrite — 823 commits across 1,912 files — has now merged. Elvis Saravia frames this wave as a shift toward malleable harnesses. Ready-made agent workflows for n8n and Make live in the AI SKILLS automation templates catalog.

Decision models land in llama.cpp, and Cloudflare releases the open Clef

llama.cpp added a /v1/systemone endpoint for running decision models locally. A decision model is a model that, instead of writing free text, picks exactly one option from a given list: an action, a class or a continuation. Georgi Gerganov announced the endpoint, and Clément Delangue showed a one-command launch with llama serve -hf ggml-org/Kev-4B-GGUF. Cloudflare released Clef, an open-weights decision model post-trained from Qwen3.8-27B, plus a lighter clef-flash built on Qwen3.5-9B; Clef models are already on Ollama. Perplexity claims pplx-decider-v1-27b averages 85.7% across 11 benchmarks. Skeptics question the novelty: Merve Noyan calls decision models a rebrand of zero-shot classifiers.

Meta publishes six math papers produced with Muse Spark in a plain chat

Meta published six papers on open math problems produced with Muse Spark 1.1 and 1.2 through the ordinary meta.ai chat, with no custom scaffold. AI at Meta announced the release, and Alexandr Wang shared the list. Each paper labels which passages were drafted primarily by humans and which by AI, and a second group of mathematicians reviewed the work. In parallel, Google presented Cogentic, a Gemini-based multi-agent system that produced new results on five open theory problems: every draft must pass two adversarial verifiers, and the agents share a ledger of verified lemmas. Most problems took Cogentic about 100 model calls; the hardest took about 1,000.

Numbers and facts

AI SKILLS summary table: cost per task on Agent Arena as of October 2, 2026. Per-task costs for GPT-6 Sol and GPT-6 Astra are our own calculation from the percentages Arena published; all other figures come from Arena and DL Weekly.

ModelAgent Arena rankCost per taskPrice per 1M tokens (in/out)
Claude Opus 5.52$1.58—
Claude Sonnet 5.5 [Max]3$2.74—
GPT-6.1 Sol [Max]5$0.56$2/$10
GPT-6 Sol—≈$0.92 (calculated)—
GPT-6 Astra—≈$2.95 (calculated)$10/$50

Different perspectives: has Claude Opus 5.5 gotten worse since launch?

In early October 2026, Reddit users complained en masse about a quality drop in Claude Opus 5.5, and the sentiment tracker modelsentiment.com recorded the model's score falling from 71–73 to 55 out of 100. Yet none of the complainants offered a reproducible measurement, and independent leaderboards this same week keep Opus 5.5 at #2 on Agent Arena.

Tools and techniques

In brief