8 of 13 AI Agents Upsell 'Wealthy' Users by $284 — AI Digest

8 of 13 models acting as personal agents pick pricier options for users who seem wealthy, with Opus 4.8 showing gaps up to $284 a month. Devin cuts prices by up to 70%, and OpenAI retires its GPT-3-era models.
Top stories
Personal agents upsell "wealthy" users: 8 of 13 models, with gaps up to $284 a month
Eight of thirteen language models acting as personal agents picked pricier options when the context suggested the user was well-off. The finding comes from a study of 325,000 experiments shared by @omarsar0. A personal agent is a model that compares offers and makes choices on a person's behalf: an insurance plan, a subscription, a purchase. Opus 4.8 showed the widest gap: for a "wealthy" user it chose insurance that cost $284 a month more. For GPT-5.5, the gap grew by 40% when the user's employment details were hidden. The takeaway for anyone handing purchases to an agent: the model builds its own profile of the buyer from indirect signals, and that profile shapes the price you end up paying.
Opus 5.5: Anthropic's official prompting guide says to start at medium effort
Anthropic published an official prompting guide for Claude Opus 5.5 that recommends starting at medium or low effort and leaving enough max_tokens for hidden reasoning. Changing effort at the top level invalidates the prompt cache unless you use the per-message effort beta. For agent builders, the guide says to parse responses by content block type, not to treat a text-only end_turn update as task completion, and to pace work with an elapsed-time budget string such as elapsed 340s / 1200s. Pasted user content should be wrapped in tags with a random ID to reduce prompt injection. Prompt injection is an attack in which text hidden inside data overrides the model's instructions. Instead of a generic "think carefully," the guide favors operational instructions like "look around first." In the Reddit discussion, users report that a time budget stops the model from wrapping up long refactors early. Ready-made templates live in the AI SKILLS prompt library.
Coding agents get cheaper: Devin cuts prices by up to 70%, Cline ships a desktop app
Cognition cut Devin's prices by 30–40% in Fusion and Normal modes and by up to 70% in Review mode, and Devin Mobile is now in beta. The change was announced by Cognition. The same day, Cline shipped a desktop app and Base44 launched Base Code, a model-agnostic cloud development environment. A harness is the scaffolding around a model: the tools, prompts and loop an agent uses to work on code. An Arena study of the "harness tax" found that choosing between Claude Code, Codex and Pi changes cost more than it changes accuracy. Skills for coding agents are collected in our Claude Code skills guide.
OpenAI retires GPT-3-era models: davinci-002, babbage-002 and gpt-3.5-turbo-instruct
On September 28, 2026, OpenAI shut down gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106 in its API and pointed developers to gpt-5.6-terra as the replacement. The deprecation notice is being discussed on Reddit, where the poster calls it the end of the GPT-3-era API lineage. Smaller models such as babbage may have no true drop-in replacement, because newer models differ in behavior, cost, latency and how they respond to prompts. Commenters argue that retired closed models should have their weights released, at least for research and behavior comparisons. The open alternatives from that period, GPT-J and GPT-Neo, remain available precisely because their weights are public. For running open-weight models yourself, see our open-source LLM guide.
Opus 5.5 rebuilt Pokémon Red: 25,000 lines of JavaScript and no image files
A developer released Pokémon Claude Red, a browser remake of Pokémon Red that, by the author's account, Claude Opus 5.5 wrote in Claude Code as roughly 25,000 lines of JavaScript over about three days. The game is playable in the browser and the source is on GitHub. All 151 Pokémon are drawn in code on a 320×180 canvas, while maps and data come from a disassembly of the original: 224 maps, the gyms, the Elite Four and even the MissingNo. glitch. On Reddit, most comments focus on legal risk, since a public remake built on someone else's intellectual property is likely to draw a takedown. A second showcase is an engineering bridge test: five models designed a 3D-printed bridge from 500 g of plastic, and the poster says Opus 5.5's design held about 130 lb, nearly 5x the runner-up.
Numbers and facts
- 89.6% on ARC-AGI-2 at $0.44 per task for GPT-6 Sol, according to ARC Prize; on ARC-AGI-3 it scored 4.6% on the standard harness and 23.0% on the provider harness.
- 1,414 skills in the vibe coding category and 21 for automated code review sit in the Skills for Claude Code section of AI SKILLS as of September 30, 2026 (platform catalog data).
- +14 points on BFCL v4 missing-function tasks came from Salesforce's Critical-State RL, which trains only the tool call whose action changes the outcome, per @dair_ai.
- 1,325 Elo on image-to-video for Pruna's P-Video-2 Pro, built on MiniMax H3, at 4.5–8 seconds per clip, according to Pruna.
Different views: does penalizing "wait" and "maybe" make models smarter?
A logit bias is a manual shift in the probability of specific tokens during generation; a −2 penalty on hesitation words lifted Qwen3.5-4B from 74% to 84% on 50 MATH-500 problems while cutting reasoning tokens by 19.4%. The experiment is written up on Reddit and reproduces the idea of a Meta paper in llama.cpp.
- For. The gain holds across quantizations: Q3_K_M rose from 52% to 66% and Q2_K from 12% to 24%, with 11–19% fewer reasoning tokens.
- Neutral. Liquid AI calls the behavior "doom looping" and released an antidoom tool that trains it out of the model instead of suppressing it at inference time.
- Against. Commenters point out that 50 problems is a small sample: the gain may simply come from more answers running to completion, and on open-ended tasks the penalty could quietly degrade results.
Tools and techniques
- Watch for silent model regressions. LiveNerf runs Opus 5.5 daily on a frozen set of 78 questions, uses Opus 5 as a control and flags deviations above 7.5 points. MarginLab keeps a similar tracker for Claude Code.
- Keep secrets out of the agent. According to @_philschmid, the Gemini Managed Agents Credentials API injects secrets on the wire only for trusted domains, so the agent itself never sees them.
- Put a human before tool calls. METR built a tool-call monitor that blocks an agent's action until a person reviews it. Pipelines with an approval step are in our automation templates.
In brief
- Gemini 3.8 Flash TTS took first place on Hume's index, clones a voice from 30 seconds of audio and covers more than 100 languages.
- Arrow 2 Telos set a new high for SVG generation at 1,624 Elo.
- Meta launched Meta Enterprise Platform to sell Muse and its AI APIs to companies and hired MongoDB's CEO to run it; MongoDB's stock fell about 20%, Ben's Bites reports.
- NVIDIA released OpenShell, an open-source sandbox whose policies govern agents' files, network and tools; more than 100 companies joined the safety stack, and OpenAI is not among them, Reddit users note.