OpenAI ships GPT-6.1 Sol and dots, its always-on agents

GPT-6.1 Sol lists at Sonnet 5.5's price and leads it at medium effort in an independent test; dots are agents that run around the clock. Four more stories inside.

Share
OpenAI ships GPT-6.1 Sol and dots, its always-on agents

OpenAI's DevDay brought two releases on Tuesday: GPT-6.1 Sol, at the same list price as Anthropic's Sonnet 5.5, and dots, agents that run around the clock on GPT-6 Astra. Pi, the minimal coding agent, now supports MCP after months of saying it didn't need it. Codex got a cloud version and a CLI refresh, though the CLI's changelog doesn't yet show everything OpenAI's recap promises. ChatGPT's $200 Pro plan buys less for new subscribers, and a new $500 tier is the only Pro plan with OpenAI's faster Ultrafast mode. And a new benchmark is tracking whether Opus 5.5 quietly got worse; it won't have an answer until late October.


What changed this week

OpenAI ships GPT-6.1 Sol at Sonnet 5.5's list price

OpenAI released GPT-6.1 Sol on Tuesday and says it matches GPT-6 Astra on DeepSWE, a test of real software-engineering tasks, "at roughly one-fifth of the cost." It lists at $2 per million input tokens and $10 per million output, the same as Anthropic's Sonnet 5.5 from Monday; cached input is $0.10 against Sonnet 5.5's $0.20. Artificial Analysis, an independent benchmarking firm, tested both at medium reasoning effort: Sol scored 48 on its Intelligence Index to Sonnet 5.5's 41, at $0.21 per task against $0.59. At maximum effort, Sonnet 5.5 leads, 56 to 52. Medium is Sonnet 5.5's default in Claude Code; Anthropic's API defaults to High, and OpenAI states no default for Sol. OpenAI's system card rates Sol "Critical capability in Cybersecurity" and gives it a 1.50% rate of misrepresentation on coding tests built to provoke it, above GPT-6 Sol's 1.30%. Compare them at the effort you run.


OpenAI launches dots, agents that run around the clock

OpenAI's dots are "always-on agents built to handle everything," running on GPT-6 Astra with their own cloud computer. They connect to "over 4,000 apps" and are rolling out on Pro, Business Premium and Enterprise plans. Work a dot does itself "uses nothing" from your plan, says Tibo Sottiaux of OpenAI's Codex team, but a Codex task it creates draws usage "as usual." Dots aren't a coding tool. Coding work one hands to Codex counts against your plan.


Pi adds MCP to its core after publicly rejecting it

Pi, Earendil's minimal coding-agent harness, now supports MCP, the protocol agents use to reach outside tools, in its core. Its site once carried "a proud declaration that Pi does not support MCP," and Mario Zechner, one of its makers, titled a post last November "What if you don't need MCP at all?" The team says a lot has improved in MCP, though not everything. Pi runs MCP tools through Codemode, a JavaScript sandbox where the agent chains tool calls in code. Using Pi? Upgrade to try it.


OpenAI brings Codex to the cloud and refreshes its CLI

Codex now runs "on a computer, remotely from a phone, or in the cloud from any device," OpenAI's DevDay recap says. The CLI got a "New full-screen interface" and "ways to manage parallel work," its developer account posted. Codex Security Cloud scans whole GitHub repositories on a schedule and prepares fixes. The recap also promises voice control and an "/agents" view; the changelog through version 0.159.2 lists neither. Check your version first.


ChatGPT's $200 Pro plan buys less; a $500 tier launches

OpenAI reopened its $200 Pro plan on Tuesday with "a lower usage allowance" for new subscribers, its help page says. "The monthly price remains $200." If you already had Pro 200 at the cutoff, you keep your old allowance "through Oct 29, 2026." The new Pro 500 gets "our highest usage allowance at 25 times the ChatGPT Plus allowance" and is the only Pro plan with Ultrafast, which generates tokens up to 8× faster in Codex. OpenAI gives no size for the cut. Budget for the smaller allowance after October 29.

About ChatGPT Pro tiers | OpenAI Help Center
Information about our paid subscription plan, Pro.

A new benchmark tracks whether Opus 5.5 got worse, with no verdict until around October 24

livenerf, a GitHub project, runs the same 78 questions against Opus 5.5 every day through headless Claude Code, to catch a model getting quietly worse after launch. Its README says "the first possible call is around 2026-10-24"; six of 30 days are in. It already names a limit: in validation it couldn't tell Opus 5 from Opus 5.5 at 99% confidence. It topped Hacker News on the question r/ClaudeAI keeps asking. No verdict yet. Check back after October 24.

GitHub - ninjahawk/livenerf: Benchmark for tracking model capability after release.
Benchmark for tracking model capability after release. - ninjahawk/livenerf

Also worth your time


Know someone who'd want this in their inbox? Forward it — that's how this grows. And if we got something wrong, or you think we buried the real story today, hit reply. A person reads every one.


The New Way is written with AI. It gathers the day's stories, checks them against their sources and drafts every summary. A person decides what runs and reviews every issue before we hit send.