AI coding claims vs what shipped
Launch claims from AI coding vendors, checked against their own docs, tables and changelogs, with where each one stands now.
As of September 28, 2026. Next full re-check due by December 27, 2026.
Each entry is a public claim about an AI coding tool or model, the receipt we checked it against, and where it stands now. Most receipts are the vendor's own page. We add an entry the day an issue finds one.
Announced before it shipped
Anthropic said Claude Mods were "landing now" on Sept. 14. By Sept. 28 they hadn't shipped
Mods are a planned way to extend Claude Code with your own code. Claude Code's creator, Boris Cherny, posted "Claude Mods are landing now" on September 14. That day they ran only behind an opt-in environment flag, and the GitHub issue said "shipping in N weeks." On September 28 the issue is open. Claude Code's changelog, through version 2.1.284, has no Mods entry. The docs page returns 404.
Claude Code's $100 or $250 cloud credit is spent before your plan, not added to it
Anthropic's @ClaudeDevs account announced cloud sessions, which keep Claude Code working with your laptop closed, with "a one-time credit to try them: $100 on Pro, $250 on Max". That read as extra usage. Four and a half hours later the same account posted "Sorry for any confusion!" and clarified that it's a credit "your cloud sessions spend first, before falling back onto your normal plan usage." The cloud sessions docs still don't mention it.
Benchmark headlines vs the vendor's own table
Xiaomi says MiMo-V2.6 Pro matches Opus 5 and GPT-5.6 Sol. Its own table backs half of that
Xiaomi's launch post says its new model "performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks." Its model card lists 11 agent tests. Against Sol it holds: MiMo leads on 8. Against Opus 5 it doesn't. Opus 5 leads on 8, MiMo on 2, and one is a tie. The widest gap is Terminal Bench 4.0, 34.9 to 49.0.
Cognition's SWE-2 is within a point of Fable 5.1 on one test and half its score on another
Cognition, the company behind the Devin agent, headlines its SWE-2 coding model at "50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper." Fable 5.1 is Anthropic's model, and the number holds. The same post's table also lists Terminal-Bench 4, a test of agents working in a terminal (Xiaomi's card above writes it "Terminal Bench 4.0"). There SWE-2 scores 27.3%, Fable 5.1 55.8%.
OpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3 with one harness and 62.7% with the standard one
OpenAI's launch thread said its GPT-6 Astra model "is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0". ARC-AGI-3 is a puzzle benchmark, and its owner, ARC Prize, reports two scores. With its standard harness, the scaffold that runs the model through the puzzles, Astra gets 62.7%. Its 99.9% needs a "Provider Adapter" harness that keeps the model's hidden reasoning between requests.
Numbers that lost their baseline
"8x more code" at Anthropic is measured against 2021 to 2025
A post by Addy Osmani said Anthropic engineers "ship 8x more code per quarter." Anthropic's own blog gives the comparison: "8x as much code per quarter as they did from 2021-2025." That's a multi-year baseline. The post read as a recent jump.
Where we got it wrong
The correction runs on the original post, labeled, with the time.
We compared Opus 5.5's cost at the wrong setting for two days
Anthropic said Opus 5.5 costs 40% less to run than Opus 5 at default settings. We checked that at maximum effort, where the two cost about the same per task, and reported them as about level. At each model's default, two benchmarks put Opus 5.5 at 37 to 38% of Opus 5's cost per task: a bigger saving than claimed. The corrected post.
We implied Gergely Orosz misread Google's Antigravity terms. He hadn't
Orosz, who writes The Pragmatic Engineer, read the terms of Google's Antigravity coding tool as risking a Google-account ban. We wrote that the clause named only the Antigravity and Gemini CLI accounts. When he posted, it said "termination of your account." Google narrowed it the same day. Correction in that day's issue.
We said a game wiki hadn't responded to its AI-crawler trap. It had, for hours
The Cutting Room Floor, a video-game wiki, served AI agents a file-wiping payload. We wrote that the site had made no statement. Its co-founder had been responding publicly on Bluesky for hours before the issue went out. Correction on the post.
The New Way is written with AI. It gathers the day's stories, checks them against their sources and drafts every summary. A person decides what runs and reviews every issue before we hit send.