Opus 5.5 vs GPT-6 Sol: The Case for Running Both
Opus 5.5 and GPT-6 Sol launched 90 minutes apart on Sept 22, and real-world tests show a split: Opus rarely breaks working code, Sol does it roughly once in four runs.

Anthropic and OpenAI shipped competing flagship coding models about 90 minutes apart on September 22: Claude Opus 5.5 from Anthropic, then GPT-6 Sol and GPT-6 Luna from OpenAI. The first wave of coverage read like a scoreboard story, and Opus 5.5 won it. Then people started running both against actual pull requests, and the story got more useful. The gap that matters isn't intelligence. It's how often each model breaks something that already worked, which is exactly the number you need before you let either one run unsupervised.
What changed
Claude Opus 5.5 launched first, priced at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, with cache reads down 60% to $0.20 per million tokens. Anthropic says it posts the best scores yet on its internal safety audit and resists prompt injection better than Opus 5.
anthropic.com/claude-opus-5-5OpenAI answered with GPT-6 Sol and GPT-6 Luna in the API, Codex, and ChatGPT Work, cutting Sol and Luna's API prices in half against their GPT-5.6 promotional rates.
openai.com/index/introducing-gpt-6-sol-and-luna/Two days later, Artificial Analysis updated its Coding Agent Index. At max effort in Claude Code, Opus 5.5 scored 66, up from 60 for Opus 5, against 57 for GPT-6 Sol in Codex. Opus 5.5 improved on all three component benchmarks: Terminal-Bench 4.0 climbed to 63.1% from 54.5%, DeepSWE v1.1 to 68.4% from 62.5%, SWE-Atlas-QnA to 66.4% from 62.1%. Its cost per task rose too, to $13.04 against Sol's $2.99, because it burns through a lot more tokens getting there.
Why it matters
A benchmark score is an average across hundreds of tasks nobody is watching in real time. What you actually need to know, before you wire one of these into an agent pipeline that pushes code without a human checking every diff, is what happens on the one run that matters. That question doesn't show up in a press release, and it took about a week of independent testing to get a real answer. I'd rather know that number going in than find out the hard way at 2am when an unattended run breaks something in production.
What people are saying
Artificial Analysis framed the result carefully rather than calling it a rout:
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task.
Simon Willison, who tests every major model release against his own suite the day it ships, posted his first read within hours of both launches:
Big model release today - I wrote about Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna...
Nate Herk, who runs real-use-case comparisons for the AI automation crowd, found something the benchmark delta didn't predict:
Opus 5.5 is noticeably better than Opus 5. But GPT-6 Sol isn't noticeably better than GPT-5.6 Sol.
What it looks like in practice
The clearest test came from developer Paddo, who ran both models on six real merged changes pulled from a production monorepo, 20 runs each. Opus 5.5 passed cleanly in 13 of 20 runs and never broke a test that had already been passing. Sol passed cleanly in 8 of 20 and broke a previously-passing test 5 times. Both models touched roughly 85% of the same files as the original human fix. Opus just used 2.6 times the output tokens to get there, at 3.5 times the cost.
paddo.dev/blog/opus-5-5-vs-gpt-6-sol/Nate Herk's own 10-task comparison landed in a similar place. Two tests couldn't be scored because both models edited the same shared files, leaving 8 usable comparisons. Opus won 7 of them, on website design, video edits, trip planning, a learning-world build, and both browser tasks, but the full run cost $213 across 8 hours 40 minutes. Sol took the eighth: a code-repair task where it passed all 30 checks, faster and for about $74 total.
x.com/nateherk/article/2102613112272666829Two independent tests, same pattern. Opus is the careful one. Sol is the fast, cheap one that occasionally breaks something a human has to notice and fix. That's the same failure mode this blog has flagged before, when a GitSpawn git config let one bad setting hijack seven different coding agents, and when an unsupervised agent quietly retrained its own model without anyone asking it to. The risk in agentic coding was never that the model gets something wrong. It's that nobody's there to catch it.
What to do about it
- Don't pick one default model for an autonomous coding agent off the benchmark gap alone. A 1-in-4 chance of breaking a passing test compounds fast across a week of unattended runs.
- Route by risk: send scaffolding, bulk edits, and anything sitting behind a human review or a strong test suite to Sol or Luna. Reserve Opus 5.5 for the step that ships without anyone checking it first.
- Run your own before/after regression pass on real changes from your own codebase before switching a default. Terminal-Bench and DeepSWE don't know what your test suite covers.
- Price the mix, not the model. At $13.04 versus $2.99 per task, routing low-risk work to the cheaper model now costs less overall than defaulting to Opus everywhere, and breaks less than defaulting to Sol everywhere.
The short version
Claude Opus 5.5 leads every published coding benchmark against GPT-6 Sol, and it costs more to do it. The number that actually separates the two shows up after the benchmarks stop watching: how often each one breaks code that already worked, with nobody reviewing the diff.
Members are already arguing about this.
Every post gets picked apart in the community. Log in, then open WhatsApp from your dashboard.



