Best AI Model for Coding: GPT-6 Sol vs Claude Opus 5.5
7 min read
Photo by Igor Omilaev on Unsplash
Best AI Model for Coding: GPT-6 Sol vs Claude Opus 5.5
OpenAI and Anthropic picked the same week, almost the same day, to update their flagship coding models. On September 22, OpenAI shipped GPT-6 Sol and GPT-6 Luna at half the price of their predecessors. Hours later, Anthropic answered with Claude Opus 5.5, a 40 percent cheaper Opus that drops the old five-hour usage caps entirely. If you build with these models for a living, the question on your mind is probably simple: which one should you actually be running your coding workloads on right now?
There's no single answer, but there is a clearer picture than there was a week ago. Here's what actually changed, what the two companies are claiming, and where the claims hold up.
What OpenAI Shipped: GPT-6 Sol and Luna
Sol and Luna aren't brand new names. Both were introduced back in July 2026 as the workhorse tier below GPT-6 Astra, OpenAI's top-end model. What shipped on September 22 is a refreshed generation of both, positioned specifically at cost-sensitive, high-volume use.
Sol is built for complex, professional work: coding, technical writing, long-horizon agent tasks. Luna is the cheaper sibling, meant for the high-volume grind, summarizing, extracting, answering, the kind of work you'd never want a $20-per-million-token model doing.
The headline is pricing. GPT-6 Sol now runs $2 per million input tokens and $10 per million output tokens, a straight 50 percent cut from GPT-5.6 Sol. Luna is even cheaper at $0.10 input and $0.50 output per million tokens, also half of what GPT-5.6 Luna cost. OpenAI credits the drop to "improvements in caching and inference" rather than a smaller model, and on the benchmarks it's shared, the new models don't look diminished for it.
On DeepSWE v1.1, a coding benchmark, Sol at max reasoning effort scored 68.8 percent, within about a point of Claude Fable 5's 69.9 percent, at roughly 80 percent lower cost by OpenAI's accounting. Luna, the budget model, hit 66.6 percent on the same benchmark, which OpenAI frames as 93 percent cheaper than running Claude Opus 5. On OSWorld 2.0, a computer-use benchmark, Sol at its highest effort setting scored 60.5 percent, which OpenAI says matches Claude Opus 5 at medium effort for a fraction of the price. Sol is also reportedly making about half as many factual errors as its predecessor, which OpenAI describes as approaching the reliability of its flagship Astra model.
Both are now live in ChatGPT Work, Codex, and the API as gpt-6-sol and gpt-6-luna, rolling out gradually to Plus, Pro, Business, Enterprise, and Edu accounts.
The Same-Day Answer: Claude Opus 5.5
Anthropic wasn't going to let that pricing move sit unanswered. Claude Opus 5.5 launched the same day, cutting the price of Opus by 40 percent: $4 per million input tokens and $20 per million output, down from Opus 5's rates, with a faster mode available at $8 in and $40 out for latency-sensitive work. Cache reads dropped even further, from $0.50 to $0.20 per million tokens, a 60 percent cut that matters a lot if your workflow leans on long, reused context. Anthropic also did away with the five-hour usage caps that had been a running complaint from heavy users of Claude Code and similar tools.
On the benchmark side, Anthropic published a comparison against its own previous model and against a competitor it calls Fable 5.1. Opus 5.5 scored 66.4 percent on Terminal-Bench 4.0 versus 52.3 percent for Opus 5 and 55.8 percent for Fable 5.1. On FrontierCode v1.1, it hit 54.4 percent against Opus 5's 48.0 percent. On CursorBench 4.0, a benchmark built around real editor workflows, Opus 5.5 reached 57.8 percent versus 46.6 percent for the previous Opus. Context window is unchanged at 1 million tokens, with synchronous output capped at 128,000 tokens and batch output up to 300,000 in beta.
Anthropic's real-world examples lean toward scale rather than speed: one user reportedly pushed a 680,000-line code migration through in under a day, and another audited and repaired a 200,000-line codebase in about three hours.
So, Which Is the Best AI Model for Coding Right Now?
Here's the part that's easy to miss if you just skim the press releases: OpenAI's benchmark comparisons for Sol and Luna were run against Claude Opus 5, not Opus 5.5, because Anthropic's new model landed the same day the comparisons went out. That doesn't make OpenAI's numbers wrong, but it does mean the "beats Opus at a fraction of the cost" framing is already measuring against a model that's no longer Anthropic's current one. A fair Sol-versus-Opus-5.5 benchmark comparison doesn't exist yet from either company, and it's worth waiting for independent evaluations (LMSYS-style leaderboards, or coding-specific ones like the DeepSWE or SWE-bench maintainers) before treating either vendor's own numbers as the final word.
What we can say with the numbers in hand: for raw coding-task accuracy on a shared benchmark like DeepSWE v1.1, GPT-6 Sol and Claude Fable 5 are within a point of each other, and Sol is doing it at a noticeably lower price per token. If your workload is high-volume and cost-sensitive, comparing Luna's $0.10/$0.50 pricing against anything in Anthropic's current lineup isn't close; Luna is built for a different job than Opus is. And if you're already deep in Claude Code and were waiting for the usage caps to loosen, Opus 5.5 answers that directly in a way price cuts alone don't.
In practice, the honest advice for most teams is to treat this as a testing moment, not a switching moment. Run your own eval set, the specific repo, the specific task types, against both gpt-6-sol and claude-opus-5-5 before moving spend. Vendor benchmarks are directionally useful and consistently favor the vendor that published them.
What This Means for Developers
The bigger trend here matters more than either individual release. Epoch AI's cost analysis, published around the same week, put the decline in AI inference cost at roughly 47 percent per quarter since 2023, something like a 13x annual drop when compounded. Two flagship-adjacent models cutting prices 40 to 50 percent in the same 24 hours isn't a coincidence; it's what that curve looks like from the outside. For anyone building products on top of these APIs, the practical upshot is that architecture decisions you made six months ago around cost, batching heavy work to a cheaper model, avoiding long context because of price, are worth revisiting now that the economics have moved again.
Key Takeaways
- GPT-6 Sol and Luna launched September 22 at half the price of their GPT-5.6 predecessors: $2/$10 per million tokens for Sol, $0.10/$0.50 for Luna.
- Claude Opus 5.5 launched the same day at 40 percent lower cost ($4/$20 per million tokens), with cache reads cut 60 percent and the old five-hour usage caps removed.
- OpenAI's Sol-vs-Opus benchmark comparisons were made against Opus 5, not the same-day Opus 5.5, so treat "beats Opus" claims as provisional until independent, apples-to-apples evaluations appear.
- On DeepSWE v1.1, GPT-6 Sol (68.8%) and Claude Fable 5 (69.9%) are close enough that price, not raw accuracy, is likely to be the deciding factor for most teams.
- The underlying trend, roughly 47 percent quarterly inference cost declines since 2023 per Epoch AI, is why two major vendors cut prices in the same week.
FAQ
Is GPT-6 Sol better than Claude Opus 5.5 for coding? There's no independently verified head-to-head yet. OpenAI's public benchmarks compare Sol against the older Claude Opus 5, not Opus 5.5, so the fairest answer right now is "run your own eval," not "yes" or "no."
How much cheaper is GPT-6 Sol than the previous version? Fifty percent. GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, half of GPT-5.6 Sol's pricing.
What's the difference between GPT-6 Sol and GPT-6 Luna? Sol is aimed at complex, professional-grade work like coding and long-horizon agent tasks. Luna is the high-volume, lower-cost option for simpler jobs like summarization and information extraction, priced at roughly a twentieth of Sol's rate.
Did Claude Opus 5.5 get a bigger context window? No. Context window stays at 1 million tokens, the same as Opus 5. The changes are pricing, cache costs, usage limits, and benchmark performance, not context length.
- GPT-6
- Claude Opus 5.5
- AI coding tools
- LLM benchmarks
- OpenAI
- Anthropic