Claude Opus 5.5 vs GPT-6 Sol: Price and Speed

On 22 September 2026, Anthropic and OpenAI each released a flagship model on the same day. Anthropic shipped Claude Opus 5.5, and OpenAI launched GPT-6 Sol and GPT-6 Luna. Neither announcement leads with a benchmark chart. Both lead with a price cut. That is the story worth reading past the model names.
What actually shipped
Claude Opus 5.5 is Anthropic’s new flagship, priced at $4 per million input tokens and $20 per million output tokens, a 20 percent cut against Opus 5. Cache reads drop 60 percent to $0.20 per million. Anthropic reports it performs comparably to Claude Fable 5.1 on most tasks while costing roughly 40 percent less to run, with output roughly 30 percent faster than Opus 5.
GPT-6 Sol and GPT-6 Luna extend OpenAI’s GPT-6 line with updated smaller models rather than a new frontier size. Sol lists at $2 per million input tokens and $10 per million output, half of the prior generation’s promotional pricing. Luna is priced for high volume work at $0.10 and $0.50 per million tokens. OpenAI says the newer models make about half as many factual errors as their predecessors on its internal accuracy test, and attributes the price drop to improvements in caching and inference rather than a smaller model.
| Claude Opus 5.5 | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| Input / output price (per million tokens) | $4 / $20 | $2 / $10 | $0.10 / $0.50 |
| Positioned as | Flagship, quality first | Mid tier, cost per task | High volume, simple jobs |
| Headline claim | Matches Fable 5.1 quality at lower cost | Fewer factual errors, cheaper per task than GPT-6 Astra | Cheapest tier for simple, repeatable work |
| Availability | Claude Platform, AWS, Google Cloud, Microsoft Azure | OpenAI API | OpenAI API |
How do the full Opus 5.5 and Sol families compare on price and speed?
Neither Opus 5.5 nor Sol ships alone. Each sits inside a three tier family that runs from a fast, cheap model up to the flagship, and the family view is what actually determines a routing strategy rather than a single headline price.
| Model | Tier | Input / output price (per million tokens) | Speed positioning |
|---|---|---|---|
| Claude Haiku 4.5 | Fast, high volume | $1 / $5 | Fastest and cheapest in the Claude lineup, built for high volume simple tasks |
| Claude Sonnet 5 | Balanced | $2 / $10 | Mid tier cost, the default for everyday agentic work |
| Claude Opus 5.5 | Flagship, daily use | $4 / $20, cache reads $0.20 | Anthropic’s fastest flagship yet, output over 30 percent quicker than Opus 5 |
| Claude Fable 5.1 | Frontier, hardest work | $10 / $50 | Slowest and most expensive Claude tier; Opus 5.5 now matches it on most tasks at under half the price |
| GPT-6 Luna | Fast, high volume | $0.10 / $0.50 | Cheapest OpenAI tier, built for simple, repeatable work |
| GPT-5.6 Terra | Balanced | $2.00 / $12.00 | Mid tier, cut roughly in half against GPT-5.5 list pricing; not yet refreshed to GPT-6 |
| GPT-6 Sol | Flagship, daily use | $2 / $10 | Reasoning heavy flagship; its GPT-5.6 predecessor reached up to 750 tokens per second on Cerebras hosting in preview |
| GPT-6 Astra | Frontier, hardest work | $10 / $50 | OpenAI’s most capable and aligned tier, positioned for computer use, browsing, software engineering, cybersecurity and science work that Sol is not tuned for |
Why the tier names. Both naming schemes encode a strategy, not just a brand. OpenAI’s tier names are drawn from the solar system on purpose: Sol for the sun, the brightest and most capable tier; Terra for earth, the solid middle; and Luna for the moon, the light and fast tier, with the generation number left free to advance on its own schedule so a tier name stays stable across refreshes, which is why Sol and Luna moved to GPT-6 in September while Terra is still on GPT-5.6. Anthropic’s names come from literary form instead: Haiku for the smallest, most economical model, Sonnet for the balanced mid tier, and Opus, meaning a major work, for the flagship. Fable sits a class above Opus, named from the Latin fabula, a story told, the same root behind Mythos, the restricted sibling model built on the same underlying system. In both schemes, the tier name is the durable product a customer builds against, and the version number is just how current the model behind it is.
Which tier to compare with which
Adding Astra to the table matters for a reason beyond completeness: it is the model that makes each vendor’s lineup line up. Compare by role, not by name similarity or by whichever pair a headline happens to put next to each other.
| Compare this Claude tier | With this OpenAI tier | Because |
|---|---|---|
| Claude Haiku 4.5 | GPT-6 Luna | Both the fast, cheap, high volume tier in their lineup |
| Claude Sonnet 5 | GPT-5.6 Terra | Both the balanced mid tier, and the closest price match of any pair here: $2/$10 against $2.00/$12.00 |
| Claude Opus 5.5 | GPT-6 Sol | Both the tier each vendor positions as the one you reach for by default on demanding work |
| Claude Fable 5.1 | GPT-6 Astra | Both the top, most capable tier held back for the hardest work, and price matched exactly at $10/$50 |
This is why the rest of this post benchmarks Opus 5.5 against Sol rather than against Astra: Opus 5.5 and Sol are the peer comparison by role, even though Sol is priced at half of Opus 5.5. Comparing Claude Sonnet 5 against GPT-6 Astra, or GPT-6 Luna against Opus 5.5, is comparing a mid tier or entry tier model against a flagship two or three tiers above it, which will make the cheaper side look worse than it is on capability and the pricier side look worse than it is on cost. Price and tier name do not always move together across vendors either: Sol is Astra’s default sibling but costs a fifth of it, while Fable 5.1 and Astra happen to be priced identically. Match by what each model is built to be reached for, then check the price, not the other way around.
Two things are worth flagging before you route traffic off this table. Anthropic has said Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, so the Sonnet and Haiku rows above are the prior generation still on sale next to a new flagship, not a matched set. On the OpenAI side, only Sol and Luna moved to the GPT-6 generation on 22 September. Terra was cut in price at the end of July but has not yet been rebadged, so it is still running as GPT-5.6 Terra beside a GPT-6 flagship and a GPT-6 entry tier. Read the table as each vendor’s current shelf, not as two cleanly aligned generations.
The shape holds regardless of the version numbers involved. Both vendors price a roughly 10x to 20x spread between their fastest and flagship tier, and both now sell the flagship close enough to the mid tier’s price that the mid tier’s main advantage is latency and volume, not cost per token.
Which model wins on key benchmarks, and by how much?
Claude Opus 5.5 leads the two benchmarks that are actually comparable across vendors, AutomationBench and Terminal-Bench 4.0, by margins of 6.8 and 8.6 percentage points. GPT-6 Sol leads the one benchmark where it edges ahead, Agents’ Last Exam, but only against the prior generation Opus 5, by a margin of 0.5 points that is effectively a tie.
What each benchmark actually measures
A score means little without knowing what it is a score on. Here is what each of the three actually tests, in one line each.
| Benchmark | What it measures |
|---|---|
| AutomationBench | Whether an agent can complete realistic, multi app SaaS business workflows end to end, across sales, marketing, operations, support, finance and HR, scored against the resulting state of the simulated environment rather than what the agent claims to have done |
| Terminal-Bench 4.0 | Whether an agent can operate a real terminal to complete real software engineering and systems work, writing and debugging code, configuring environments and recovering from failure, across domains from ML training to security to hardware |
| Agents’ Last Exam | Long horizon, economically valuable professional work across 13 industry clusters and 55 subfields, scored against real deliverables and milestones rather than short question and answer or isolated tool calls |
Cross vendor benchmark comparisons are noisy in general, because agentic evaluations measure a model plus a scaffold, and the scaffolds differ between labs. AutomationBench is the one row several independent trackers treat as directly comparable, because both vendors publish identical scores for the same older models, Opus 5 and Fable 5.1, inside it. Terminal-Bench 4.0 and Agents’ Last Exam are less clean: the scores below come from each side’s best publicly reported effort setting, and those settings are not always matched.

The chart covers only the two benchmarks where both vendors’ reported numbers line up on the same scale. Agents’ Last Exam is left out of it deliberately: the only public score there pits Sol against Opus 5, the prior Claude generation, not Opus 5.5, so plotting it next to the other two would visually imply a matchup that has not actually been run.
| Benchmark | 1st place | Score | 2nd place | Score | Margin | Relative gap |
|---|---|---|---|---|---|---|
| AutomationBench | Claude Opus 5.5 | 40.0% | GPT-6 Sol (max effort) | 33.2% | 6.8 pts | Opus 5.5 ahead by 20.5% |
| Terminal-Bench 4.0 | Claude Opus 5.5 (medium effort) | 52.5% | GPT-6 Sol (max effort) | 43.9% | 8.6 pts | Opus 5.5 ahead by 19.6% |
| Agents’ Last Exam | GPT-6 Sol (max effort) | 56.4% | Claude Opus 5 (prior generation, best) | 55.9% | 0.5 pts | Sol ahead by 0.9%, within noise |
Two caveats the table does not carry on its own. Opus 5.5 wins Terminal-Bench 4.0 at medium effort against Sol’s maximum effort, so the comparison already favors Sol’s best case against Opus 5.5’s default; no publicly reported score exists yet for Opus 5.5 at its own maximum effort on this benchmark, which means the true gap is at least 8.6 points and plausibly wider. And the one row a GPT-6 model wins compares Sol against Opus 5, not Opus 5.5. No Opus 5.5 score on Agents’ Last Exam has been published, so that result is best read as open rather than as a win over the current Claude flagship.
Cost changes the read entirely. Sol reaches its AutomationBench score at roughly 27 cents per task, and third party trackers describe Sol as the cheaper way to reach any score up to roughly the mid 40s, after which it has no further headroom to spend. Above that ceiling, Opus 5.5 is the tier that can go further, at a proportionally higher price. So the honest summary is not “Opus 5.5 is better” or “Sol is cheaper.” It is that Opus 5.5 currently holds the higher ceiling on the benchmarks that compare cleanly, and Sol is the cheaper way to clear a moderate bar, which is a different question and the one most production workloads should actually be asking.
A note on how much to trust these numbers. A benchmark score is not a guarantee of reliability, and none of the three above are run by an independent party against both vendors under identical conditions. AutomationBench’s own published analysis found that agents frequently report a task done when the underlying environment state is actually wrong, which affected the majority of each tested model’s failures in that study, Opus included. The figures in this post are a mix of vendor self reported results and third party trackers, gathered close to each launch, and every one of them can move as labs patch scaffolds, fix task bugs, or publish a later effort setting. Read the table and chart above as a snapshot worth validating against your own workload, not a settled ranking.
Opus 5.5 vs GPT-6 Sol at a glance
Pulling the price, speed and benchmark sections together into one view is where the trade-off actually lands: Opus 5.5 costs more and wins the benchmarks that compare cleanly, Sol costs less and is the cheaper way to clear a moderate bar.

Nothing in that scorecard is new information. It is the same numbers from the sections above, laid out so the trade-off reads in one glance instead of across three tables.
Why both labs are competing on cost per task now
A year ago, flagship launches led with a benchmark chart. Both of these lead with a price sheet and a cost per task claim. Three forces explain the shift.
- Inference cost has become the gating factor for agentic use cases. An agent that calls a model dozens of times per task multiplies the per token price by every intermediate step. At that volume, a 20 percent price cut changes what is economically viable to automate, in a way that a benchmark improvement alone does not.
- The frontier gap keeps closing faster than pricing power can hold. When Opus 5.5 matches competitors’ maximum effort performance at default settings, and Sol claims parity with a costlier predecessor at a fraction of the price, neither lab can rely on a capability lead to justify last year’s pricing. Cutting price is how you defend share when the model behind you is only a few months from catching up.
- Enterprise buyers evaluate on total cost of a workflow, not a leaderboard rank. A procurement conversation now asks what a thousand support tickets or a full codebase migration costs end to end, not which model scores highest on a held out exam. Both vendors are answering the question buyers are actually asking.
What this means for architecture and platform teams
Neither release changes what good practice looks like. It changes how much slack you have to be wrong about a model choice.
- Route by task, not by brand loyalty. A tiered strategy, Luna or a comparable small model for high volume simple jobs, Sol or Opus 5.5 for work that needs deeper reasoning, and the most capable tier reserved for the smallest slice of genuinely hard tasks, is now cheap enough to be the default rather than an optimization you get to later. This is the same argument behind context engineering: what you route to the model, and which model you route to, both shape the answer you get.
- Re-run your cost per task numbers, not just your cost per token numbers. A cheaper model that needs more retries or a larger scaffold can cost more end to end than a pricier model that gets the task right the first time. The benchmark comparisons above make that trade-off explicit; your own evaluation harness should too.
- Treat model portability as a live requirement, not a one time decision. Pricing moved twice in a single week across two vendors. An architecture that hard codes one model’s API shape, rather than routing through an abstraction layer, pays for every future repricing in engineering time rather than in a config change.
- Validate accuracy claims against your own workload before trusting vendor benchmarks. A halved error rate on an internal factuality test is a meaningful signal, but it is measured on the vendor’s own distribution of tasks. Confirm it holds on the kind of work you actually send the model before shifting production traffic.
The risk worth naming
Price competition this aggressive is good for buyers in the near term and worth watching in the medium term. Inference has real infrastructure cost behind it, and a sustained price war funded by continued capital raises is a different situation than one funded by genuine efficiency gains in caching and serving. If pricing resets upward once the current funding cycle tightens, workflows built assuming today’s per token economics will need re-costing. Build the routing layer that lets you absorb a repricing event without a redesign, rather than assuming today’s numbers hold indefinitely.
Key questions
Q1) What are Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna?
Claude Opus 5.5 is Anthropic’s new flagship model, released 22 September 2026, priced at $4 per million input tokens and $20 per million output tokens. GPT-6 Sol and GPT-6 Luna are OpenAI’s updated mid tier and high volume models released the same day, priced at $2/$10 and $0.10/$0.50 per million tokens respectively.
Q2) Is Claude Opus 5.5 or GPT-6 Sol better for coding tasks?
It depends on the task and the budget. On the two benchmarks that compare cleanly across vendors, Opus 5.5 leads: 40.0 percent versus Sol’s 33.2 percent on AutomationBench, and 52.5 percent versus 43.9 percent on Terminal-Bench 4.0, margins of 6.8 and 8.6 points. Sol is the cheaper way to reach a moderate capability bar, at roughly 27 cents per task, and has no further headroom above that bar, so the right choice depends on whether the task needs Opus 5.5’s higher ceiling or Sol’s lower cost per call.
Q3) Why are both companies cutting prices instead of only improving benchmarks?
Agentic workflows call a model many times per task, so price per token compounds quickly at scale, making cost per task as important as raw capability. Neither lab can hold last year’s pricing when the capability gap to competitors keeps narrowing, and enterprise buyers increasingly evaluate total workflow cost rather than a single benchmark score.
Q4) What should an architecture team do in response to these launches?
Build a routing layer that sends simple, high volume work to the cheapest adequate model and reserves the most capable tier for tasks that genuinely need it, and measure cost per completed task rather than cost per token alone. Keep the model choice abstracted behind an interface so a future repricing or capability shift does not require a redesign.
Why it matters
The headline is not which model wins a benchmark this week. It is that both Anthropic and OpenAI now treat cost per task as a first class competitive dimension alongside raw capability, and that changes what an enterprise can afford to automate rather than only how well it performs. Teams that build a routing and evaluation layer now will absorb the next repricing cycle as a configuration change instead of a migration.
I would be interested to hear how other teams are structuring their model routing given how frequently pricing and relative capability are shifting between vendors this year.
Disclaimer:
Pricing, benchmark figures and product details in this post are drawn from Anthropic’s and OpenAI’s own announcements and contemporaneous media and third party benchmark reporting as of publication, and may change as vendors update pricing or release further models. Benchmark comparisons in particular were not independently re-run for this post, mix vendor self reported and third party sourced figures, and are not always matched on effort setting or scaffold between vendors; treat them as directional rather than authoritative. All data and information provided on this blog are for informational purposes only. All the image sources used are for reference only. The author makes no representations as to the accuracy, completeness, correctness, suitability, or validity of any information on this blog and will not be liable for any errors, omissions, or delays in this information or any losses, injuries, or damages arising from its display or use. This is a personal view and the opinions expressed here represent my own and not those of my employer or any other organization.