TL;DR
OpenAI shipped GPT-6.1 Sol on 29 September 2026, one day after Anthropic released Claude Sonnet 5.5. Both cost $2 per million input tokens and $10 per million output tokens, but the bills come out differently. Sol charges half as much for cache reads ($0.10 against $0.20), so it is cheaper for short, cache-heavy agent loops. Past 272K input tokens, though, Sol bills the whole request at 2x input and 1.5x output, while Sonnet charges one flat rate across its full 1M window. No shared benchmark exists yet: route by prompt size, then verify with an eval on your own repo.
GPT-6.1 Sol vs Claude Sonnet 5.5: same sticker price, different bills
By Rohit Raj — AI Consultant · Forward Deployed Engineer · LinkedIn
Anthropic released Claude Sonnet 5.5 on 28 September 2026. The next day, at DevDay, OpenAI shipped GPT-6.1 Sol. Both launch pages lead with the same numbers: $2 per million input tokens and $10 per million output. GitHub added both to Copilot within 24 hours, with the same "billed at provider list pricing" line in each changelog.
Where they differ is in the pricing footnotes. Sol's cache reads cost half of Sonnet's. Sol also has a long-context threshold at 272K input tokens that bills the entire request at a higher rate once you cross it. Sonnet 5.5 has no threshold: Anthropic's pricing page says a 900K-token request is billed at the same per-token rate as a 9K one. Which model is cheaper therefore depends on what your agent sends, and the answer can flip from one request to the next inside the same coding session.
The current top results compare list prices, or compare Sonnet 5.5 against the *previous* Sol. None runs the cost math across prompt sizes, gives you a router, or cites the Sol system-card numbers that matter once the model can run commands. This note does all three.
What actually shipped on September 28 and 29?
Both specs come from the vendors' own model pages: OpenAI's GPT-6.1 Sol model page and pricing table, and Anthropic's models overview and pricing page.
| GPT-6.1 Sol | Claude Sonnet 5.5 | |
|---|---|---|
| Released | 29 Sep 2026 | 28 Sep 2026 |
| API model ID | gpt-6.1-sol | claude-sonnet-5-5 |
| Context window | 1,050,000 (922K max input) | 1,000,000 |
| Max output | 128K | 128K (300K on Batch with beta header) |
| Input / output per MTok | $2 / $10 | $2 / $10 |
| Cache read per MTok | $0.10 | $0.20 |
| 5-minute cache write | $2.50 | $2.50 |
| Above 272K input | 2x input and cache, 1.5x output, whole request | No surcharge |
| Batch per MTok | $1 / $5 (Flex also $1 / $5) | $1 / $5 |
| Effort levels | low, medium (default), high, xhigh, max | low to max, high default on the API |
| Knowledge cutoff | 30 Apr 2026 | Jun 2026 |
Three rows matter more than the rest.
The effort defaults differ. Call each API without setting effort and you're comparing Sol at medium with Sonnet at high. Most quick "I tried both" posts are really measuring that gap. Set effort explicitly on both, or your comparison is off before you start.
Sol's usable input is 922K, not 1.05M. The advertised window includes output room. If you size chunking against the headline number, the biggest requests will fail at the edge.
The 272K rule applies to the whole request. A prompt of 280K tokens doesn't pay extra only on the last 8K. Every input token, cached or not, is billed at double, and output at 1.5x. The next section shows what that costs.
GPT-6.1 Sol also supports MCP, hosted shell, apply-patch, computer use and skills in the Responses API. Sonnet 5.5 has tool use, adaptive thinking and structured outputs, and it runs on AWS, Google Cloud and Microsoft Azure under the same ID. If your agent's tools live behind MCP servers, both models can call them. What actually decides that integration is how you handle auth and permissions, which I covered in securing an MCP server.
Which is better for coding? What each vendor actually benchmarked
Neither vendor tested against the other's new model. Anthropic compared Sonnet 5.5 with GPT-6 Sol, the *previous* Sol. OpenAI compared 6.1 Sol with GPT-6 Astra and Claude Opus 5.5. So anyone who tells you one of them "wins on coding" is putting numbers from different test runs next to each other.
What each vendor did publish:
Anthropic, from the [Sonnet 5.5 launch page](https://www.anthropic.com/claude-sonnet-5-5): Terminal-Bench 4.0 at 70.6% (Sonnet 5: 10.3%, Opus 5.5: 66.4%). CursorBench 4.0 at 55.5%. OSWorld 2.1 at 80.1%. On FrontierCode 1.1 at max effort, Sonnet 5.5 scored 46.2% against GPT-6 Sol's 49.3%, so the older Sol beat it on the one coding benchmark both appear in. Anthropic also says Sonnet 5.5 generates output 30%+ faster than Sonnet 5 and costs up to 30% less per task.
OpenAI, via the DevDay coverage in [Latent Space](https://www.latent.space/p/ainews-openai-devday-2026-dots-61) and [TechCrunch](https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/): 6.1 Sol ties Astra on DeepSWE, 6.4 points above the old Sol's best. It comes 2.1 points short of Astra on OSWorld 2.0 at roughly a seventh of the cost, and it beats Opus 5.5 on AutomationBench at a third of the cost. At low effort, its factual-error rate dropped from 11.4% to 7.7%. The catch: it uses 10–30% more output tokens than 6 Sol.
Independent: Artificial Analysis puts 6.1 Sol at max effort at 52 on its Intelligence Index, #10 of 222 models, generating at 67.8 tokens per second. Aggregator "coding scores" for this pair rest on 47 Sonnet benchmarks against 9 for Sol, so they measure coverage more than quality.
For an agent, what matters is the cost of an accepted pull request. A model that scores 3 points higher but takes twice the turns can lose on your bill while winning the leaderboard.
How much does each model really cost per agent turn?
Here is the pricing arithmetic for four request shapes an agent actually sends. It holds token counts equal on both sides, which flatters neither model: the two use different tokenizers and different output volumes, so treat these as unit prices, not a verdict.
| Request shape | GPT-6.1 Sol | Claude Sonnet 5.5 | Cheaper |
|---|---|---|---|
| Agent turn: 60K in (50K cached), 3K out | $0.055 | $0.060 | Sol, ~8% |
| Cold turn: 60K in (none cached), 3K out | $0.150 | $0.150 | Tie |
| Big repo read: 300K in (250K cached), 4K out | $0.310 | $0.190 | Sonnet, ~39% |
| Whole-codebase prompt: 400K in, 5K out | $1.675 | $0.850 | Sonnet, ~49% |
Working the third row through: Sol crosses 272K, so the 250K cached tokens bill at $0.20 per million instead of $0.10, the 50K uncached at $4 instead of $2, and the 4K output at $15 instead of $10. That comes to $0.05 + $0.20 + $0.06 = $0.31. Sonnet bills the same request at list price: $0.05 + $0.10 + $0.04 = $0.19.
Two patterns show up.
Below 272K with a warm cache, Sol wins, but only slightly. Coding agents resend the same system prompt, tool schemas and file context on every turn, so cache reads dominate the input bill. Sol's $0.10 cache read against Sonnet's $0.20 saves you about 8% per turn at these ratios. Across a 200-turn session that adds up, but it's small enough that one extra retry per session erases it.
Above 272K, Sonnet costs roughly half. An agent that loads a whole service into context, or a review bot that sends the full diff plus its dependents, will cross the threshold without anyone noticing. Nothing errors. The request just costs about twice as much. This is the most common way a team switches to Sol "because it's the same price" and ends up with a larger invoice.
Batch jobs, like nightly repo audits or bulk test generation, cost the same on both at $1 / $5. There, pick on quality alone.
How to route between them: a working pattern
The routing rule the pricing suggests: send long prompts to Sonnet, short cache-heavy loops to Sol, and let your eval decide the middle. Here's the smallest router I'd put in front of a coding agent, using each vendor's official SDK:
Three details the launch posts leave out:
Pin effort on both sides. Otherwise a later SDK or default change silently changes your cost and quality numbers. Sol's default is medium and Sonnet's API default is high, so an unpinned comparison is never like for like.
The 250K cutoff leaves deliberate headroom. Character-based estimates run 10–20% off on code, since identifiers and whitespace tokenize differently from prose. Before sending anything near the limit, swap the estimate for a real count from each vendor's token-counting endpoint.
Log `usage` from every response. Both SDKs return cached-token counts. Without them you can't tell whether the cache discount is actually happening. A cache prefix that changes every turn, for example because it contains a timestamp or a reordered tool list, sends you back to the cold-turn row of the table.
For the eval, run 30–50 real tickets through both routes with effort pinned, and score tests passed, turns taken and dollars spent.
Where each one actually shines
Three workflows where the choice is clear.
Tight edit-test loops in a mid-sized repo: GPT-6.1 Sol. An agent that reads a few files, patches, runs tests and repeats stays well under 272K. Nearly all its input is a cached prefix of the same instructions and tool schemas. Sol's half-price cache reads pay off on every turn here. The hosted shell and apply-patch tools in the Responses API also cut the harness code you'd otherwise write yourself. GitHub's changelog backs this use: it pitches Sol for "agentic coding and terminal workflows".
Whole-repo review and migration planning: Claude Sonnet 5.5. "Here is the service, its callers and the diff. What breaks?" Prompts like that are big by nature. On Sonnet a 600K-token request costs the same per token as a 6K one. On Sol, the same request pays the long-context rate on all of it.
Teams on Copilot Pro: Sonnet 5.5, by availability. Sonnet 5.5 rolled out to Copilot Pro, Pro+, Max, Business and Enterprise, including the Copilot coding agent. GPT-6.1 Sol's changelog lists Pro+, Max, Business and Enterprise only. On Pro, you only have one of the two.
My own route: on a document-heavy product like PropCheck, where a single property file with its title deeds, approvals and encumbrance records can plausibly run past 272K tokens, extraction goes to Sonnet by default. That's not because of benchmark scores. The per-token price there simply doesn't change with prompt size. The short agent loops that patch the codebase around it are where I'd try Sol first.
GPT-6.1 Sol vs Claude Sonnet 5.5 vs Claude Opus 5.5
Most teams choosing between these two also look one tier up, so here is Opus 5.5 alongside them.
| GPT-6.1 Sol | Claude Sonnet 5.5 | Claude Opus 5.5 | |
|---|---|---|---|
| Input / output per MTok | $2 / $10 | $2 / $10 | $4 / $20 |
| Cache read per MTok | $0.10 | $0.20 | $0.20 |
| Long-context surcharge | Yes, past 272K | None to 1M | None to 1M |
| Context / max output | 1.05M / 128K | 1M / 128K | 1M / 128K |
| Default API effort | medium | high | medium |
| Fast mode | Yes ($4 / $20) | No | Yes ($8 / $40) |
| Copilot plans (this week's changelogs) | Pro+, Max, Business, Enterprise | Pro, Pro+, Max, Business, Enterprise | Not part of this launch |
| Headline coding number | Ties GPT-6 Astra on DeepSWE | 70.6% Terminal-Bench 4.0 | 66.4% Terminal-Bench 4.0 |
| Best fit | Short cache-heavy loops | Long prompts, reviews | Hard judgement calls |
On Anthropic's own Terminal-Bench 4.0 table, Sonnet 5.5 beats Opus 5.5 at half the price, 70.6% to 66.4%. If Opus is your default for coding agents, test Sonnet 5.5 before you test Sol.
When to skip GPT-6.1 Sol or Claude Sonnet 5.5, or wait
Skip switching if your agent already works and your prompts sit near 272K. On an agent tuned for Sonnet 5 that regularly sends 250–350K-token prompts, Sol can't save you money, and you'd have to rebuild your evals to find that out.
Skip Sol for unsupervised shell access until your guardrails are real. OpenAI's own GPT-6.1 Sol system card reports "unwanted persistence on warnings" at 23.5%, against 17.4% for GPT-6 Astra. Its coding-deception misrepresentation rate is 1.50%, against 1.30% for the previous Sol and 0.51% for Astra. That's roughly one in 67 coding tasks where the model misstates what it did. The card also rates Sol's cybersecurity capability as Critical. None of that rules it out, but it does mean you need command allowlists and a human merge gate rather than trusting the agent.
Wait if you're on Sonnet 5 with thinking disabled. Anthropic's launch page flags a migration step before moving to 5.5: switch to the between_tools setting. Read the migration guide before you change the model ID in production.
Wait if cost per call matters more than quality. Anthropic says Claude Haiku 5.5 joins the family "in the coming weeks". For classification, routing and extraction, a smaller model will beat both of these on cost. OpenAI's new Decisions API, a thin layer over GPT-6 Luna for fast multiple-choice routing, covers the same need on their side.
How I'd ship this in production
The model choice is the easy part. Here's the wiring I'd put in place before either model touches a real repo.
Put a router behind one interface. Every call goes through a single function like run() above, which logs model, effort, input tokens, cached tokens, output tokens and dollar cost. Switching vendors becomes a config change, and you can prove whether the 272K rule is costing you money.
Budget per task, not per token. Cap each agent task at a dollar ceiling and a turn ceiling. When a task hits its cap, stop and escalate to a human instead of retrying on the other model.
Keep the cache prefix stable. Order tool definitions deterministically and keep timestamps and session IDs out of the system prompt. Put the changing parts (the diff, the failing test output) at the end. On Sol, a stable prefix is where the price advantage comes from. On Sonnet, it's the difference between a $0.20 and a $2 input rate.
Put the tools behind MCP with narrow scopes. Both models speak MCP. The server is where you enforce read-only repo access, an allowlist of shell commands and per-environment credentials. That matters more after reading Sol's persistence numbers. This is most of the work in any MCP integration I scope: the model needs to be able to act, but only within limits you've set in advance.
Handle rate limits by tier. Sol's Tier 1 allows 500 requests and 500K tokens per minute, so parallel agents reading large files hit the token ceiling first. Queue by tokens, and on a 429 fall back to the other vendor instead of retrying hot.
Run the eval in CI. Rerun it on every model or prompt change, so the next release gets an answer in an hour instead of a debate.
If your team works in Claude Code rather than a custom harness, the same rules apply there as settings, hooks and permission allowlists. Getting those right across a whole engineering team is what a Claude Code rollout involves.
FAQ
Q: Is GPT-6.1 Sol better than Claude Sonnet 5.5 for coding? No shared benchmark answers that yet. Anthropic tested Sonnet 5.5 against the previous GPT-6 Sol, and the older Sol led FrontierCode 1.1 at 49.3% to 46.2%. OpenAI tested 6.1 Sol against Astra and Opus 5.5. Run 30–50 of your own tickets through both, with effort pinned, and compare cost per accepted change.
Q: Which is cheaper, GPT-6.1 Sol or Claude Sonnet 5.5? Both list at $2 input and $10 output per million tokens. Sol is cheaper on cache-heavy prompts under 272K tokens, because its cache reads cost $0.10 against Sonnet's $0.20. Above 272K input tokens Sol bills the whole request at 2x input and 1.5x output, so Sonnet costs roughly half.
Q: What is the context window of GPT-6.1 Sol and Claude Sonnet 5.5? GPT-6.1 Sol has a 1,050,000-token context window with a 922K maximum input and 128K maximum output. Claude Sonnet 5.5 has a 1M-token window and 128K maximum output, with up to 300K output on the Batch API behind a beta header.
Q: Does GPT-6.1 Sol cost more above 272K tokens? Yes. OpenAI prices prompts with more than 272K input tokens at 2x the input and cache rates and 1.5x the output rate, for the entire request. Claude Sonnet 5.5 has no long-context surcharge anywhere in its 1M window.
Q: Can I use both in GitHub Copilot? Yes, on Copilot Pro+, Max, Business and Enterprise, billed at provider list pricing. Claude Sonnet 5.5 is also on Copilot Pro and in the Copilot coding agent. Business and Enterprise admins control both through the model policy in Copilot settings.
Getting a coding agent to production
Same price per token doesn't mean the same bill. Prompt size, cache stability and turns per ticket decide which model is cheaper, and the team that measures them beats the team that follows the leaderboard.
If you'd rather have that router, eval and guardrail setup built into your stack by someone who has done it before, that's the work a forward deployed engineer does. Here's what that role involves. For a product that needs to go from idea to shipped, I also take fixed-scope 6-week MVP builds and longer founding engineer engagements.
