🌏 中文版
What changed
Anthropic launched Claude Sonnet 5.5 on 2026-09-29 with a headline that reads "costs up to 30% less for most work" — the kind of line that makes you assume a price cut. Check the official pricing page, though, and the input, output, and cache-read rates are identical to Sonnet 5, down to the cent: $2.00 / $10.00 / $0.20 per 1M tokens. This isn't a price cut. It's the same price for a model that needs fewer tokens to do the same work — a 30% cost reduction earned through efficiency, not a smaller number on the rate card. Those two things mean very different things for budget planning.
Before / after
| Item | Sonnet 5 | Sonnet 5.5 | Change | Effective |
|---|---|---|---|---|
| Input | $2.00/1M tokens | $2.00/1M tokens | 0% | 2026-09-29 |
| Output | $10.00/1M tokens | $10.00/1M tokens | 0% | 2026-09-29 |
| Cache read | $0.20/1M tokens | $0.20/1M tokens | 0% | 2026-09-29 |
| Cache write (5m) | $2.50/1M tokens | $2.50/1M tokens | 0% | 2026-09-29 |
| Batch input | $1.00/1M tokens | $1.00/1M tokens | 0% | 2026-09-29 |
| Batch output | $5.00/1M tokens | $5.00/1M tokens | 0% | 2026-09-29 |
The flat rate across the board is itself the story worth noting: Anthropic chose to position Sonnet 5.5 as "do more at the same price point," rather than compete on rate cuts the way OpenAI did with GPT-6 Sol/Luna (see our 09-28 pricing writeup).
Cost math
Anthropic doesn't give a single formula for the "30% less" claim — it's illustrated per scenario. The most concrete number in the launch post comes from Slack's own testimonial: Sonnet 5.5 beat Sonnet 5 on "almost all offline Slackbot evals, in fewer steps and with about 14% fewer output tokens."
Scenario: A Slackbot-style agent handling 10,000 tasks/day, averaging 1,500 input tokens + 800 output tokens per task, switching from Sonnet 5 to Sonnet 5.5 and matching Slack's reported 14% output-token reduction.
| Sonnet 5 | Sonnet 5.5 | Monthly savings | |
|---|---|---|---|
| Input cost/month (same volume) | $900 | $900 | $0 |
| Output cost/month | $2,400 | $2,064 | $336 |
| Total | $3,300/month | $2,964/month | $336 (↓10.2%) |
That conservative estimate lands at 10.2% — well short of the "up to 30%" headline. The gap comes from where the 30% figure is measured: specific benchmarks at specific effort settings. On FrontierCode at High effort, Sonnet 5.5 matches a 10-point-higher score than Sonnet 5 at roughly one-fifteenth the cost per task. On Terminal-Bench 4.0 at Medium effort (the Claude Code default), Sonnet 5.5 beats Sonnet 5's best score for less than a tenth of the cost per task. These are cases where a lower effort setting reaches a higher score — not an average across all workloads. Tasks that already need Sonnet 5's high-effort, long-context settings will see savings well below 30%.
Impact on developers and teams
Who benefits most
Workloads that are high-volume, repetitive, and can run at a lower effort setting benefit the most: classification, document/slide formatting, routine bug fixes. These tasks didn't need Sonnet 5's high-effort setting to begin with, so switching to Sonnet 5.5 lets you drop the effort level while matching or beating quality — capturing the token-reduction effect twice over. Complex tasks that already required Sonnet 5's high-effort, long-context mode will see a much smaller improvement, and might even cost more overall once you account for the extra verification work.
Competitive landscape
Current flagship pricing (input/output, USD per 1M tokens, standard short-context rate):
| Model | Input | Output | Note |
|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | Cheapest high-capability model (see our 09-28 pricing post) |
| Claude Haiku 4.5 | $1.00 | $5.00 | Anthropic's cheapest general-purpose model |
| GPT-6 Sol | $2.00 | $10.00 | Matched Sonnet-tier pricing since its 09-22 launch |
| Claude Sonnet 5.5 (new) | $2.00 | $10.00 | Rate-for-rate tie with GPT-6 Sol — the fight is now over token efficiency, not price |
| Claude Opus 5.5 | $4.00 | $20.00 | Launched 09-22, 20% below Opus 5 (see our 09-25 model card) |
With Sonnet 5.5, Anthropic and OpenAI are now going head-to-head at the exact same $2/$10 price point: GPT-6 Sol got there by cutting 50% off its predecessor's promotional rate, while Sonnet 5.5 got there by holding price flat and claiming fewer tokens per task. Two different routes to the same conclusion — both vendors are now competing on total cost per task rather than the sticker price per token. That means the number worth watching going forward is less "$/1M tokens on the rate card" and more "how many tokens did this task actually burn."
What to do
- If you're running customer-support or document-generation workloads on Sonnet 5: test a switch to Sonnet 5.5. Same price, likely fewer tokens, and low migration cost (API-compatible — just swap the model ID to
claude-sonnet-5-5). - If your workload already relies on Sonnet 5's high-effort, long-context settings: run your own A/B test before assuming you'll hit the advertised 30% — the real-world gap can be much larger than expected.
- If you're comparing Sonnet 5.5 against GPT-6 Sol: the two are priced identically ($2/$10), so the decision now comes down to benchmark fit for your task, not price.
- If your team used "thinking off" as a cost-saving lever: Sonnet 5.5 replaces the old toggle with a new
between_toolssetting — check the official migration guide before upgrading to avoid a compatibility break.
Takeaway
Seeing a headline like "costs up to 30% less" makes you want to go check the pricing page for a rate cut — this time the rate hadn't moved a cent, and the entire 30% came from "the same task using fewer tokens." That's worth remembering: a vendor's "cheaper" claim doesn't automatically mean a lower list price. Efficiency gains can be packaged into a pricing headline just as easily as an actual rate cut, but the two have very different reliability for long-term budget planning — a rate cut is a guaranteed constant, while efficiency gains swing wildly by task type and can shrink toward zero in the wrong workload.
References
Loading...