Claude Sonnet 5.5 Hits 70.6% on Terminal-Bench With Opus-Grade Limits

Anthropic released Claude Sonnet 5.5 on September 28, 2026 at $2 per million input tokens and $10 per million output — unchanged from Sonnet 5 and half of Opus 5.5’s $4/$20. Anthropic reports 70.6% on Terminal-Bench 4.0 against Opus 5.5’s 66.4%, and shipped it with the cyber restrictions previously reserved for its flagship models.

What is Claude Sonnet 5.5?

Claude Sonnet 5.5 is Anthropic’s mid-tier coding and agent model, released September 28, 2026. It carries a 1M-token context window, 128K maximum output, five effort levels, and a June 2026 knowledge cutoff. The price did not move from Sonnet 5.

The headline is that a mid-tier model now outscores the flagship on the benchmark Anthropic uses to sell coding performance. Terminal-Bench 4.0: 70.6% for Sonnet 5.5, 66.4% for Opus 5.5.

The more consequential change is invisible in the benchmark table. This is the first Sonnet to ship with frontier-grade cybersecurity restrictions, and the first with classifiers that block reasoning extraction.

Where Claude Sonnet 5.5 runs

The model is live on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure. Weights are closed.

Three-hyperscaler day-one availability matters more than it sounds. It means enterprise procurement can route the model through existing cloud commitments rather than a new vendor contract, which compresses the sales cycle from quarters to hours.

How much does Claude Sonnet 5.5 cost?

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 per million and cache writes at $2.50 per million, per MarkTechPost. That is identical to Sonnet 5 and exactly half of Opus 5.5’s $4 and $20 rates.

Anthropic also claims the model generates output “more than 30% faster” than Sonnet 5 and costs “up to 30% lower” per task, driven by reduced token consumption rather than a rate cut.

Why cost per task is not actually half

Token rates and task costs are different numbers. At maximum effort on Terminal-Bench 4.0, mixed-news reports Sonnet 5.5 consumes roughly 193,000 tokens per task at about $7.60. Opus 5.5 uses roughly 120,000.

So the cheaper model thinks longer to get there. A 50% discount on tokens against a 60% increase in tokens consumed is not a 50% saving — it is a much thinner one, and it only shows up in production bills.

The efficiency gain against Sonnet 5 is real, though. Anthropic cited hedge fund Balyasny measuring 121,000 tokens per answer on Sonnet 5.5 versus 497,000 on Sonnet 5 — a 76% reduction on the same workload.

How does Claude Sonnet 5.5 score on benchmarks?

Anthropic published gains across six suites. The pattern: Sonnet 5.5 lands at or just above Opus 5.5 on scoped coding and agent work, and just below it on open-ended professional tasks. Here is what Anthropic reported.

Benchmark Sonnet 5.5 Comparison
Terminal-Bench 4.0 70.6% Opus 5.5: 66.4% · GPT-6 Astra: 59.1% · Sonnet 5: 10.3%
GDPval-AA v2.1 1844 Opus 5.5: 1846 · GPT-6 Sol: 1487
OSWorld 2.1 (computer use) 80.1% —
CursorBench 4.0 55.5% —
FrontierCode 1.1 52.1% (xhigh) / 46.2% (max) —
Humanity’s Last Exam 64.5% (with tools) —

The GDPval-AA line is the honest one. 1844 against Opus 5.5’s 1846 is a tie, not a win, and Anthropic did not dress it up.

What independent testing shows

Artificial Analysis ran Terminal-Bench 4.0 separately and got lower numbers for both models: 63.6% for Sonnet 5.5 and 59.6% for Opus 5.5, per Decrypt.

The 4-point gap between the two models held. The absolute scores dropped 7 points. Both facts matter — the ranking survives third-party testing, the marketing number does not.

Anthropic was unusually direct about the limits of its own table: “Benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.”

The company’s positioning statement is narrower still. Sonnet 5.5, Anthropic says, “is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets” — the same surface it built out when it launched Claude Docs and Slides in mid-September.

Why did Anthropic put Opus-grade cyber limits on a mid-tier model?

Because the model’s cyber capabilities reached Opus 5 levels. The Next Web reports Sonnet 5.5 carries “restrictions of the kind previously reserved for its most capable models,” and that higher-risk cyber requests “will visibly fall back to Sonnet 5.”

That is a first for the Sonnet line, and it is a capability disclosure disguised as a safety note. Anthropic is saying its cheap model is now dangerous enough to need the expensive model’s leash.

There is a commercial cost to that. A developer on a legitimate security workload can now be silently downgraded to a model scoring 10.3% on Terminal-Bench 4.0 instead of 70.6%. Anthropic has not published how often that routing fires.

The anti-distillation defense is new too

Sonnet 5.5 is the first Sonnet shipping with classifiers that block reasoning extraction, with preserved thinking cryptographically tied to the account that produced it.

The motivation is documented. The Next Web reports that in August 2026, researchers decoded 315,320 thinking blocks from 6,708 public agent traces and recovered 62 API keys, 33 passwords, and seven private keys.

That is a supply-chain problem, not a model problem. Reasoning traces leak whatever the agent touched, and agent traces get posted publicly by default in a lot of tooling.

Regulation is pushing the same direction. Article 55 of the EU AI Act obliges providers of systemic-risk general-purpose models to protect infrastructure and prevent serious incidents — duties in force since August 2025, with enforcement powers arriving in August 2026.

Who wins and loses financially?

The clearest winner is any buyer currently paying Opus 5.5 rates for scoped coding work. The clearest loser is Opus 5.5 itself, which Anthropic just undercut with its own product at half the token price on its own flagship benchmark.

  • Winners: coding-tool vendors and agent startups whose gross margin tracks tokens per resolved ticket. A 76% token reduction against Sonnet 5, as measured by Balyasny, drops straight to the bottom line.
  • Winners: AWS, Google Cloud, and Azure, which get day-one distribution and keep the compute spend inside existing commitments.
  • Losers: Opus 5.5 revenue on scoped work. Anthropic’s own table gives buyers a reason to downgrade.
  • Losers: vendors reselling premium reasoning with no cost advantage of their own, now competing against a $2 model that ties the flagship on professional-work scores.
  • Exposed: anyone running unattended security automation, who now faces silent fallback to a far weaker model with no published trigger rate.

The cannibalization problem is deliberate

Anthropic could have priced Sonnet 5.5 between Sonnet 5 and Opus 5.5. It did not. Holding at $2 while claiming the Terminal-Bench lead is a decision to trade Opus margin for volume and for switching costs at rival labs.

The timing fits a company buying capacity aggressively. Anthropic signed an $11.6 billion cloud deal with Akamai in late September, attached to a 5% warrant. Cheap inference is how you fill contracted capacity.

It also lands two days before OpenAI’s DevDay counterpunch. GPT-6.1 Sol shipped September 29 at the same $2 entry price, and OpenAI put Opus 5.5 in three of its five benchmark comparisons. Both labs are now competing to be the cheapest model that scores like an expensive one.

What is the skeptical read on Claude Sonnet 5.5?

The Terminal-Bench comparison is not apples to apples. Anthropic’s 70.6% for Sonnet 5.5 and 66.4% for Opus 5.5 were not produced at identical settings — the Opus figure is an xhigh-effort result, per mixed-news. A 4-point lead across different effort configurations is a soft claim.

Second, the Sonnet 5 baseline of 10.3% deserves scrutiny. A jump from 10.3% to 70.6% inside one point release is a seven-fold improvement on a benchmark that did not change versions. That usually indicates the older model was poorly configured for the harness, not that the new one is seven times better.

Third, the “up to 30% lower cost per task” figure is a ceiling with no floor disclosed. Against the ~193,000 tokens Sonnet 5.5 burns at max effort, the saving depends entirely on which effort level a workload actually needs.

Fourth, every safety number is missing. Anthropic disclosed that high-risk cyber requests fall back to Sonnet 5 but published no false-positive rate. Buyers are being asked to accept an unquantified availability risk on exactly the workloads where a silent downgrade does the most damage.

Frequently asked questions

How much does Claude Sonnet 5.5 cost?

$2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50 per million. That is the same as Sonnet 5 and half of Opus 5.5’s $4/$20.

Is Claude Sonnet 5.5 better than Opus 5.5?

On Terminal-Bench 4.0, Anthropic reports 70.6% versus 66.4%. On GDPval-AA v2.1 it is effectively a tie at 1844 versus 1846. Anthropic itself says Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.”

What is the Claude Sonnet 5.5 context window?

1M input tokens with a 128K maximum output, and a June 2026 knowledge cutoff.

Where can I use Claude Sonnet 5.5?

On the Claude Platform as claude-sonnet-5-5, and through AWS, Google Cloud, and Microsoft Azure. The weights are closed.

What are the new cyber restrictions?

Sonnet 5.5’s cyber capabilities are comparable to Opus 5, so it carries restrictions previously reserved for flagship models. Higher-risk cyber requests visibly fall back to Sonnet 5.

Why does Claude Sonnet 5.5 block reasoning extraction?

It is the first Sonnet with classifiers that stop reasoning extraction, with preserved thinking tied to the producing account. Researchers decoded 315,320 thinking blocks from 6,708 public agent traces in August 2026, recovering 62 API keys, 33 passwords, and seven private keys.

Do independent benchmarks match Anthropic’s numbers?

Directionally, not absolutely. Artificial Analysis measured 63.6% for Sonnet 5.5 and 59.6% for Opus 5.5 on Terminal-Bench 4.0 — both about 7 points below Anthropic’s figures, with the ranking unchanged.

The bottom line

Claude Sonnet 5.5 is the most aggressive pricing move Anthropic has made this year, and it is aimed inward as much as outward. Holding at $2 per million input tokens while claiming a 4-point Terminal-Bench lead over a model that costs twice as much is a deliberate decision to cannibalize Opus 5.5 revenue rather than let OpenAI take it.

The verdict: migrate scoped work, keep Opus for judgment. Anthropic’s own guidance says exactly that, and the GDPval-AA tie at 1844 versus 1846 supports it better than the Terminal-Bench headline does. Check your token consumption before you assume a 50% rate cut is a 50% bill cut — at max effort it is not close.

The part of this release that will still matter in six months is the security posture, not the score. A mid-tier model carrying flagship cyber restrictions and account-bound reasoning traces is Anthropic telling the market that capability diffusion has outrun its old tiering, and that reasoning traces are now a credential-leak surface with 315,320 documented examples behind it.

The unanswered question is the one Anthropic did not price: how often the cyber fallback fires on legitimate work. Until that number exists, the discount comes with an availability risk nobody can model.

Sources

Comments

One response to “Claude Sonnet 5.5 Hits 70.6% on Terminal-Bench With Opus-Grade Limits”

  1. […] normally arrives within days of an API launch. Here there is no API to reproduce them on. Our coverage of Claude Sonnet 5.5’s Terminal-Bench result had third-party runs to compare against within hours. Argon has […]

Leave a Reply

Discover more from Wealth Engine

Subscribe now to keep reading and get access to the full archive.

Continue reading