Sonnet 5.5

457 points - today at 5:58 PM

Source

Comments

simonw today at 6:44 PM
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Here's how the thinking effort levels compare:

  low
  27 input, 1,623 output, thinking_tokens: 0
  1.6284
  Duration: 10138ms (10s)
  
  medium
  27 input, 1,796 output, thinking_tokens: 0
  1.7914 cents
  Duration: 11266ms (11s)

  high
  27 input, 2,334 output, thinking_tokens: 745
  2.3394 cents
  Duration: 17376ms (17s)

  xhigh
  27 input, 5,730 output, thinking_tokens: 2535
  5.7354 cents
  Duration: 41882ms (41s)

  max (failed to return response)
  27 input, 128,000 output, thinking_tokens: 128000
  $1.28
  Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.
Sol- today at 6:13 PM
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

abejora today at 6:17 PM
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card

wongarsu today at 6:11 PM
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5

Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models

MisterMunchkin today at 7:16 PM
It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

And my job won’t even pay for Claude now because it’s so ruinously expensive.

wkcheng today at 6:15 PM
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?

Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.

Jcampuzano2 today at 6:25 PM
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.

Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.

Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.

robertclaus today at 9:19 PM
These benchmark results keep getting more questionable without error bars.
ghoshbishakh today at 6:27 PM
So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
johnmlussier today at 6:08 PM
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.

This is bollocks. Their safeguards are shit.

gregwebs today at 6:51 PM
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.

From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.

OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.

guilhermeasper today at 8:41 PM
AI companies these days releasing models every week like Netflix episodes.
heyjstn today at 6:49 PM
Have anyone tried a workflow that:

- Fable 5.1 for planning/adversarial reviewer

- Opus 5.5 for well-scoped tasks break down

- Sonnet 5.5 for these well-scoped tasks implementation

I think the blocker might be how efficient the context is compacted and sending around between these agents

nanook today at 7:38 PM
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
avree today at 6:14 PM
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
yapfrog today at 6:29 PM
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
alansaber today at 6:31 PM
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
sergdigon today at 9:38 PM
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
mchusma today at 8:55 PM
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
dom96 today at 7:05 PM
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.

https://bench.killswitch-lang.org/

    Claude Sonnet 5    17.8%
    Claude Sonnet 5.5  7.4%
onlyrealcuzzo today at 6:14 PM
> In our testing, it costs up to 30% less per task than its predecessor.

> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.

They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.

Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.

I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).

For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.

This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.

Hopefully they release a Haiku that actually has a reason for existing.

AM1010101 today at 6:50 PM
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.

On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.

If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.

trvz today at 8:44 PM
Maybe they could put Fable onto creating a website that doesn't use 80% of the GPU on an Apple M2.
solenoid0937 today at 6:30 PM
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.

At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.

tombert today at 6:39 PM
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
sajithdilshan today at 7:26 PM
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
bastawhiz today at 8:18 PM
I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?

Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose

alasano today at 6:42 PM
I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
ChickeNES today at 6:20 PM
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
low_tech_punk today at 8:18 PM
The documentation mentions error code "frontier_llm": The request could assist the development of competing AI models.

I'm very curious how do they know what requests could assist competing AI models.

s3p today at 6:10 PM
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
croemer today at 6:40 PM
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
swingboy today at 7:07 PM
Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.
ghoshbishakh today at 6:33 PM
So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.

In xhigh effort it is a lot cheaper and possibly lot less impressive?

bayesianbot today at 6:20 PM
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
itishappy today at 8:05 PM
Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.
pookieinc today at 6:08 PM
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
s314 today at 6:36 PM
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
taurath today at 6:12 PM
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
a13o today at 7:47 PM
This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.
__jl__ today at 7:51 PM
artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.
mroche today at 7:22 PM
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
takerofnaps today at 6:11 PM
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
pavitheran today at 6:24 PM
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
square_usual today at 6:25 PM
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
rtuin today at 6:14 PM
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
ramish94 today at 5:59 PM
In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family

_fw today at 6:16 PM
I still can’t find a place for Sonnet models, I never have.

I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.

Give me the frontier, or give me the cheapest form of good enough.

limsungkee today at 6:33 PM
Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
jtrn today at 6:52 PM
Here's my purely academic initial impression based on only what they have released from the blog and the system card:

If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.

BUT

It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.

Some of the more interesting things I found from scanning the system card:

- It is the only model tested that shows no preference for rude or polite style.

- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.

- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).

- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.

- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.

- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.

- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).

Clinical behaviour:

Suicide and self-harm handling is reported as weaker in the API because it

It sometimes called a wish to die understandable.

It sometimes validated self-harm as functional.

It sometimes suggested harmful substitute behaviours.

As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."

Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be

yipinwong today at 8:10 PM
Sticking with OPUS 5.5 for resume/STAR generation for me. Tried Sonnet 5.5 but worse than OPUS for thinking for sure, less error/inconsistency check.

I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md

It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.

---

For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.

I use AI the way I want, you don't force me not to use SOTA for this

Alifatisk today at 6:44 PM
In other news

> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

iagocc today at 6:09 PM
Waiting for the pelicans
SeriousM today at 6:47 PM
Next will be haiku 5.5, surpassing opus 4.8
deleted today at 6:08 PM
laurenz-bauer today at 7:10 PM
Oh yes. I think you might get a lot for what you pay with Sonnet 5.5.
jdw64 today at 8:17 PM
Sonnet 5.5 is way better than GPT 6 Sol. Does that even make sense?

Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.

On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...

BoorishBears today at 8:03 PM
I know this isn't a model thing, but why do all the labs blow at product outside of models?

Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.

How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"

AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.

AtNightWeCode today at 7:58 PM
Took forever to load this garbage site in both FF and Chrome. Sometimes I wonder if these corps really are corps trying to sell a product.
system2 today at 7:44 PM
Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.
ahriad today at 6:21 PM
Time to switch team to Claude from OpenAI again.
dude250711 today at 6:56 PM
It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".
enraged_camel today at 6:43 PM
Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.

If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.

OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.

It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...

dack today at 6:34 PM
very annoyed they aren't showing fable on the graph.