In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.
OpenAI and Anthropic's lead is vanishingly small at this point.
OpenAI and Anthropic's lead is vanishingly small at this point.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
Really then what is the point of the Cyber Verification Program?
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase.
Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
In our testing, it costs up to 30% less per task than its predecessor.
Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
I believe you meant to cite the Opus 5.5 System Card which states:
Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.
OpenAI and Anthropic's lead is vanishingly small at this point.
Yeah, I did kind of feel like the step down from Opus 5.5 was so large as to never make it appealing.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
Your takeaway from "Sonnet 5.5 matches and sometimes exceeds the SoTA worldwide" is "their lead is vanishingly small"...?
This is crazy, what is the point of all these equivalent models?
Those are just 3 particular technical benchmarks. Presumably Opus is a larger model and has greater world knowledge.
Price going down on each release
Yup. Recursive self improvement presented in hard numbers.
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
This is bollocks. Their safeguards are shit.
Yup. As far as I can tell, the Cyber Verification Program does absolutely nothing.
lol I got flagged for using the word fuzz, not even in a security context (it was a parser so security adjacent but still).
Really then what is the point of the Cyber Verification Program?
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
You should read the actual docs for the CVP. At the very top:
https://support.claude.com/en/articles/14604842-real-time-cy...
You obviously should not expect the CVP to cover this model either.
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
cutting edge fable is for them not you and they're not going to share the metrics until they give you access.
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
Fable 5.5 probably drops soon so it would just be confusing.
Waiting for the pelicans
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Daybreak Blue is not bad and the bar to get into OpenAI's program is reasonable.
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
5 was definitely bad. 5.5 seems a lot better so far. But still not close to Fable in terms of quality.
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
Yes. Our career is over, as is our economy. Soooo... FYI :(
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
Useful for API requests, when using AI in the product rather than to build the product.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
t/s maybe? IDK, because their token speed comparison was against Sonnet 5.
At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks pointless at xhigh.
I still can’t find a place for Sonnet models, I never have.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
I believe you meant to cite the Opus 5.5 System Card which states:
I cannot find a Sonnet 5.5 system card.
It was linked in another HN post: https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c85...
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
Time to switch team to Claude from OpenAI again.
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.