In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.
OpenAI and Anthropic's lead is vanishingly small at this point.
OpenAI and Anthropic's lead is vanishingly small at this point.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
Really then what is the point of the Cyber Verification Program?
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
Does it also take 2 seconds of critical thinking to realize that the models that are covered should be accurately named by the people making the decisions?
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
It can't be unlimited because you can spawn parallel streams.
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
Unlimited but account-level throttled tps (more parallel streams means more throttling across all) is OK IMO, as long as it isn't too crazy. The thing that makes subscriptions really suck is having to watch the quotas, because prompt cache maintenance.
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
How well it is documented or known that how they use the passport information and so on. Current blocker for EU citizen is to share that data for AI company…
Between OpenCode and Openrouter which one would you suggest? Sometimes I have pure grunt work to be done on non-sensitive data for which I want to use Chinese models. For example, tasks like extracting something from publicly available large pdf files.
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
It was really verbose and pedantic. I'm sure that made it more thorough. But compared to Fable (which it wasn't much cheaper than) where you could get the same rigour and more with a lot more concision, it was a tough sell. 5.5 is a lot cheaper and seems a lot better balanced.
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
In short, seems to describe vibe-coding to me?
What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
It sorta feels to me like extrapolating from "the internet has all the knowledge for free" to "we won't need tradespeople anymore"
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
I wonder how others are coping with it aside from denial and cynicism.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
Exactly. I'm saying that Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase.
Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
I think going for more understanding is the way you need less understanding. The more solid your core understanding of your codebase is the less you need to know the details, the less missunderstandings the less iterations needed, the less mental capacity consumed
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.
A lot of old-school software engineering is about how to deal with this reality.
No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
There is more software to be written than there were programmers so lots of people do indeed want more software, for example small tools and one off projects that aren't worthwhile to make pre LLM.
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
I think that's what a future dev team is going to look like.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Plan longer chains of work / higher level goals that can be broken down into multiple chains of work. This will allow you to automate more work units to be worked on.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
I think this is valid now, but not guaranteed to be valid forever. For engineers, there was a period where more checks, more tests, more auto code reviews improved results quite a bit. People were consuming tokens like crazy (including me). Then things improved via better effort/thinking levels, where you could see repeated code reviews plateaued, so now people don't really do that quite as much.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
In our testing, it costs up to 30% less per task than its predecessor.
Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
I did like GLM 5.3 Flash but it's just way too often I'd run it on some long running task and come back to it repeating the same tokens or tool calls endlessly, just doing nothing. It wasn't unusable, but I couldn't trust it. That's really frustrating and I think new models have to do better not just on benchmark scores but general reliability and user experience as well.
At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
It depends entirely on its capabilities. If it is significantly smarter than Luna, which frankly is quite likely, then a lot of people won't mind paying more.
Or if its luna at 1000 tok/sec. Speed is what most of my peers care most about these days since less intelligent models can do most grunt work just fine.
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
I believe you meant to cite the Opus 5.5 System Card which states:
Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
I saw that with corporate software too. What works is creating a skill with the task steps, it fades its initial reasoning. (I am not talking about observer safe guards, but the safety RTL).
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
Open AI is terrible at diversifying their offering. We use Anthropic models via AWS bedrock where inference is deployed in EU regions due to strict compliance reasons. We've been wanting to try out the new Open AI models for ages, but they don't offer the models in any EU region. Open AI is losing a ton of money they can milk from corporations because of that.
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
Just added! I had Opus 5.5 look at them, not a perfect way to score them but it's close-ish -- Best would be a ELO, where people play both and rank a winner, but I don't know if people want to bother doing that.
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Score Tokens Reason Cost
Kimi K3 Max 44 48k 32k $2.00
Half reason 44 32k ? 16k ? ?
Opus Med 51 26k 12k $1.34
Opus High 54 36k 18k $1.82
Opus Max 58 119k 84k $5.98
Sonnet Med 41 ? ? $0.59
Sonnet High 47 ? ? $1.08
Sonnet Max 56 193k 142k $7.60
Medium is Anthropic's default.
Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.
I’m not sure whether that’s a feature or a bug at this point though.
I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
On the contrary, subagents save context overall, when the task is sufficiently large.
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.
Agree with afro88. Opus 5.5 as competent as Fable 5.1 on complex adversarial review of material and equations planted with errors. I still use both for a bit of variety.
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.
On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
Here's my purely academic initial impression based on only what they have released from the blog and the system card:
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.
It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.
The pricing model confuses me though (I presume by design, Hanlon be damned).
For what it's worth, the subscription plans of Anthropic are also 20x cheaper per token than the API prices. The API prices seem to have a very healthy margin.
Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
American AI lost the game already, people just can't see it.
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
The difference is that Microsoft did that while relying on the extremely load bearing windows ecosystem. These AI companies have no equivalent lock in, nothing even close to it honestly.
I can guarantee you 80% of people will call any llm "a chatgpt", most have never heard of claude, even less of opus, or sonnet, "deepseek" probably reminds them of a brand of toothpaste or something like that, "GLM" might make them think of the new mercedes SUV perhaps. 99.9% will never self host, nor send a single sent to a chinese model provider.
I don't know a single company using deepseek internally in any capacity, and I have friends in a lot of tech/tech heavy companies, virtually all of them use claude, the lucky ones get cursor with claude/chatgpt/grok.
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
You're a decent sized company and wants to manage the SW + security on all your employee's PCs. They need to be able to update/remote SW on your machine remotely, see your settings, etc.
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
The usage numbers tell a different story. Not only is the great majority of AI users on American frontier providers, they are also willing to pay the premium price for the premium product. It's not like tech is oblivious to the Chinese models. It's that the industry is aware they are always 6 months to a year behind.
You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.
You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.
Say that to Luna's face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing "cost" as the only factor.
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
my feeling is you're most likely wasting money using Opus for implementation. The plan and task breakdown should be specific enough that Sonnet can implement without you noticing a difference.
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
Wouldn't make sense to use anything below 5.5 from Anthropic at the moment. But pretty sure this is just an awkward transition phase of at most a week or two until Fable 5.5 is out.
It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.
I know this isn't a model thing, but why do all the labs blow at product outside of models?
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.
Well, always watch also the number of tokens used (and price). Intelligence scales with tokens so you might make Luna as smart as Sol with crazy amount of them :)
I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?
Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
unwittingly said 'yaaay' when I saw the Sonnet 5.5 entry. lol I love Sonnet so much. Loved it since 3.5 and never really liked Opus even when I tried to use it for technical work. Once we had two back to back anthropic releases neither of which was a Sonnet upgrade from 4.5 (4.6?) and I was getting kinda sad that they're considering discontinuing it. These are weird reactions I'm having to these tools even when I mostly use them for technical work considering I prefer talking to GPT and Gemini is just a blunt, very powerful hammer.
In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.
OpenAI and Anthropic's lead is vanishingly small at this point.
Yeah, I did kind of feel like the step down from Opus 5.5 was so large as to never make it appealing.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
Your takeaway from "Sonnet 5.5 matches and sometimes exceeds the SoTA worldwide" is "their lead is vanishingly small"...?
This is crazy, what is the point of all these equivalent models?
Those are just 3 particular technical benchmarks. Presumably Opus is a larger model and has greater world knowledge.
Price going down on each release
Yup. Recursive self improvement presented in hard numbers.
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
This is bollocks. Their safeguards are shit.
Yup. As far as I can tell, the Cyber Verification Program does absolutely nothing.
lol I got flagged for using the word fuzz, not even in a security context (it was a parser so security adjacent but still).
Parsers are security adjacent until they aren't.
Very true.
Really then what is the point of the Cyber Verification Program?
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
Can confirm. Its worthless
Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
it did exactly what a human would do - "I can't look at this shit"
Yuh—- no. 4.8 can handle this bullocks lol
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
Bingo. And to prove your point, after switching to cheap open models (I think Qwen?) it did indeed find a bug in my WAL implementation.
This is the way.
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
You should read the actual docs for the CVP. At the very top:
https://support.claude.com/en/articles/14604842-real-time-cy...
You obviously should not expect the CVP to cover this model either.
He did mention Sonnet...
It takes about 2 seconds of critical thinking to realize that if Opus 5.5 isn't covered yet, neither will a model that just launched an hour ago.
Does it also take 2 seconds of critical thinking to realize that the models that are covered should be accurately named by the people making the decisions?
Sure, the documentation should be up to date but it's obviously not? That doesn't excuse not thinking critically.
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
Are you sure they aren't already doing that for certain organizations?
https://support.claude.com/en/articles/14604842-real-time-cy... says that Cyber Verification Program doesn't apply to Opus 5.5 yet, they hope to roll it out for 5.5 "soon"
Give them a break, they got into a War with Trump over this..it will come soon enough
Funnily enough the Fable safeguards are the worst and testing Sonnet 5.5 didn't trigger them as much as it even did for Opus on my benchmarks[1].
1 - https://bench.killswitch-lang.org
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
cutting edge fable is for them not you and they're not going to share the metrics until they give you access.
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
Fable 5.5 probably drops soon so it would just be confusing.
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
Waiting for the pelicans
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
Can't wait for it to get to the point where it's like an internet subscription: unlimited tokens 24/7/365 at a low, fixed monthly price.
It can't be unlimited because you can spawn parallel streams.
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
Unlimited but account-level throttled tps (more parallel streams means more throttling across all) is OK IMO, as long as it isn't too crazy. The thing that makes subscriptions really suck is having to watch the quotas, because prompt cache maintenance.
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Daybreak Blue is not bad and the bar to get into OpenAI's program is reasonable.
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
Yeah, I did say the bar is low :)
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
How well it is documented or known that how they use the passport information and so on. Current blocker for EU citizen is to share that data for AI company…
what is the easyest way to use the chinese models and which harness does work with them well?
OpenCode harness with their subscription would be my recommendation.
Between OpenCode and Openrouter which one would you suggest? Sometimes I have pure grunt work to be done on non-sensitive data for which I want to use Chinese models. For example, tasks like extracting something from publicly available large pdf files.
Opencode Go, with the Deepseek 4.1 Flash model feels like a bottomless pit, which is great for grunt work.
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
5 was definitely bad. 5.5 seems a lot better so far. But still not close to Fable in terms of quality.
what was the problem with 5? to me, it seemed like the first Opus since 4.6 that was a real step up in intelligence without any obvious downsides
It was really verbose and pedantic. I'm sure that made it more thorough. But compared to Fable (which it wasn't much cheaper than) where you could get the same rigour and more with a lot more concision, it was a tough sell. 5.5 is a lot cheaper and seems a lot better balanced.
Read its output, and especially comments
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
It sorta feels to me like extrapolating from "the internet has all the knowledge for free" to "we won't need tradespeople anymore"
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
This made me a bit emotional. Thank you for the beautiful response.
You're welcome, and I'm glad it helped. Now more than ever, we have to try to connect with actual humans and take care of each other.
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?
The proof of the pudding.
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
Yes. Our career is over, as is our economy. Soooo... FYI :(
Hackernews' neuroticism remains undefeated
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
Useful for API requests, when using AI in the product rather than to build the product.
Most business/enterprise accounts also have to pay API rates.
Exactly. I'm saying that Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
I think going for more understanding is the way you need less understanding. The more solid your core understanding of your codebase is the less you need to know the details, the less missunderstandings the less iterations needed, the less mental capacity consumed
When they inevitably drop allocation after post-launch hype dies down.
Cache read is the same as Opus as well where most agentic workflow cost comes from.
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium, so what's the point of sonnet?
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
Or just use a better harness.
You can also create custom agents with defined effort levels and use those.
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
Because humans are currently the bottleneck.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.
A lot of old-school software engineering is about how to deal with this reality.
No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
There is more software to be written than there were programmers so lots of people do indeed want more software, for example small tools and one off projects that aren't worthwhile to make pre LLM.
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.
[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
lol yeah, our review bot does a cost based analysis and pauses itself until you re-resume if it goes over a threshold.
I think that's what a future dev team is going to look like.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
Plan longer chains of work / higher level goals that can be broken down into multiple chains of work. This will allow you to automate more work units to be worked on.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
I think this is valid now, but not guaranteed to be valid forever. For engineers, there was a period where more checks, more tests, more auto code reviews improved results quite a bit. People were consuming tokens like crazy (including me). Then things improved via better effort/thinking levels, where you could see repeated code reviews plateaued, so now people don't really do that quite as much.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
I find that I can do 4 to 7 sessions in parallel, and still review everything in depth and co-design.
I mostly use Fable though, Opus only via sub-agents.
Subagents.
Try telling an agent to go through your backlog.
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
Agreed. I’ve found that Anthropic [dot] com at least honors “reduce motion” accessibility settings, and that makes their site a bit more useable.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
From TFA:
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
I read they identified a training bug and were going to push out an updated release to fix the looping. I really like it overall.
GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.
I did like GLM 5.3 Flash but it's just way too often I'd run it on some long running task and come back to it repeating the same tokens or tool calls endlessly, just doing nothing. It wasn't unusable, but I couldn't trust it. That's really frustrating and I think new models have to do better not just on benchmark scores but general reliability and user experience as well.
At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
It's been out for an hour and you've already concluded this?
Would Jev-type functionality be a reason to dust off Haiku?
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
t/s maybe? IDK, because their token speed comparison was against Sonnet 5.
At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks worse at xhigh.
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
i'd love to see them re-enter that space but given haiku 5 never happened I wouldn't bet on it
i think they see what openai charges for luna and just don't want to try and compete
But they literally stated that they would release Sonnet 5.5 and Haiku 5.5 after Opus 5.5 was released
Haiku 5.5 is DOA without a massive price cut. Luna is literally 10x cheaper at current pricing
It depends entirely on its capabilities. If it is significantly smarter than Luna, which frankly is quite likely, then a lot of people won't mind paying more.
Or if its luna at 1000 tok/sec. Speed is what most of my peers care most about these days since less intelligent models can do most grunt work just fine.
they've already said in both the opus 5.5 and sonnet 5.5 blog posts that haiku 5.5 is coming
They did mention in the Opus 5.5 announcement blogpost that Sonnet and Haiku 5.5 will follow soon.
It literally does not?
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
at that point, you can switch to dirt cheap open models
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
I still can’t find a place for Sonnet models, I never have.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Claude, is that you?
you're absolutely right to push back
Damn, HN commenters starting to talk in claudisms now
this is human writing...
this is claude writing...
corporate needs you to find the difference
Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
Well presumably now it’ll fall back to Sonnet 5.5 lol
That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
Fallback was usually Opus 4.8
I believe you meant to cite the Opus 5.5 System Card which states:
I cannot find a Sonnet 5.5 system card.
It was linked in another HN post: https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c85...
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
this "feature" is one of the primary that caused me to cancel and move to exclusively open weight based systems
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.
Anthropic made it that way, and I'd say the lower score is accurate.
So I suppose the easy fix for Anthropic would be to have Opus 5.5 now fall back to Sonnet 5.5, right?
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
Time to switch team to Claude from OpenAI again.
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
Cost are bigger than Opus 5.5 for that effort
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
Important to note that lower model + higher reasoning gives a different (not higher) quality of response than higher model + lower reasoning.
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Subagents.
So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
doom marketing at its finest
Not really, they never released Mythos. And they never said Fable was dangerous. They've been very consistent
Well, it's performance "surface" (is there a better term for this?) is probably very narrow compared to Fable :)
Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
Mythos and Fable are the same llm. Fable has an extra tool that's in front of it that decides to accept the promp or not.
Hence why it is no longer dangerous, yes
If you consider security through obscurity a safe route, yes
How is this security through obscurity?
That term is about hiding a system's design in order to secure something, rather than having robust defences.
And Sonnet has a similar classifier in front:
You have to actually spend some effort reading the article they published to answer your own question.
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
I mean, they worked really hard for this. Back in February everybody loved them.
I think all the positive people have just stopped commenting.
The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.
Astra is still the uncontested #1 code generator.
Astra is amazing, I love it.
Yeah, especially coupled with Opus for alternative reviews. A massive token burn though.
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.
In xhigh effort it is a lot cheaper and possibly lot less impressive?
very annoyed they aren't showing fable on the graph.
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
I saw that with corporate software too. What works is creating a skill with the task steps, it fades its initial reasoning. (I am not talking about observer safe guards, but the safety RTL).
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
Open AI is terrible at diversifying their offering. We use Anthropic models via AWS bedrock where inference is deployed in EU regions due to strict compliance reasons. We've been wanting to try out the new Open AI models for ages, but they don't offer the models in any EU region. Open AI is losing a ton of money they can milk from corporations because of that.
In other news
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
Low and medium both used 0 thinking tokens.
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
If it was trained on HN, there would be a 60% chance of it just saying "I'm so tired of this request, can we please move on"
I'd say:
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
That's max output tokens per response limit, separate from context length
It's the output token limit, which has been 128,000 for Claude models for quite a while note
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
I think this is a bug. I've not seen this problem from any of the other frontier models.
Do other models put a hard cap on the output tokens it can generate?
Yes, the OpenAI GPT-6 Astra limit is 128,000 as well: https://developers.openai.com/api/docs/models/gpt-6-astra
Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
I would also consider this a bug. I think ajy reasonable consumer would.
I think the next models will be benchmaxxing on the Pelican benchmark tbh
wow nobody but you has ever thought of this and certainly simonw has never addressed this
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
Yes! The point being that up until yesterday, every model struggled with this, and now they don't.
Your welcome!
"this" being recreating Pacman specifically, or games?
I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
Other models fail at oneshot creation of similar games?
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
Thanks!
Thank you for doing this. It is very helpful not just for capabilities but also for costs.
I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
Just added! I had Opus 5.5 look at them, not a perfect way to score them but it's close-ish -- Best would be a ELO, where people play both and rank a winner, but I don't know if people want to bother doing that.
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
oh! How about pengo, dig dug, & defender?
I'm afraid those will be too easy. I'm not sure what the next game should be...
Interesting. Sonnet 5 was horrible, and Opus 5 was unplayable, but both Sonnet 5.5 and Opus 5.5 were about as close to the real thing.
GPT models really have no taste huh.
“Pelicans are solved.”
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Medium is Anthropic's default.
Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?
I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
that pelican one-pedaling
This is great news because it means the model has not been benchmaxxed on stupid metrics.
PS: the next human that brings up pelicans on bicycles should try to draw them.
Does anyone really still care about these pelicans?
Any model release it’s the top comment, I do not understand why.
Karma farming by parent commenter and HNers’ tendency to upvote low quality content (not dissimilar to other social media networks)
Is this low quality content relative to most HN comments?
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.
In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
You can click the little [-] icon next to the post to collapse the entire sub-thread. I do that all the time.
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.
I’m not sure whether that’s a feature or a bug at this point though.
what's most surprising is the difference between high and xhigh
I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
Next will be haiku 5.5, surpassing opus 4.8
Have anyone tried a workflow that:
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
You might as well use Opus for everything there.
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?
Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
On the contrary, subagents save context overall, when the task is sufficiently large.
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
If you use subagents your main agent won't need to compact as often, with the loss of information that entails.
Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
Using advisors doesn't break anything
Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.
Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.
Agree with afro88. Opus 5.5 as competent as Fable 5.1 on complex adversarial review of material and equations planted with errors. I still use both for a bit of variety.
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.
On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
Here's my purely academic initial impression based on only what they have released from the blog and the system card:
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.
https://bench.killswitch-lang.org/
Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.
Oh yes. I think you might get a lot for what you pay with Sonnet 5.5.
It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
Obviously not as "intelligent" but almost 10x cheaper
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
You don't have to pay for cache write if prompt isn't part of conversation.
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.
The pricing model confuses me though (I presume by design, Hanlon be damned).
For what it's worth, the subscription plans of Anthropic are also 20x cheaper per token than the API prices. The API prices seem to have a very healthy margin.
What model are you using ?
Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
What harness are you using with those?
opencode
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
The difference is that Microsoft did that while relying on the extremely load bearing windows ecosystem. These AI companies have no equivalent lock in, nothing even close to it honestly.
I can guarantee you 80% of people will call any llm "a chatgpt", most have never heard of claude, even less of opus, or sonnet, "deepseek" probably reminds them of a brand of toothpaste or something like that, "GLM" might make them think of the new mercedes SUV perhaps. 99.9% will never self host, nor send a single sent to a chinese model provider.
They are not the ones spending API money.
I don't know a single company using deepseek internally in any capacity, and I have friends in a lot of tech/tech heavy companies, virtually all of them use claude, the lucky ones get cursor with claude/chatgpt/grok.
Most businesses I know switched to Google Sheets or Google Docs.
Never heard about anyone using google sheets and docs for messaging, email, presentations, etc.
Gmail, Meet and Presentations do all of those
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
You’ve never heard of anyone using gmail for email?
You're a decent sized company and wants to manage the SW + security on all your employee's PCs. They need to be able to update/remote SW on your machine remotely, see your settings, etc.
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
cfengine is 33 years old
From my experience GLM 5.3 is at most 6 months away from the frontier models, and good enough for most tasks already
Even 5.2 is doing really well in comparison here: https://labs.scale.com/leaderboard/sweatlas-refactoring
Europe didn't even start in the race. Europe is the biggest losers of all it seems. What a shame.
People dont want iPhone 14s in 2026. People want the latest and greatest. Chinese companies desperately trying to get western usage of their models
They would if those phones were 899/10=89.9 bucks.
How are these models so cheap?
What are they censoring that matters to programmers?
Does AI-style writing start bleeding into the comments, or is HN now also full of bots like reddit?
This ignores two things.
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
The usage numbers tell a different story. Not only is the great majority of AI users on American frontier providers, they are also willing to pay the premium price for the premium product. It's not like tech is oblivious to the Chinese models. It's that the industry is aware they are always 6 months to a year behind.
You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.
You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.
Say that to Luna's face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing "cost" as the only factor.
If it's for personal use, is there a reason you don't want the subscription?
Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model.
And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
my feeling is you're most likely wasting money using Opus for implementation. The plan and task breakdown should be specific enough that Sonnet can implement without you noticing a difference.
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
Benchmarks were always barely useful to begin with. Gotta actually try the model.
Wouldn't make sense to use anything below 5.5 from Anthropic at the moment. But pretty sure this is just an awkward transition phase of at most a week or two until Fable 5.5 is out.
Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.
This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.
artificialanalysis.ai benchmarks are [here](https://artificialanalysis.ai/articles/claude-sonnet-5-5). Anthropic is back at spot 1, 2, 3 and 5. Impressive even if these benchmarks are problematic in many ways.
It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.
Took forever to load this garbage site in both FF and Chrome. Sometimes I wonder if these corps really are corps trying to sell a product.
I know this isn't a model thing, but why do all the labs blow at product outside of models?
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.
Sticking with OPUS 5.5 for resume/STAR generation for me. Tried Sonnet 5.5 but worse than OPUS for thinking for sure, less error/inconsistency check.
I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md
It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.
---
For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.
I use AI the way I want, you don't force me not to use SOTA for this
Sonnet 5.5 is way better than GPT 6 Sol. Does that even make sense?
Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.
On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...
Well, always watch also the number of tokens used (and price). Intelligence scales with tokens so you might make Luna as smart as Sol with crazy amount of them :)
Also, these are benchmarks...
The documentation mentions error code "frontier_llm": The request could assist the development of competing AI models.
I'm very curious how do they know what requests could assist competing AI models.
I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?
Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose
AI companies these days releasing models every week like Netflix episodes.
the (Ai) factory must grow!
Maybe they could put Fable onto creating a website that doesn't use 80% of the GPU on an Apple M2.
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
These benchmark results keep getting more questionable without error bars.
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
unwittingly said 'yaaay' when I saw the Sonnet 5.5 entry. lol I love Sonnet so much. Loved it since 3.5 and never really liked Opus even when I tried to use it for technical work. Once we had two back to back anthropic releases neither of which was a Sonnet upgrade from 4.5 (4.6?) and I was getting kinda sad that they're considering discontinuing it. These are weird reactions I'm having to these tools even when I mostly use them for technical work considering I prefer talking to GPT and Gemini is just a blunt, very powerful hammer.
We're still on Sonnet 4.6 as we found Sonnet 5 to perform worse across all of our evals, especially against time.
username does NOT check out. so confused.