Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Pentagon: Palantir AI Overreliance Led to Strike Killing 123 Iranian Children(gizmodo.com)
    discuss
  2. The Curious Power of Punctuation(newyorker.com)
    discuss
  3. Show HN: FreeCoffee – Self-hostable donation tool with crypto support(github.com/freecoffee-bio)
    discuss
  4. The Climate May Have Some Surprises in Store(nytimes.com)
    discuss
  5. Hacking group ShinyHunters claims it breached the FBI, stole agents' data(techcrunch.com)
    1comments
  6. Countries propose global oversight body to manage AI dangers(aljazeera.com)
    discuss
  7. Explaining to business people why building software is still hard(manager.dev)
    discuss
  8. Null Diary: Confessions from the Machine(nulldiary.io)
    1comments
  9. SAML: A Fractal of Bad Design(trailofbits.com)
    discuss
  10. Let's Stop Debating the Gnu Imp Manipulation Program(aria.dog)
    discuss
  11. Ask HN: What's Cool Now in Software?
    discuss
  12. GPT-6 Sol (Max) Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    discuss
  13. GPT-6 Luna (Max) Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    discuss
  14. Ask HN: Are you are a coward or a traitor?
    1comments
  15. Piclang – Programming language purpose built for malware(github.com/zarkones)
    1comments
  16. Show HN: Lightdrift / Image search for AI agents through API and MCP(lightdrift.ai)
    discuss
  17. RegEx 101 for Data Pipelines(expanso.io)
    discuss
  18. Douglas Davis: The First Collaborative Sentence(whitney.org)
    discuss
  19. Goose Robot(twitter.com/robotworksgoose)
    discuss
  20. Morpion Solitaire(morpionsolitaire.com)
    discuss
  21. High-mileage electric cars are more reliable than petrol ones, study finds(theguardian.com)
    2comments
  22. A Call for Control of Frontier AI Models(government.nl)
    discuss
  23. Why banning AI in the classroom will not be enough(theconversation.com)
    1comments
  24. Show HN: Cookbook – a workspace for your team and agents(cookbook.team)
    1comments
  25. Does an open-weight decision model beat a hosted one? Jev vs. Laya(astgl.com)
    discuss
  26. Zero-downtime Linux kernel zero-day mitigation via eBPF and SECCOMP(github.com/mc493)
    discuss
  27. Show HN: a0flow — paid micro-APIs for AI agents, x402/USDC, no signup(a0flow.com)
    discuss
  28. International Conference on Functional Programming (ICFP) 2026 talks released(youtube.com)
    discuss
  29. What does this assembly code do?(pagetable.com)
    1comments
  30. AgentsView(agentsview.io)
    discuss

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

96 pointsby 2h agoartificialanalysis.ai
34 comments
1h agoHN ↗

That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.

1h agoHN ↗

Greed is not what's driving these prices, its cost. They considered very much in the red.

1h agoHN ↗

Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5

12m agoHN ↗

The UI/UX seems impressively bad. DeepSWE's cost curve has a better, more obvious way to sort by only the top level of reasoning to avoid 80% of the graph just being the same 3-5 models at their 8 different reasoning levels...

It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.

I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.

An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.

1h agoHN ↗

Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...

1h agoHN ↗

I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.

1h agoHN ↗

Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.

1h agoHN ↗

This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

1h agoHN ↗

How did you get the reasoning trace? Is it the actual one or the summarized one?

1h agoHN ↗

"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

47m agoHN ↗

Just want to say: you’re such a legend, please do not stop sharing your pelicans, it’s always fun to see how they change over the months :)

1h agoHN ↗

The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.

29m agoHN ↗

Of course it was exposed - not sure it's explicit or not. Why wouldn't HackerNews comments be part of the training data? And Simon's blog and the many discussions about Pelicans? It'd be hard to miss. Doesn't mean Anthropic has made this an explicit goal in training.

1h agoHN ↗

I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).

1h agoHN ↗

For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.

My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.

47m agoHN ↗

That was true for me four weeks ago, but 2-3 weeks ago Luna turned into drivel in essentially the same complexity of task. I feel it came back somewhat in recent days but does feel like it's being manipulated.

1h agoHN ↗

I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.

1h agoHN ↗

I've asked Opus 5 Max for what I thought were easy tasks at work to be completed. It always fails after reaching a tool limit.

I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.

1h agoHN ↗

I am begging you on my knees to please stop posting this cringe.

The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

1h agoHN ↗

That it's better in specific ways? What difference does it make when it came out? The benchmark results are not going to change unless they're messing with the model.

1h agoHN ↗

Yes, but crucially, in ways that are increasingly decoupled from any practical pattern of usage, considering that I wouldn't see how you can argue that Astra is worse than Opus 5.

One man's modus ponens is another's modus tollens I guess.

1h agoHN ↗

What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.

1h agoHN ↗

You forgot to include whatever you're proposing instead.

"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.

Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.

1h agoHN ↗

This index doesn't have "Astra" and "Opus 5". Every entry with corresponding data is a `(model, reasoning)` tuple.

So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.

32m agoHN ↗

I don't try to say that the parent commenter is right in any way, but the two models' "high" settings probably doesn't mean the same thing. So probably comparing only them is not useful.

1h agoHN ↗

so its more intelligence than fable?

can anyone help me?

1h agoHN ↗

Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?

1h agoHN ↗

All anthropic launches are like this. They just post it and don't particularly put out the PR sprint that OpenAI does with videos, livestreams or whatever.

(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)