Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Faster and local Jev like model for Mac(github.com/mizorewww)
    discuss
  2. Show HN: Botzilla – Web automation using visual scripting(botzilla.dev)
    discuss
  3. Show HN: Notes on Agentic AI – A text-first guide for practicing engineers(github.com/mrsachindixit)
    discuss
  4. Enjoy Every Sandwich(bradmontague.substack.com)
    discuss
  5. Shipping our game to twelve platforms on day one(m2h.nl)
    discuss
  6. Show HN: Last Internet Connection(github.com/rubinoslaw)
    discuss
  7. Harper Lee first edition novel found in Oxfam shop(bbc.com)
    discuss
  8. Meta Tests Muse AI Agent Calls That Are Made by Humans in a Call Center(404media.co)
    discuss
  9. Unreal Agent(github.com/unreallabsai)
    1comments
  10. Show HN: Parametric, low-poly assets for BIM/CAD(3dassetstudio.com)
    discuss
  11. Show HN: Dragonfly – an API client that understands your Express/Next.js routes(usedragonfly.xyz)
    discuss
  12. ADIE – Beyond Prediction. Toward Verifiable Decisions(robertbudai.github.io)
    discuss
  13. An update on how we confirm your age group on Discord(discord.com)
    discuss
  14. Linux at Long Last and GPLv3(obscura.com)
    1comments
  15. Denuvo sues discord user for breaking DRM to games that they didn't make [video](youtube.com)
    discuss
  16. I Changed My Mind About AI Risk(persuasion.community)
    discuss
  17. The Rise and Fall of Tech Worker Power(bostonreview.net)
    discuss
  18. Ask HN: Honey Pots in 'who wants to be hired'?
    discuss
  19. Betel nuts are a top mind-affecting substance. There's evidence of ancient use(npr.org)
    discuss
  20. LiveCrew – a persistent AI exec team that argues back (with receipts)(livecrew.tech)
    discuss
  21. Jev Creator: System One Models for Prod, Not God [video](youtube.com)
    discuss
  22. Show HN: RUUN – Tor Browser with a second, non-Tor mode in the same app(gitlab.com/varad_7)
    discuss
  23. Ask HN: Which is cheaper: ChatGPT tokens or a 2nd Pro 20x subscription?
    discuss
  24. Untangling Lifetimes – Arena Allocators in C(dgtlgrove.com)
    discuss
  25. Nano Banana 2.5: Faster, Better AI Images(nanobanana25ai.net)
    discuss
  26. Fast-CLI: CLI tool written in Zig for testing internet speed(github.com/mikkelam)
    discuss
  27. Sauna Jungle Expansion(saunajungle.com)
    discuss
  28. GPT-6 Sol and Luna(openai.com)
    59comments
  29. Cookie started its life as a plastic bottle(acs.org)
    discuss
  30. GPT-6 Sol(developers.openai.com)
    1comments

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

69 pointsby 1h agoartificialanalysis.ai
29 comments
53m agoHN ↗

That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.

51m agoHN ↗

Greed is not what's driving these prices, its cost. They considered very much in the red.

47m agoHN ↗

Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5

54m agoHN ↗

Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...

52m agoHN ↗

I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.

48m agoHN ↗

Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.

47m agoHN ↗

This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

45m agoHN ↗

How did you get the reasoning trace? Is it the actual one or the summarized one?

43m agoHN ↗

"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

33m agoHN ↗

The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.

41m agoHN ↗

I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).

40m agoHN ↗

For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.

My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.

39m agoHN ↗

I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.

19m agoHN ↗

I've asked Opus 5 Max for what I thought were easy tasks at work to be completed. It always fails after reaching a tool limit.

I asked Opus 5 High for the same task and requested it to minimize tool usage. It produced an answer in a few minutes that I was deploying to my target platform about 30 minutes later.

45m agoHN ↗

I am begging you on my knees to please stop posting this cringe.

The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

36m agoHN ↗

That it's better in specific ways? What difference does it make when it came out? The benchmark results are not going to change unless they're messing with the model.

30m agoHN ↗

Yes, but crucially, in ways that are increasingly decoupled from any practical pattern of usage, considering that I wouldn't see how you can argue that Astra is worse than Opus 5.

One man's modus ponens is another's modus tollens I guess.

30m agoHN ↗

What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.

20m agoHN ↗

You forgot to include whatever you're proposing instead.

"Trust me bro, Astra is better" isn't perhaps as useful as you seem to believe. I'm not even saying it is right or wrong, just that my opinion on this topic is still just one additional subjective data-point.

Only thing I wish with these benchmarks is that they would run repeat tests every couple of months. Then re-rank based on that too. We've seen a lot of performance fall-off after a couple of weeks with new releases.

19m agoHN ↗

This index doesn't have "Astra" and "Opus 5". Every entry with corresponding data is a `(model, reasoning)` tuple.

So I'm unclear what you're actually saying and wondering if you've missed that. Are you saying that at every reasoning level it says Opus 5 beats Astra? I just compared Opus 5 high to Astra high and it has Astra as generally better than Opus.

43m agoHN ↗

so its more intelligence than fable?

can anyone help me?

42m agoHN ↗

Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?

35m agoHN ↗

All anthropic launches are like this. They just post it and don't particularly put out the PR sprint that OpenAI does with videos, livestreams or whatever.

(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)