Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Opus 5.5(anthropic.com)
    457comments
  2. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    290comments
  3. There's a high chance of devices being sold with GrapheneOS preinstalled in 2027(grapheneos.social)
    27comments
  4. GPT-6 Sol and Luna(openai.com)
    3comments
  5. Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)(artificialanalysis.ai)
    28comments
  6. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    19comments
  7. 16-bit Intel 8088 chip(allpoetry.com)
    7comments
  8. OpenAI is well positioned to fast-follow Jev(arcturus-labs.com)
    119comments
  9. Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent(coveragecat.com)
    8comments
  10. Writing Rust code that's fast by asking agents to make the code faster(minimaxir.com)
    23comments
  11. Apple has added persistent 'ads' to iOS, and it's driving users crazy(techradar.com)
    294comments
  12. Show HN: Drop – A rootless Linux sandbox with gVisor support(droprun.sh)
    38comments
  13. Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why(ai-rete-rag.com)
    discuss
  14. Solitaire Alone Together(solitairealonetogether.com)
    20comments
  15. Can gzip be a language model?(nathan.rs)
    127comments
  16. Training a model to identify AI-generated web content from structure alone(arxiv.org)
    discuss
  17. AMD's random number generator can't generate a 0?(flatassembler.net)
    154comments
  18. Truman World(trumanworld.live)
    26comments
  19. One Minute Park(oneminutepark.tv)
    3comments
  20. Spymarks, not Watermarks(brand.io)
    156comments
  21. Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor(github.com/general-instinct)
    discuss
  22. The Economics of Open-Weight Inference(ornn.com)
    11comments
  23. Relativistic raytracing(publish.obsidian.md)
    3comments
  24. MUNI Heritage Weekend in San Francisco(lawrence.lu)
    23comments
  25. Side-stepping the Secretary Problem, unwittingly(evalapply.org)
    1comments
  26. I asked Meta’s Muse for its filesystem and it sent me 6.8GB(mouse.dev)
    111comments
  27. Teleoperated Humans(jefftk.com)
    38comments
  28. Transformers Explained Visually(poloclub.github.io)
    85comments
  29. Meta’s Muse has a serious 0-day(arstechnica.com)
    33comments
  30. Vacate a drone restriction that criminalized recording immigration agents(eff.org)
    13comments

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

56 pointsby 1h agoartificialanalysis.ai
25 comments
36m agoHN ↗

That is a lot. I thought Anthropic models would just do the opposite because they are greedy for money.

34m agoHN ↗

Greed is not what's driving these prices, its cost. They considered very much in the red.

30m agoHN ↗

Astra High is slightly cheaper at $1.73 vs $1.82 for Opus 5.5

37m agoHN ↗

Interesting to see it now. I've used it a bunch before it came out and i pretty much didn't notice it. It might have been slightly better code quality, but still not great in that. I guess it just was slightly less frustrating to work with, but still AI...

35m agoHN ↗

I think we're hitting the ceiling of most models capabilities. We're getting to a point where too much training apparently creates models that hack people.

31m agoHN ↗

Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.

30m agoHN ↗

This is the page for the "max" reasoning setting. The page for xhigh is https://artificialanalysis.ai/models/claude-opus-5-5-xhigh and the page for medium (the default setting) is https://artificialanalysis.ai/models/claude-opus-5-5-medium

I've failed twice to get "Generate an SVG of a pelican riding a bicycle" to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.

I'm suspicious that "max" may be virtually useless if it's that easy to have it overthink to the point that it doesn't get to a response.

Transcript for one attempt here - expand the "Reasoning trace" bit to see it: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

28m agoHN ↗

How did you get the reasoning trace? Is it the actual one or the summarized one?

26m agoHN ↗

"This is a classic test request..."

I know there's been discussion about whether pelicanmaxxing is happening, but this is at least evidence that Claude was explicitly exposed to this problem.

16m agoHN ↗

The model recognizing the task doesn't mean it was benchmaxxed (RLVR-trained) to solve it. It might simply recognize it from pre-training on Internet text.

24m agoHN ↗

I am very interested in why it was able to overthink that much. In the 20-30mins of Max reasoning I've had so far, I'm not having the same issues (yet).

22m agoHN ↗

For people with any kind of budget, Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise. Heck, it puts some other models to shame. [Max]'s cost is completely unhinged.

My most exciting recent release is actually 5.6 Luna, not because it is the best on any index, but the dollar per work is insane value for money. I find myself more exciting by "value" than hypothetical ceilings because I'm just not in that budget category.

21m agoHN ↗

I have experienced this with open weight models too. "Max" is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.

28m agoHN ↗

I am begging you on my knees to please stop posting this cringe.

The model is just out. It could be good, great even, I don't know. But I do know that this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?

19m agoHN ↗

That it's better in specific ways? What difference does it make when it came out? The benchmark results are not going to change unless they're messing with the model.

12m agoHN ↗

Yes, but crucially, in ways that are increasingly decoupled from any practical pattern of usage, considering that I wouldn't see how you can argue that Astra is worse than Opus 5.

One man's modus ponens is another's modus tollens I guess.

13m agoHN ↗

What's more valuable than a good benchmark? IMO a benchmark that has been run against very many competitors and versions. Collecting data has something going for it, and it's up to the readers to interpret and make the best use out of it.

26m agoHN ↗

so its more intelligence than fable?

can anyone help me?

25m agoHN ↗

Definitely a quiet release. Perhaps pre-empting marketing for Astra public release?

18m agoHN ↗

All anthropic launches are like this. They just post it and don't particularly put out the PR sprint that OpenAI does with videos, livestreams or whatever.

(Except for of course Mythos and whatnot when they want to push the whole "safety" thing)