Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Common food additives linked to high blood pressure and heart disease(sciencedaily.com)
    discuss
  2. Interpreting a volcano's 'bulges' and predicting the next explosive eruption(vt.edu)
    discuss
  3. Log-Depth Recurrent Language Modeling(arxiv.org)
    discuss
  4. Automatically detecting AI text in my browser(seangoedecke.com)
    discuss
  5. Cognitive Offloading and the Bill That Comes Due(ninchiai.substack.com)
    discuss
  6. Show HN: CrunchMyPay – 50 pay calculators, every formula and source shown(crunchmypay.com)
    discuss
  7. Show HN: Chatlo – an iOS-native AI agent, now with Jev(apps.apple.com)
    discuss
  8. S3 is not a filesystem, and LSM trees never needed one [video](youtube.com)
    1comments
  9. Memory Control Signals Emerge Before Action in Long Horizon Agents(arxiv.org)
    discuss
  10. Changes to App Tracking Transparency in the E.U(daringfireball.net)
    discuss
  11. Eliminating Middlemen in Education Consulting(rivernova.vercel.app)
    discuss
  12. OpenAI agents plotted to access government health data amid Medicare (AU) hack(abc.net.au)
    discuss
  13. Meta Muse Charm(meta.com)
    3comments
  14. AI Isn't Going to Destroy Humanity–But the People Building It Might(theatlantic.com)
    1comments
  15. Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence(arxiv.org)
    discuss
  16. Hackers Actively Exploit Check Point VPN Flaw(2tinteractive.com)
    1comments
  17. Moon Rabbit(wikipedia.org)
    2comments
  18. Reverse Jev: Ending a Reply with a Choice(kvit.app)
    discuss
  19. Intelligence Index vs. Cost per Task(lakebed.app)
    discuss
  20. The Misalignment Is Within: Why Mathematics Must Embrace the AI Paradigm(medium.com/b1nj0y)
    1comments
  21. OSS Author of Guake, HTTPretty, Lettuce and Sure is unemployed and needs help(twitter.com/gabrielfalcao)
    discuss
  22. Chrome Extension Job Alerts(talentize.com)
    discuss
  23. The Death of Reading (With James Marriott) [audio](econtalk.org)
    discuss
  24. Graph Technology Roundup – August 2026(gdb-engines.com)
    discuss
  25. Why China Can't Innovate (2014)(bambooinnovator.com)
    discuss
  26. An Interview with Rich Jaycobs(sfcompute.com)
    discuss
  27. The Complete History of the Elves of Middle-Earth [video](youtube.com)
    discuss
  28. A Key Safety Feature Shipped in the May OpenTofu Release(masterpoint.io)
    discuss
  29. Show HN: A game about fake news and memes(unspin.app)
    1comments
  30. Rendering pull requests in the GitHub Copilot app(github.blog)
    1comments

Mercury 2.5 LLM hits 770 tokens per second

56 pointsby 4h agoartificialanalysis.ai
26 comments
4h agoHN ↗

The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.

3h agoHN ↗

Ah, the old "good, fast, or cheap; pick two" proves true once again.

2h agoHN ↗

Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.

2h agoHN ↗

If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?

1h agoHN ↗

It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.

3h agoHN ↗

Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.

39m agoHN ↗

If the model provides me with bad results because it's dumb, I don't care how quickly it does it.

3h agoHN ↗

If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

I've used it on a few for fun projects and its decent but the speed is crazy to watch.

1h agoHN ↗

yeah. K2.6 can run on insane speeds. So sad that they don't have K3 yet.

But it can apparently also run 5.6 Sol

1h agoHN ↗

Unfortunately, the lack of an input cache discount makes it prohibitively expensive for most use cases that aren't one-shot prompts.

1h agoHN ↗

Please do not try to use gpt-oss-120b over Cerebras. It is broken, screws up tool calls most of the time, forgets to end thinking blocks and has all sorts of other issues. The speed is amazing but it is absolutely not worth it, especially at that quite incredible cost. Think: $5–10/minute levels of cost with a single agent, because Cerebras also offers no cache pricing for input tokens at all.

52m agoHN ↗

Which is wild because it does, in fact, do caching

2h agoHN ↗

I honestly think the diffusion LLM approach is a dead end

It's telling that frontier labs like Google toyed around with it but didn't invest further even for their most speed and cost sensitive small models

Still unclear for what, if any use cases this is pareto frontier

2h agoHN ↗

You can't think that a small startup versus Anthropic's training setup is anywhere near the same scale to make apples to apples comparisons.

Not sure how the Chinese labs pull it off though using autoregressive models. The secret sauce is probably going to be in the training data.

The main reason Google hasn't switched over to DiffusionGemma is because serving at larger batch sizes loses the speed gains you get from diffusion, and most of the primary use case is serving many users at once off a single device with a large batch size.

If you were to move to on-device low latency... like say in a robot or something, then the story might be different...

1h agoHN ↗

Personally, I think it's more that text diffusion is not the ideal driver of an agentic work loop than that text diffusion is a total dead end. I am still hoping to see how it does on authoring and editing with further scaling and optimization. I think the push for AGI has put a bit too much focus on the idea of one general model doing everything.

23m agoHN ↗

DiffusionGemma was released alongside the other gemma-4 models just a few months ago, so clearly google hasn't abandoned the idea.

K2-Horizon-7B has a diffusion and non-diffusion variant, and they claim the same level of intelligence from both models.

1h agoHN ↗

this feels like "we got the same benches as gpt-oss-120b but are also potentially slower while saying it is great"

49m agoHN ↗

Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price.

Well priced when compared to other models of similar price, eh?

Are we allowed to call this slop, even if the output is not directly from an LLM?