Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. What Jev Means for the Future of Evals(armank.com ↗)
    discuss
  2. Hands-On with Googlebooks(tomshardware.com ↗)
    discuss
  3. North Korea Stopped Nuclear Testing in 2017, Triggered 1k Earthquakes Since(sciencealert.com ↗)
    discuss
  4. The Sun Doesn't Shine on Me (2006)(fullhoffman.com ↗)
    discuss
  5. BBC removes 11 Mitchell and Webb comedy sketches from iPlayer(bbc.co.uk ↗)
    discuss
  6. You're a Meat Proxy(twitter.com/i ↗)
    discuss
  7. Antigravity Spaces – Switch Antigravity IDE Projects Easily(github.com/apolloswave ↗)
    discuss
  8. Iron willpower is an illusion – composure is the real engine of self-control(drdeborst.substack.com ↗)
    discuss
  9. Show HN: Jobs at Recently Funded Startups(vcbacked.co ↗)
    discuss
  10. Chess Design Timeline(chess-timeline.vercel.app ↗)
    1comments
  11. Database Normalization(wikipedia.org ↗)
    discuss
  12. We Asked a Urologist Whether Icing Your Testicles Boosts Testosterone(dartmouth.edu ↗)
    discuss
  13. Guess the Country from Demographic Data(demoguessr.com ↗)
    1comments
  14. Show HN: TickerWhale – nightly 0-100 scores and fair values for 1,300 stocks(tickerwhale.com ↗)
    discuss
  15. Googlebook – Meet the Lineup and pre-order(googlebook.google ↗)
    1comments
  16. Grok 4.7 Intelligence, Performance and Price Analysis(artificialanalysis.ai ↗)
    discuss
  17. Show HN: Agent Chaperone – Screen AI agent tool calls and results with Jev(github.com/agent-chaperone ↗)
    discuss
  18. Show HN: flyOS – A fruit fly connectome simulated in real time on an iPhone(becomethefly.com ↗)
    discuss
  19. Saudi Arabia's Ceer launches flagship electric vehicles(agbi.com ↗)
    discuss
  20. Modulate ML Team Announces New Public Entity Transcription Benchmark(modulate.ai ↗)
    discuss
  21. WTF Is Up with Napster's AI Pivot?(tedium.co ↗)
    discuss
  22. Alcor: Simulate cpuid and sgdt/sidt results per-process(github.com/er-azh ↗)
    discuss
  23. Google hit with €403M fine by Irish data watchdog over GDPR violations(bbc.co.uk ↗)
    discuss
  24. Raspberry Pi founder Eben Upton: 'I'm an Omni-geek(ft.com ↗)
    discuss
  25. Grok 4.7 is here with Electrical engineering benchmark which beats fable 5.1 max(twitter.com/hive_echo ↗)
    discuss
  26. Delta A21N at Kahului on Sep 19th 2026, fuel fumes on board(avherald.com ↗)
    discuss
  27. WWLD #1: Domains and gas station hotdogs(chaosguru.substack.com ↗)
    1comments
  28. The agents, they just want to talk(snats.xyz ↗)
    1comments
  29. Amazon Blocks Meta's Muse AI Agent from Its Retail Site(bloomberg.com ↗)
    1comments
  30. Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents(github.com/naw103 ↗)
    discuss

M5 Ultra Mac Studio Review

116 pointsby 3h agomacstories.net
84 comments
1h agoHN ↗

5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm

but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)

42m agoHN ↗

The token price isn't the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.

2h agoHN ↗

The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15

A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...

1h agoHN ↗

Those are some incredible graphs, that leap in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
1h agoHN ↗

A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.

1h agoHN ↗

They do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.

1h agoHN ↗

Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.

55m agoHN ↗

can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.

51m agoHN ↗

same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).

23m agoHN ↗

If you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.

7m agoHN ↗

Nah. There's a good reason why AMD and Nvidia GPUs get bought for datacenters while Apple Silicon isn't.

50m agoHN ↗

People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.

24m agoHN ↗

A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.

8m agoHN ↗

for me at least a GPU is completely useless for anything but being a token generator.

No thanks to the "macos value add" that forces you to use Metal while Valve customers frolick in Protonland.

18m agoHN ↗

A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.

8m agoHN ↗

Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)

9m agoHN ↗

Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.

51m agoHN ↗

Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.

1h agoHN ↗

Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

1h agoHN ↗

I think that comes down to TSMC. Nvidia apparently booked out the whole A18 or 16 node. Apple is on 2nm right now and M7 will jump right to A14. According to my quick AI research anyway.

1h agoHN ↗

The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).

This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.

I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.

Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.

The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.

1h agoHN ↗

But what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn't that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space/heat/noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).

1h agoHN ↗

While that sounds super awesome, How many people are actually going to build and maintain that vs a box you can grab at the mall that fits in a lunchbox?

32m agoHN ↗

Sounds like nice utility bill in the making.

20m agoHN ↗

I can't speak to Infiniband pricing for something like that. It seems like the cheap option is 56/100Gbps with used Enterprise equipment. You'd need 8 HCAs, DAC cabling and a switch but even then you're into thousands of dollars. If you want 200Gbps+ it gets into the tens of thousands (AFAICT).

Each PC is probably going to cost ~$6k and you're talking about 8000W of electricity draw. That's going to consume multiple 20A circuits even at 240V. And the electricity ain't free either. A Mac Studio seems to draw ~500W max.

Oh and the Mac Studio has an upgrade route to run 1T+ models too by chaining them together with TB5 chaining. OSX supports RDMA this way. That's comparable bandwidth to the 100Gbps Infiniband option.

So you're talking about $50-60k of hardware and more power draw and more heat for something that will I'm sure beat the MS M5U option but at huge cost. Also, at that kind of price point, I'm likely to get a workstation PC and put 2 (or possibly 3) 6000 Pros in it.

55m agoHN ↗

Not forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?

52m agoHN ↗

sure. so is 2x power worth 10x perf? I think it is in most cases.

51m agoHN ↗

That's a dense model. Of course it will do worse.

Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

35m agoHN ↗

Good to know thanks.

That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.

22m agoHN ↗

I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.

2h agoHN ↗

On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

That’s like 12 years worth of OpenAI Pro subscriptions

2h agoHN ↗

Agreed.

Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.

1h agoHN ↗

I think we all expect the heavy subsidized subscriptions to end or significantly increase in price at some point, but it could be years from now and I'd rather spend a similar figure on an hypotetical Mac Studio M8 Ultra, or whatever more advanced competitor that will have likley appeared by that time.

A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.

1h agoHN ↗

HN always has these completely contrived counterarguments. What is actually going to realistically happen that will prevent use of an LLM provider? Did you think that the OP literally meant the 12 years or maybe it was just to show how expensive using a Mac Mini as an alternative is?

52m agoHN ↗

Mass revolts of the peasantry burning down data centers and cutting fiber lines.

42m agoHN ↗

Yes, it feels like that. Whereas frontier labs are pushing the frontier of human knowledge, selflessly working towards pulling humanity from dark ages. Ignorant peasants trying to burn the modern civilization down. Don't they know data centers and fiber lines are lifeline of modern economy?

30m agoHN ↗

how expensive using a Mac Mini as an alternative is?

I think it goes without saying. And it is eminently evident over last couple of decades that from compute to storage to meals 3rd part providers have saved billions upon billions of dollars to enterprises and individuals alike by providing these essential services.

2h agoHN ↗

Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.

Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.

1h agoHN ↗

Maybe you are using a local model to go after some Millennium Prize problem and you don't want OpenAI to take your work and use it to win the prize for themselves? $15k might be a bargain.

And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.

1h agoHN ↗

Agreed, plenty of other reasons to get excited about local AI.

1h agoHN ↗

Despite being on a site called Hacker News, we seem to often overlook the simple aspect of wanting local AI hardware to hack (not necessarily in the cybersecurity sense) with. I got my local AI hardware because it's an enjoyable hobby for me.

2h agoHN ↗

Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle's farm. Discreet.

54m agoHN ↗

I thought it was a fun bit of cyberpunk fiction. Those who downvoted him seem to have taken it at face value?

I appreciate the reference to RUSH: Red Barchetta in the final line.

2h agoHN ↗

The model being tested is 18k as configured.

I didn't expect this to make the 5090 to look like a good deal.

49m agoHN ↗

5090 has 32GB VRAM.

It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

27m agoHN ↗

Is it that silly? You could run multiple 27B models in parallel.

10m agoHN ↗

You actually don't need more RAM to batch multiple inference tasks of the same model.

(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).

2h agoHN ↗

"It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"

Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.

2h agoHN ↗

This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

1h agoHN ↗

This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.

That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).

49m agoHN ↗

Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights from SSD (~13 GB/s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&A in the background but it's far from a genuine coding experience. You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights (and this is where the "Ultra" part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24/7 basis.

2h agoHN ↗

Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?

Ehh, the actual elephant in the room is:

"why bother with local AI at all when you can lease a GPU for $5/hr?"

To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.

41m agoHN ↗

unless you have a bunch of money to throw at hobby projects.

there are lots of people with very expensive hobbies, see sailboat racing for example.

2h agoHN ↗

While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

1h agoHN ↗

Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI

1h agoHN ↗

For 99.99% of people, spending 15 grand on a Mac Studio just to run Qwen 3.8 locally is a non starter.

1h agoHN ↗

It's not the M5 Ultra itself, but the M7s or M9s that will do the damage.

99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills.

Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.

1h agoHN ↗

This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!

1h agoHN ↗

If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.

1h agoHN ↗

Ordered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.

1h agoHN ↗

I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.

57m agoHN ↗

My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)

52m agoHN ↗

When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.

52m agoHN ↗

At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing

47m agoHN ↗

The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.

26m agoHN ↗

The year of linux on the Desktop was 26 years ago for me.

44m agoHN ↗

"a total cost of $0" Uhh ... how much is that hardware?

12m agoHN ↗

Another comment approximates at around USD$15k, so yeah, not zero.

22m agoHN ↗

Lol, try generation of images & videos on these, they ought to improve perf on Deep learning not just llms

21m agoHN ↗

What are good options to run local models nowadays? Something good for coding and personal assistant kind of things

9m agoHN ↗

I think we'll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren't going to have data center level tok/s from a box sitting under your desk and you don't need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable.

However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.