- 111comments
- 67comments
- 10comments
- 117comments
- 3comments
- 90comments
- —discuss
- 28comments
- 171comments
- 813comments
- 1058comments
- 232comments
- 16comments
- 90comments
- 92comments
- 165comments
- 200comments
- 140comments
- 109comments
- —discuss
- 82comments
- 6comments
- 428comments
- 290comments
- 190comments
- 659comments
- 189comments
- 85comments
- 243comments
- 175comments
It would have been funny if the "how" had been dropped to end with "Did AMD Ryzen get 50% faster in two years?"
Am I missing something? Max boost went up by only 15% but the base frequency went up 38%
Better thermals maybe? Less throttling?
its manufactured on a smaller node so probably has more leeway for higher freq on all core loads.
I’d also be weary of single core scores,Geekbench is known to be favoring specialized instruction sets like AVX or encryption extensions more and more as the version number progresses. Though I don’t know if GB6 is also like this I wouldn’t be surprised it it was caused by it.
One of the biggest improvements to the Ryzen 7 X3D series (e.g. 9800) is that the caches have been moved from one side of the die (to the other), which places the major heat source closer to the heat sinks.
SO yes, less heat throttling.
Smaller node size means less voltage.
Power is heat, and the amount dissipated is the square of the voltage- a processor that is reliable at less voltage means you can get a lot more frequency in the same heat envelope.
Yet reliability decreases as frequency goes up, and that can only be stabilized by adding more voltage- so the faster you run the processor, the more voltage you ultimately have to give it, so the power/heat produced grows exponentially until you can't get rid of the heat fast enough (at which point your only option is to actively cool the chip).
This is overclocking 101.
Note that classic overclocking was viable because of arbitrage- buying a processor, pushing it to the point it got too hot, and stress-testing it at that temperature to ensure reliable operation. Processor manufacturers all do that formal verification at the factory now as they'd be uncompetitive otherwise, especially in laptops.
As ProllyInfamous mentioned; bringing the cores to the top helped reduce heat, which allows them to boost higher for longer periods then the 7800x3d.
2 years sounds very fast. It seems to be because the 3d cache SKUs of each generation were released at different stages. The Zen 3 one was a later variant.
First zen 3 November 5, 2020 with desktop processors. First desktop Ryzen 9000 processors on August 8, 2024. So the generations were about 4 years apart.
Yes I stated that as well in the previous comment [1] on how Apple got 50% faster in 3 years.
[1] https://news.ycombinator.com/item?id=49772502
2022 to 2024
I suspect the improvements are even more dramatic going from Zen 1 through to Zen 5. AMD has really hit the jackpot with how scalable the Ryzen CPU is considering how they're able to improve the performance from year to year. This is a stark difference to the FX series during the 2010s, which saw very small YoY performance increases by comparison. Ryzen really is AMD's equivalent to what Nehalem/Core was for Intel back in the mid 2000s.
That is all true but I will defend the FX series a little. Mostly now that there is a lot of software that scales across cores better now, they haven't aged as terribly as others have. They aren't great but not terrible considering.
That's pretty much irrelevant since the AMD's FX arch's issues weren't that SW at the time wasn't using all the 8 cores. Intel dropped the Core 2 Duo and Quad into the era where most SW was still stuck in single threaded for a long time and those CPUs still ripped single-threaded SW tasks regardless.
Here's the big reasons why the FX sucked back then and why they still suck today in the multi-thread SW era:
So unless you're into collecting vintage CPUs as display pieces, this one definitely belongs in the e-waste pile instead of burning electricity, because it did not age like wine with the adoption of SW multi threading like people were hoping.
It blows my mind that AMD watched Intel try to do basically the same thing only a few years prior with NetBurst, and fail so badly that they had to scrap that entire evolutionary branch and start over – and AMD still went and did it again themselves anyway.
I remember rumblings a decade-ish ago that basically their hand was possibly forced to release the thing to avoid a full on revolt about abandoning all of their work; after all, Intel had been investing in deep pipelines for a while before, they were strapped for cash after the ATI Acquisition, and other 'server-ish' CPUs had done CMT type things in the past (keeping in mind that AMD was seeing a huge surge in server market share due to Hammer.)
Hyperthreading/SMT is a significant boon for heavily threaded workloads. What makes that such a win while Bulldozer's implementation of "two integer units sharing a front-end, cache and FPU" is supposedly so bad? Because that description makes the 8 core Bulldozers sound exactly like a 4 core with SMT.
That's hugely debatable and depends on SW workloads and the SMT implementation + CPU pipeline design.
In SMT the execution engines, ALUs, FPUs, and caches are completely shared. When one thread stalls waiting for RAM, the second thread sneaks into the idle execution units. At best, SMT yields a ~10% to 20% throughput boost over a single thread.
It's not the same thing. Bulldozer arch sits between a true 8-core and 4-core + SMT implementation.
Exactly, which is a significant benefit for how marginal the costs are.
Then surely it should be even better than 4 cores with SMT?
If you're going to argue that the problem with Bulldozer was it's weird semi-SMT solution, you need to explain how it would've been better without it (aka as a regular quad core). Because even if it just gets the 10-20% performance improvements from being a form of SMT it would be better to have it than to not. And if you have lots of integer unit-bound threads, it should be even better than that.
I explained all the bottlenecks of the architecture in a comment above, that the issue was more than 4-core +SMT instead of true 8 cores. Please read it.
I did read it. It's not clear from it why you think having 2 integer units per core in a SMT-like configuration makes it worse. If you have <=4 threads it doesn't matter, just schedule the threads on different proper cores. If you have >4 memory/FPU/front-end heavy threads you should see the same benefit as SMT. If you have >4 integer arithmetic-bound cores, you should see a significant benefit beyond what SMT would give.
Now the very long pipeline and high memory latency are obviously significant issues with the architecture but those seem disconnected from the 4-core+SMT issue? I'm not questioning those issues at all, it's just not the part of your comment which interested me
EDIT: okay so in this comment: https://news.ycombinator.com/item?id=49809017, you explain that there's actually a fairly large part of a core that's duplicated, not just two integer units. If each "core" gets its own integer unit, register file and L1 cache, you're actually paying a ton of die space for it, unlike SMT which is "free". I can totally get how that can be a terrible trade-off for most workloads if it all ends up mostly starved due to front-end/FPU/memory throughput.
Bulldozer was twice the die size of Intel 4C/8T but with lower performance.
AFAIR Steamroller was a big 'correction' of the Shared resource issues in the arch (I can't remember if other revisions had other improvements).
AMD was also having to deal with the fact GloFo split off and was relying more on general 'bulk' lithography, which kneecapped them for some time especially due to yield issues on the FX series and overall cost of that deal.
Intel also very quickly after, released Sandy Bridge and aggressively scaled it up and down; the 2500K was so cheap yet powerful I know of at least one setup that ran for a decade an only got replaced because they needed to upgrade to windows 11 for compliance-esque reasons. My own 2500K I replaced in 2017-2018-ish, only because either the motherboard took an unfortunate dive and it was easier to replace both at once.
FWIW, I did do a cheapie FX build in 2015ish for my then-girlfriend as a DVR and light gaming/emulation 'under the TV box', and it did the job well for the price, but it definitely wasn't anything amazing.
It was a tough time for AMD for sure. I think the 'split' between the Cat cores (Bobcat/Jaguar) also hurt them from a resource standpoint, although one could argue that it also kept them alive to recover (i.e. Jaguar in XBox One and PS4 being a volume contract part) [0]. They did a lot of moves that caused short term pain (that glofo spinnoff helped pay off the ATI Acquisition AFAIR) but helped them become the company that is still surviving today.
[0] - One odd side note, I still find it odd that they never did a dual channel Jaguar laptop part. I still ask whether it was because it would have made the FX look that bad...
Yeah that's hyper-threading intel was doing it as well and all modern CPUs do it as well. Where AMD dropped the ball, was they did not disclose that in their marketing as clearly as they should.
All CPUs today are marketed as x cores 2x threads, back then some AMD marketing genius in their infinite wisdom put 8 cores on the box, instead of the honest 4 cores with hyperthreading.
No, AMD's FX "fake" 8-core was more than just 4-cores + hyperthreading. In SMT(hyperthreading) the execution engines, ALUs, FPUs, and caches are completely shared, whereas on FX design, they built two completely separate integer pipelines (schedulers, register files, ALUs, and L1 data caches) inside one module. Only the instruction fetch/decode front-end, the FPU, and the L2 cache were shared. So the FX design would be an in-between a 4-core + SMT and a true 8-core.
Fair, I had not delved into the details, but still they were not full cores and the marketing did not make a real distinction.
I had a pilledriver one, it was a perfectly good cpu, I would buy it again. If I remember back then it was the best overall performance per dollar, the alternatives if I remember correctly were i7-39.. and i7-38.. and were at best 50% more expensive for 10-15% more performance.
Depends what you were doing with it. The Piledriver only beat the Intels in heavily multi threaded (preferably integer) workloads like media encoding, which is why it was popular with media creator workstations on a budget, but for most consumer real world tasks at the time, like video games, Intel was way ahead in performance even though it was more expensive.
The Piledriver would win the consumer bang/buck mindset back then because of the 6-core part was reasonably priced and unlocked for overclocking, so people would overclock them to beat the more expensive (locked?) 4-c/8-t Intels at a lower price, but that ignored the costs of massive extra power draw(100+ W) over the Intel, the need for beefier more expensive coolers and power supplies, more expensive AMD motherboards with beefier MOSFET power delivery stages built to withstand the higher power draws of the Piledriver, so in the end the actual bang/buck gain of the AMD system wasn't remotely as big as people were making it out to be, they were just happy to get a "6-core" AMD cheaper than a 4-core Intel thinking more cores = more "better", same how having more mega-herz was also more "better" a decade before that.
The Team AMD VS Team Intel wars on forums on these topics were wild back then.
Media encoding is actually an FPU workload, and a pretty brutal one at that.
Media encoding might not use much floating point arithmetic, but it does use massive amounts of packed integer SIMD. And all SIMD instructions (both integer and floating) execute on the shared FPU, not the integer unit. It's only scalar integer instructions that execute on the integer unit.
Which leads me to believe that Bulldozer's shared FPU is not a bottleneck at all. Most evidence seems to point to the shared frontend being the primary bottleneck (which is why steamroller puts some effort into duplicating the instruction decoding, for some pretty large IPC wins)
No. it's just the opposite of Hyper-thereading. SMT it's about maximize the resource usage of a CPU, running 2 or more threads at the same time. AMD used the opposite technique of SMT. It had a technical name that I can't remember now, and wasn't invented by AMD.
They did just fine in parallel workloads, so I think this is not accurate. The design scaled just fine. The problem was that each core was weak.
Depends how you define "doing just fine in parallel workloads". The contemporary competition from Intel that was 4-core + SMT was beating AMD's 8-core FX CPUs in most real-world tasks and benchmarks at the time. The 8-core AMD broke even and rarely won only in >4-thread strictly integer benchmarks and some >4-thread media encoding tasks/benchmarks. So if you wanted a prosumer media encoding workstation a budget then yeah, the AMD was better, but for most real world task, it really wasn't.
Can you elaborate and be more exact? What you wrote is technically vague and doesn't mean anything in technical dissection/terms.
When we say a design "scales", that means that increasing the size of the workload does not incur a lot of overhead. If contention between shared resources meant that the design was not able to achieve an ~8x speedup when run with eight parallel threads, that would mean the design was not scalable. But we did in fact see a roughly 8x speedup with eight threads, so the design scaled just fine. The problem with the design was that each core was individually crummy, so even eight cores running in parallel had lackluster performance.
The myth that each two-core module functioned more like one core with hyperthreading would suggest that these CPUs would have much higher per-core performance when lightly loaded than when fully loaded. That is not what happened. Each core was crummy even when lightly loaded, but under full load you would have eight crummy cores, which would beat four Intel cores on a lot of workloads.
The only time contention was a serious problem was with workloads that were dominated by floating point, which were relatively rare.
By that definition it definitely was not scalable.
Care to share a source? Because AFAIR there definitely was no 8x linear speedup with 8 threads even in benchmarks, let alone in real world use cases. The only benchmarks where those 8 threads would scale best and beat Intel were archival compression/decompression and media encoding. At everything else Intel wiped the floor with it.
Many real-world compute workloads, especially gaming related, are floating point.
This argument has never made very much sense to me. Yes, the decision to group cores into modules with some shared resources did introduce a bottleneck and result in lower performance than having isolated cores. But this didn't make the processor worse than if it didn't have the additional cores at all. These processors were at their best on highly parallel workloads. AMD shipped far more cores than Intel at the same price. The contention for the shared front end was not serious enough to make up for the core count advantage. How is the module architecture an explanation for why the design failed, when the bottleneck only becomes relevant in situations where the design is winning?
It's hard to find benchmarks from 15 years ago. But Phoronix finds what I describe: the design doesn't scale quite as well as a "true" eight-core design, but still scales better than the competition of the time which had lower core counts. Though it does seem straightforwardly bad at some workloads.
https://www.phoronix.com/review/amd_bulldozer_scaling/7
It's as FlowingRiver said. If we were to release the same parts again with modern software, Bulldozer would be a far stronger competitor to Sandy Bridge. FX aged better than Intel's designs of the same era. But Intel's designs were better to start with, so I'd still take the Core.
The reason they sucked is that they ran hot and their per-core performance was terrible.
I did run a multithreaded CPU only n-body problem resolvers in FX cores. So heavy double precision work and putting each core to 100% . A Fx-8370E would get a speed up around 7.5 times Vs the single threaded version A FX-4300 would get a speed up around 3.9 times Vs the single threaded version
So in a physical heavy computation task, involving double precision math (where the AMD design of share FPU units should penalize most), the FX cores where happy churning numbers with the expected speedup Vs a single core/thread version of the code. So stop saying that FX cores sucks at multithread. They fucking worked fine on that kind of tasks.
I just was to say the same thing. I had a good experience with a FX-8370E. They go for too many cores to early, and sacrificed some of the CPU performance to do it. The gamble gone wrong...
Yes, my FX 8350 was doing really well on my Gentoo machine for the price. It really depends on the use case - AMD went hard in on multi core while Intel focused on single thread perf.
Part of the reason we didn't see much in the way of YoY improvements for Bulldozer, is that AMD almost immediately abandoned it and threw resources at Zen after it launched.
Steamroller did see 30% IPC improvements over Bulldozer (all the design work would have been done before they switched to Zen), but AMD canceled the full FX version, and only ever shipped the APU version of Steamroller (with only 2 modules, aka 4 threads).
If they had shipped a Steamroller FX cpu, the generational improvements would have looked similar to many of the generational improvements that Zen received... but didn't really matter as Bulldozer started so far behind.
They never shipped the zen 5 epyc big cache chip (e.g. 9685X). :( I'd hoped they pulled it to produce a HBM integrated chip instead, but they didn't do that either.
I imagine they didn't want to cannibalize upcoming Zen 6 which will have the 3d cache
I don't think there are many fabs capable of producing HBM
My hope is on Huawei, which has recently developed a stacked chip architecture that reminds me of HBM, and I'm wondering if that's step 1 of China entering HBM production.
I don't know enough about hardware, so I might be saying absolute nonsense, happy to be corrected so I can learn where I'm wrong
The AMD Ryzen 9 9950X3D / EPYC 4585PX is a beast for 99% of compute workloads.
It has the highest base clock of any modern AMD chip (4.3 GHz), plus 16-cores, which is already more than most workloads need.
Can’t wait to see what Zen 6 brings early next year. Rumors say 24-cores and another big jump in single-core performance.
I looked at game benchmarks, and sometimes 5600 or 5600X were fastest. With all X3D processors so fast that I would not decide to upgrade above 5700X3D or anything cheaper. Now, with current RAM pricing, getting a X3D CPU with 90+ MB L3 Cache seams the way to go instead of doing a DDR4 -> DDR5 leap.
All stated for gaming, that is.
I still love my 5800X3D. Drop-in upgrade from a 3700X so I kept my board, RAM, cooler, everything.
Maybe I should do the same. I have 3800XT, I could go to 5900XT for like $300. I don't see myself upgrading anytime soon because these prices are laughable and will eventually come back down when the AI bubble pops.
I feel like most benchmarks don’t really capture CPU benefits to gaming.
It’s even difficult to find CPU benchmarks that don’t overemphasize 1080p and eSports scenarios.
My upgrade from 5600x3D to 9850x3D felt kind of dumb at the time, but I decided to do it because Micro Center’s bundle deals are so far below market pricing.
I was shocked at how much better it is. Benchmarks and FPS don’t really show things like micro stutters and little performance wrinkles like that. I’m not even sure 1% low FPS counts capture it.
A great example game for this is Oblivion Remastered. Upgrading my CPU alone with the same GPU took away the environment loading slowdown almost entirely.
The benchmark will tell you that I didn’t gain any FPS during gameplay but every time I open a door into the new environment my CPU is positively impacting the experience.
I would have said the exact same thing you are saying until I actually experienced upgrading to the best on the market. For the record, this is the first time in my life I’ve actually owned the best CPU on the market for gaming.
My old advice would have been to buy one of those sweet spot cheaper mid-range gaming CPUs, but my newer advice is really if you’ve already spent all that you’re willing to spend on a GPU (I have a 9070XT, my only upgrade paths are insanely expensive), buy the highest gaming CPU on the list for gaming benchmarks (e.g., I wouldn’t go crazy with a 9950X3D2 since it doesn’t have any gaming improvements above the 9850X3D).
And the thing about CPUs is they’re not insanely expensive like GPUs. We are talking a price delta of $200 between this beast of a CPU and something way more middling.
Oblivion remastered is an absolute dumpster fire for performance in general.
Given the massive increase in hardware prices lately, we are going to have to get by with mid range or even low end hardware for a lot longer. Studios will be forced to actually optimise games to not run like shit on a sub $8000 PC.
Every game should be targeting the switch 2 and steam deck in terms of power.
I think PS5 is a perfectly reasonable spec to be targeting, which was the higher end of mid range about 5 years ago, and still significantly out performs Switch 2 (which significantly out performs steam deck). I’m sure the GTA sales will demonstrate that the install base at this spec is wide enough for broad commercial success.
I don’t really want game experiences to be locked at a PS5 forever. That console is actually somewhat old now. There are phones that have more graphical power than a PS5. A MacBook Air has a more powerful GPU than a PS5.
I grew up with gaming systems that got rapidly better and while it can be argued that we are at a diminishing returns plateau of gaming horsepower, I don’t think the solution is to never get hardware upgrades ever again, even in the current RAM-constrained environment.
GTA VI is coming out with FPS maximum of 30FPS on the PS5 including PS5 Pro. That’s a subpar experience that is limited by hardware.
I didn’t mean to imply that we should target ps5 as an exclusive platform, but ps5 is a good hardware spec to target, as there are several platforms that can match it and it has a wide install base.
Yeah, and I see what you mean there. I think even when the PS6 comes out, game developers largely won’t be taking advantage of its capabilities just like how today they don’t really leverage the capabilities of an RTX 5090.
We saw with much of the PS5’s lifespan that many of its cross-platform games came out on PS4 and original Switch with minimal compromise.
Not enough of the install base owns those higher end solutions, so for the most part someone with a better performing rig is getting better FPS, higher resolution, and not much else. Even gaining additional draw distance is rare these days, and higher resolution textures are basically impossible to see.
If the console is going to survive rather than decline as a whole I think that Microsoft and Sony will need to get more serious about innovation rather than just making a samey spec bump at a high price that doesn’t motivate their users to upgrade. Nintendo already proved that basic and rather mild level of innovation can excite the market (original Switch).
I think it's limited by their software. The hardware buys them flexibility. Nobody is pushing it like the ps2 days.
But that’s the thing: I want to play oblivion remastered.
Having a powerful rig is the brute force way around the game’s performance problems.
My PC wasn’t $8000. My 9850X3D, 32GB of RAM, and motherboard cost $800 in a bundle. Yes, purchased in 2026 during the RAM crisis. I was also able to sell my previous CPU, RAM, and board for around $500.
Obviously this is not a budget build but about half the cost is the graphics card. I would have saved minimal amount of money if I had gotten a worse CPU.
Most benchmarks use singleplayer games with built it benchmark, you'll not see the biggest difference there. If you play online games like shooters, MMO, anything really with multiple players or entities, theres usually a much bigger difference. Also in rts games and such with many units.
Does the base clock even mean anything beyond something that may or may not vaguely indicate the performance? Its lower than what these chips can sustain with sufficient cooling from my experience. It's higher than what the chips will clock down to when not sufficiently cooled.
Base clocks from AMD are nothing more than the guarantee for any workload when the system exactly meets the listed power/temperature requirements.
They aren't the maximum or minimum but I think they are the only clocks with the specific guarantee "if you match X, you must be able to get Y". The max boost clocks usually aren't too hard to reach if you're trying though, there is just no guarantee you will in all workloads or how much it takes to get there.
I have the laptop version 9955HX3D (which I believe is a 9950 with 2.5ghz base, 5.5 boost, and lower TDP) in a Legion Pro 7 and it's a fantastic chip for a computer that I can move around. 16 full cores in a laptop is a luxury and 64gb of ram has proven to be a wise choice prior to price spikes
Let's not discuss battery life and overall size, however it helps the thermals.
I do a lot of CPU bound engineering and scientific compute on the computer and it chews through it.
I also bought the entire computer for $2500+tax new from Microcenter in December and it's appreciated 25% in value since then.
Compound this with Linux users seeing a ~8% gain (and some very big gains here and there) on performance from ongoing kernel improvements too! https://www.phoronix.com/review/linux-618-73-amd-epyc
This only goes back one year, to Linux 6.18. Wins would be even bigger if we go back another year. Also, this isn't tracking any of the rest of the improvements in userland: it's just the kernel. Some newer GCC, and upcoming new x86-64v3 targets will all have some pretty nice wins too.
Great days to be on open source. And it only ever gets better.
I love my frameworks strix halo box. 16 zen 5 cores with way more memory bandwidth than they should have for a desktop. Only missing the x3d cache
Is this what's helping drive their stock price to the moon?
They're starting to catch up with NVIDIA. Because AI.
No. They make server CPUs which is predicted/expected to be the next bottleneck after memory.
Serious question with 1GB L3 isn't it theoretically possible to boot a full Linux without ram?
Edit: dug in it, no, because cache isn't addressable and it's directly managed by the cpu.
I used to use a full Slackware Linux distribution with X-Windows on a 486 clone with 4 Megabytes of RAM.
Intel has a cache-as-ram (CAR) mode / Non-Eviction Mode, which is what you wanted but it is only available before DRAM is brought up via core initialization. Not sure about AMD.
AFAIK AMD doesn’t do this because it’s a pain in the butt to do from a hardware and firmware perspective. They basically have their equivalent of the intel ME do RAM for training and then the main cpu comes on with DRAM already up. At least it was like that five years ago when I last worked on an AMD part.
I suspect this will also be sliced across cores on a CCD, which means even less effective cache for a single core boot
I suspect someone at AMD could do this, but it would require a custom build of the CPU since the memory controller would have to change.
I'm not sure focusing on just 3 processors is enough to claim that Ryzen got 50% faster. There were already other, faster Ryzen processors available without the additional cache. The extra cache was a very new thing (and very temperature sensitive), so it makes sense that they start carefully and then get a lot of improvement quickly). And the 5800x3d was for an older chipset, with all the limitations that come with that.
It's not that temperature sensitive, 90C is plenty. If you can't keep the non-x3d processors underneath that you're just throwing away performance anyways
Are you? I'm pretty sure when I bought my 7800x3d, everybody was saying you had to be very careful about heat, because these damage more easily from high temperatures than regular CPUs. I just checked, and they do throttle at a lower temperature than others: 89 vs 95 degrees C.
It will still throttle itself so no real danger.
For performance reasons you need to be careful with heat on the X3Ds because it will throttle quicker (partially because of the lower limit, partially because the extra cache sits between the CPU and the cooler on that model making it harder to cool) but you don't have to worry about it from a damage perspective for the same reason: the limit is the point at which it'll drop performance to protect itself.
The x86 architecture still going strong. I've always though Ryzen was a fitting name, AMD had Ryzen from the ashes and struck a blow to the long dominant Intel.
When I built my Ryzen 1800X desktop, I named it "AMD Ryzing". My optimism turned out well placed.
I recently upgraded from that to a Core Ultra 7 that I got from work (e-waste recycling). It was not my intended upgrade path, and I hope Lisa isn't too mad lol (I kept my Radeon though).
How the hell does a high end CPU that's barely 2 years old end up in e-waste recycling? Was Richie Rich using them or something?
Sort of: there were about half a dozen in micro PCs in a load picked up from a hospital. I was kinda bummed that they were BIOS password locked (as were the other PCs from there). The model didn't have a password reset jumper (or any reset mechanism outside of contact Dell support with proof of purchase), so I couldn't sell them as whole systems, and stripped them for parts.
We get DDR5-based systems from time to time, but they're understandably rare. Most of the time, they still work. My rule: any day I come across DDR5 is a good day.
But why was the hospital throwing away what was essentially new high end systems that can run all the latest SW and operating systems? For what reason? It just makes no sense to me. Or am I too poor of a peasant to understand?
Are you by any chance in some super-rich part of the US? Because then no wonder your healthcare is so expensive when hospitals treat high end PCs as single use disposables. Here in Europe I sometimes see hospitals and doctors practices using 10+ year old PCs and still rocking the old 17-19 inch 1280x1024 CCFL LCD monitors from like the mid-2000s.
Regardless, I'd love to just run into high end PCs being thrown away here and pick them up for free, but where I live I see people barely throwing away their value-line Dell/HP Core 2 Duo towers with audacity to ask for 25+ Euros for their e-waste with the classic "no lowballs, I know what I got" attitude.
I don't understand it either. Two scenarios make sense to me:
Somewhat probable: everyone in a department got upgraded, no exceptions. 'This is a really nice PC, but the boss says it has to go.' (There were about 100 i5 9th gen mini PCs (still OK), and 500 t640 thin clients in that same load.)
Less probable: some department got closed and cleaned out.
Pittsburgh, Pennsylvania. Not particularly rich (rust belt), but healthcare is a significant part of the local economy. About 90% of healthcare here seems to run on HP, but Dells can appear.
My stomach turns seeing the dozens of 24" probably 1080p monitors we scrap every day. Pretty much all are significantly scratched when they come off the truck, otherwise reselling them might be good business.
Here in Europe all that is 100% resold either locally, or further to balkans/eastern-europe due to much lower purchasing power making the resale efforts of ~5 year old HW economically viable. I check the local "craigslist" often and I notice when a business is clearing house due to a flood of used HPs/Dells/Lenovos from a single user. However the list prices aren't remotely palpable to be good deals. I assume the person/business hired for clearing the house is marking them up significantly hoping to squeeze a big win from a mark.
Some of the tech companies I worked for here would first auction their older HW internally at every upgrade cycle to the employees before selling what was left to a clearing company, which makes me feel we're being robbed to be offered by employers to bid for their e-waste when in the US people find better stuff in the trash for free. Europoor indeed.
I assume those businesses don't bother reselling 1080p monitors and 2 year old PCs, and instead just throw them away, because the price of labor is so high in those parts of the US, that paying a full-time employee to take care of such resale tasks will cost them more money than they expect to make from the sale, so it's assumed to be cheaper and less hassle to just throw everything in the trash instead of wasting time on the used market dealing with tire-kickers just to gain what is essentially peanuts money for those businesses' bottom line.
Am I close with my assessment?
If an intern did it (or even a low paid full-time employee), one could easily turn a profit selling used equipment. (I work in the refurb department, reselling used stuff is literally my job!) 5 minutes per monitor @ $25/each = $300 per hour. You could sell to wholesalers for much less per unit, in exchange for not wasting time with random people.
I'm not sure how prevalent this practice is, but IT equipment is sometimes on a depreciation schedule. If an employee starts selling the company's PCs and monitors, an accountant will get very angry, because he told the government that the stuff is literally worthless, but it apparently isn't because it was sold for something. IANAL, but I'm fairly sure that's tax fraud. Once it's turned over to someone else (my company) for free, the valuation resets, or something.
Oh come on, stop felling sorry for yourself. I'm from one of those countries compared to which even Eastern Europe you offload hardware to is a paradise.
None of what you're describing is available here at all. I sometimes see Europeans willing to ship here (already very few) offer heavily used and scratched ThinkPads and Dells for money they weren't worth when they were new. Add at least $80 on top for shipping, and no warranty, you get what you get.
On the local market, you will find those Core 2 Duo with precisely the attitude you're describing, but at 2-4 times the cost.
So I buy everything new, which is expensive with our salaries, but we have no choice. At least new hardware is available now thanks to selling over the internet and improved logistics.
10 years ago you could only buy what was offered by local shops, which was always at least two generations behind, and at least twice as expensive as it was in the EU (forget the US).
For example, if things continued like this to this day, I'd estimate the newest CPU you could buy right now would be something like Ryzen 3600, at twice the cost it was in your country when it just came out. If we're lucky, maybe 7600 would already appear at a similar overprice.
I'm sure many parts of the planet are even in a worse position than we are.
My company's policy is that retired laptops are destroyed. (I suppose properly wiping them or removing storage devices should be enough but for some reason it's not considered sufficient.)
There are indeed lots of second-hand former company laptops being sold, and some companies even specialize in refurbishing and selling them. I don't know what percentage ends up being recycled but it's definitely not 100 %. I'd assume the more security-oriented the organization gets, the more likely it might be to avoid recycling devices intact.
With that said, I haven't seen two-year-old high-end devices getting retired. But if there are only a few of them, and a large fleet is being renewed, I suppose it can make sense from a fleet management point of view to renew the few perfectly current ones as well.
Single core performance still worse than an iPhone.
The x86 folks lost the ball completely. For workstation/ mobile needs you go to arm. For matmul at scale you go to gpus.
I guess windows gaming with separate gpu is the only remaining market for them.
Unless Nvidia releases a motherboard with an arm cpu on top of their RTX cards, and then that market is gone too.
The IPhone runs on a battery, with no active cooling.
There is no chance the sustained single core perf is better than a 150=200+w ryzen with a big noctua cooler on it. The power draw alone would kill the battery in a few minutes(?), and the heat would make it catch fire. lol
I don't think you are correct for single-core performance.
A20 Pro scores around 4725 at 8.9w. Geekerwan's review showed that the iPhone 18 pro could dissipate as much as 6.4w during long gaming tests. Cutting power by 50% likely still keeps around 80% of the clockspeed.
As 9950x scores around 3400-3450 in geekbench, or around 28% slower.
There's a very good chance that the iPhone single thread performance is genuinely faster even during sustained loads.
Sustained performance is usually based on cooling, not the chip. There is no reason the iPhone chip can't sustain its performance if you give it a small fan.
So whatever your argument is, it doesn't make much sense.
The M6 is based on A20 Pro and it sustains its performance in a Mac Mini indefinitely - likely longer than Ryzens.
The 9800X3D doesn't pull anywhere near that much in single threaded workloads, it'd fry the core. It's also way less efficient than the iPhone CPU design and on an older and less efficient semiconductor process from TSMC.
If you compare like for like, ie. take the process node out of the picture, then you need to compare the Ryzen 9000 series to the A16, as both use TSMC's 4NP process.
There is nothing like for like. We're comparing a fanless phone chip to machines that might have water cooling and nearly unlimited power.
By being more expensive?
It's impressive and yet it surely depends what you're doing with it? 50% better performance isn't going to make a local LLM feel quick and yet so many other aspects of computing are quite fast anyhow. I can browse the web fairly comfortably on a Raspberry Pi and it's a bit slow but manageable.
If your goal is web browsing you will want faster single core performance. Unfortunately that means Apple M series. If your goal is local LLM, I’m afraid a several-year-old GPU will smoke the fastest CPU available today.
IIRC, compared to GPUs, M-series is currently still stuck in memory bandwidths from around 2016. M7 might catch up to 2019 or so. So GPUs will be better for LLMs for a pretty decent while.
How so? I don't know of any other mainstream platform with >1TB/s memory bandwidth. Personally I don't want to deal with macOS but between the memory bandwidth and out of box Thunderbolt networking it's hard to argue that Apple doesn't have a couple significant advantages over the current alternatives.
I mean, RTX 5080 has nearly 1TB/s, 5090 has nearly 2TB/s. Maybe you are talking about CPUs / unified memory platforms? I agree nobody else does it better. But for LLMs, GPUs can still be significantly faster than even the most advanced Apple silicon on the planet. TTFT in particular is super inferior with Apple, for now.
That's probably also the reason Apple had to reluctantly give into Nvidia servers for the initial rollout of Siri AI, though they claim to use trusted computing extensions to reach an acceptable level of privacy. (I do not trust that nearly as much as the Apple Silicon nodes)
It'll be amazing five years or whatever down the line to see Apple reaching those figures. They seem to be heading in that direction lately.
I suppose it depends what models, but even heavily quantized the mid tier local models need >100GB of memory so unified memory platforms are the only option I would consider a mainstream option (doesn't require special order etc). If you want to do it in two slots that's two Blackwell RTX 6000s which is around $50k list just for the cards, and if you want to do it in four+ slots that's outside the realm of normal desktop PCs. Right now 2x DGX Spark, 2x Ryzen 395, or 1x Mac Studio are the three configurations that get to 256GB of reasonably fast memory that you can order online and plug in, for around $8k to $12k depending on the details.
I Picked a threadripper for a box that needed a lot of IO and figured that with that many PCI lanes I couldn't go wrong. But I have to admit I've been more than pleasantly surprised by the performance of the CPU as well, it - easily - outperforms all of the XEON and I7 based boxes that I have. The only thing I wished I would have done different is to max it out with 256G DIMMs when they weren't the price of a car.
Zen 5 really is nice. I swapped my RAM from a Xeon W Sapphire Rapids machine into a Threadripper Pro 9000 series machine and I get almost double the memory read performance, plus it's a heck of a lot faster in single and multi core performance, and it runs cooler and quieter. Huge win all around. Aside from the price... (I went from Xeon W5-3435X to TR Pro 9985WX, eep.)
I have the complete opposite experience. Got a 9550x with 64gb ddr5 and a fairly high end mobo about two year ago. Just running the memory at stock speed. About 50% of the time I’d reboot and one or both of the sticks would only be detected as 2gb. Would need to do a hard shutdown to get it back.
I eventually gave up and turned off memory context restore and now I just deal with the minute plus (!) time to Post.
Not sure if amd memory controllers are just garbo or what but I’m going back to intel next chance I get.
What brand of ram, and was the ram listed on the supported spec sheet for the motherboard?
Up to date on bios updates?
Yes and yes, it’s crucial ram. Nothing exotic.
Weird, just had to check.
DDR5 has been kind of a mess imo, just in general.
Yea I’ve done endless googling around the issue and it seems to be a not uncommon issue with ddr5 but there’s no way I’m buying new parts now so I’ve come to peace with it.
Doesn't mean it can be broken. Also might be a CPU contact issue.
On the desktop Intel has better memory controllers last time I looked, but it’s not that big a difference. Something in your setup is just broken. Failures will happen with any brand, so I wouldn’t chalk it up to AMD vs Intel… figure out where your problem is and get the faulty part replaced if you can.
I feel you though, regardless of the cause, that is an extremely frustrating place to be.
Yea I’m just bummed because it’s a real expensive time to be swapping memory parts around. Like I said, things seem to just work without memory context restore so I’m fine with just paying the cost when I reboot once a week or whatever. I feel like the fact that disabling MCR fixes it should narrow down the issue to some part, I’m just not sure which.
It is, but isn’t it all warranty in your scenario?
It would be painful to claim though, as what part is at fault?
That's the problem, I have no idea. The fact that disabling MCR fixes it should point to something, but I'm not knowledgeable enough to know what.
For what it's worth, I have server with 128GB of DDR4 running 24/7 on cheapo AM4 consumer board. No issues whatsoever.
I went the opposite way and went back to Intel after many many years on AMD. I snagged a Xeon 654 Granite Rapids ($850) for a new workstation geared towards local inference. I don’t need a ton of CPU cores and all Granite Rapids have 8 memory channels where as you need very high end Threadrippers (9975wx for $4000) for an equivalent due to their CCD design.
As a bonus Granite Rapids is super power efficient and runs much cooler than my previous TR and Ryzens. Intel seems to be doing good things again! My only minor complaint is that P2P doesn’t work on my multiple GPU setup because every PCIe5x16 lane has its own dedicated root to the CPU, but all that bandwidth is useful for MoE models that are offloaded to RAM.
You could get a PCIe switch backplane, stick your GPUs on there and then have them communicate locally through the switch using p2p. For some GPUs you may have to use a hacked driver (google p2p 3090 for instance).
I wanted to do that, but didn't like any of the PCIe switches / backplanes I found (mostly due to price). Now, it might be worth buying one of these instead of trying to build a fancy PC. https://c-payne.com/products/pcie-gen4-switch-5x-x16-microch...
C-Payne makes excellent stuff but it is very pricey and supply is more than a little spotty, I've used pretty much all of his stuff by now and I am a very satisfied customer if I can order what I need, which more often than not is not the case.
A good alternative is ADT, they make excellent boards, there is a 4 and a 5 slot PCIe expander with a 88096 on it that works extremely well. I have two of these connected to my rig with their own power supply and four GPUs in each (and another two in the main machine). It's not exactly a portable affair (to put it mildly). Note that these won't work in a standard PC case due to the slot spacing so you'll have to rig something for that yourself (I use 2020 + some custom 3D printed fixtures).
Yeah, I had to go high-end to get one with enough CCDs to make use of the memory channels. I'm curious to know what all-read bandwidth you see when you run intel mlc on your system. My Xeon W5 was doing ~180GB/s after a lot of tuning, which was very disappointing. The Threadripper does ~320GB/s. From what I've heard, Granite Rapids is supposed to have solved the bandwidth problems that ailed Sapphire Rapids, so your numbers ought to be similar.
I’ve been running the 64 core (128 logical cores) Threadripper PRO 3995WX in my dev machine for 4 years now (with 256gb ram).
Not sure I’ll need to upgrade my computer ever again :D
Threadripper is underrated for IO-heavy boxes. I made the same mistake with RAM prices, though; should have maxed it out when it was cheap
Ye, that's like saying one should've bought Amazon, Apple and Google stock when it was low.
That is only true if you only look at the market value of an asset.
One is a tool, the other is a speculative financial asset.
That's why I'm all in on hammers. Garage full of them.
Different kinds too? You never know when you might need a ball peen hammer or a soft faced hammer or a rubber or deadblow mallet or a framing hatchet. Why take the chance?
They bought the ram they needed when they needed it, and probably some extra too. Buying more than necessary is inherently a speculative action - you are betting on needing it in the future, and you are betting that it will be more expensive in the future. If it maintains its price it doesn't matter if you buy it now or later, if it gets cheaper or gets superseded by better alternatives then buying early is essentially a loss.
So in my view, this situation is almost exactly analogous to stock market speculation.
Hindsight remains 20/20, though.
In the computer world, wait a few years and things get cheaper. That was broken by an exceptional event. Who would have thought you can't afford DDR4 in 2026?
The upcoming shortage was public knowledge before the prices rose.
Just like an investment though - you need to have the money to play. I wasn't buying a computer at the time because [snip a long list of other toys I bought instead].
My current plan is to stick with my existing computers for a couple more years and hope this blows over. Time will tell.
We seem to be going back to the days when desktop PCs had a little lock hole.
They (in theory) could padlock them so the drives and ram couldn't be stolen?
https://francisuniverse.wordpress.com/2017/10/07/the-turbo-b...
I'm aiming for my machines to be too heavy and hard to open ;)
I put 256G in it wondering if I would ever need that much RAM and now it looks way too small. Highly frustrating. I wonder if the companies that made these deals realize that they've just given an entire industry a reason to look for alternatives. It's not like they wouldn't have sold their product otherwise.
I picked up a pre-ai-price-insanity AX162-R at Hetzner a while back and loaded it up on memory to max out the 12 channels the 48c EPYC 9454P as a "this will be the last mysql box I'll need" and have I been _wildly_ impressed with it's performance. The things I throw at it are honestly laughable at times, wildly irresponsible queries against a rather large database, the redis qps metrics are ridiculous and I just load it up with random ggufs since the memory bandwidth is... not terrible and 384GB of it is... useful.
The web app that's hosted on it deals with lots of images and text - over 100mil of each deduplicated, embedded, simhashed - it does the hashing, and the embedding in real time during ingest. Just handles it.
These things are absolutely insane.
i have 2 epyc boards from 2019 and they are still just printing money
Weird they didn't talk about the DDR4->DDR5 transition. The 5800X3d->7800X3D went from DDR4->DDR5.
DDR5 is literally twice the price per GB. It's better for sure, higher bandwidth but waaaay more expensive. The 9800X3D supports much faster and even more expensive DDR5 than the 7800X3D too (DDR5-2667 to DDR5-5600).
There was a brief and glorious time before September 2025 where DDR5 was actually pretty comparable.
Even really good DDR5 could be had for cheap; I remember paying like 300$ for 96GB of high-quality M-die hynix DDR5.
July 11 last year I got a Crucial 128GB kit for $299. Those were the days…
DDR5 has always been twice as expensive as DDR4. But yeah when they were a lot cheaper maybe it didn't matter as much as it does now.
Ram is getting better, faster etc. except for the latency. I guess we still cannot bend the rules of physics.
Latency has roughly remained the same since DDR2, no? So that has just about nothing to do with the improved performance of every generation.
Latency had to improve because of the increased CPU speeds, but it didn't, causing a regression that's easy to perceive.
Zen3 is 7nm, Zen4 is 5nm and Zen5 is n4p. That's also why Apple is ahead; TSMC is a big factor.
I think there is more to it than that.
It's not actually 2 years. It is 4 years of progress.
Zen 3: 2020 release
Zen 4: 2022 release
Zen 5: 2024 release
The author is disingenuous by looking at only the X3D variants. The Zen 3 X3D variant came out late.
For comparison, in the same time frame, the M4 is 52% faster in ST and 72% faster in MT than M1.[0] If you go by what you can physically buy in stores since Zen 3 was released to now, Apple's chips have gotten 83% faster in ST and 143% faster in MT.[1]
[0]https://browser.geekbench.com/v7/cpu/compare/433673?baseline...
[1]https://browser.geekbench.com/v7/cpu/compare/435139?baseline...
It's funny how my 3950X, which I still find to be an absolute monster for anything I throw at it, isn't even in the benchmarks anymore. Hard to believe it's 7 years old.
Perhaps better looking at it without the 3D Cache on GB6 scores.
2020 Zen 3 AMD Ryzen 7 5800 - 1852
2024 Zen 5 AMD Ryzen 9600X. - 2914
2027? Zen 6 - Est 3200?
And compare this to Apple.
M6 4000+
M5 3600+
M4 3300+
Even the Snapdragon X2 Elite X2E is better at 3300.
The gap isn't exactly shrinking. And I would love x86 to prove me wrong. But even a hypothetical Zen 7 in 2028 ( 2029 for consumer ) may only catch up to Apple's M5 released in 2025.
So far, AARCH64's advantages mostly seem to affect single core performance. AMD is also behind in the nanometer race, which is one of the ways Apple managed to improve their peformance over the years as well. If they manage to catch up in terms of die shrink, I think their multi-core perf will once again put AMD in the lead for data processing at the very least.
Although, if you count EPYC, which you might as well with the prices Apple charges for their top-end models, AMD still beats Apple with ease at multi-processing.
It's a real shame that these processors are made by Apple (and Qualcomm, I suppose). They'd be my go-to chip if they weren't.
The AMD EPYC Zen 6 is built on TSMC N2, so it is a similar comparison including clock speed, x86 just hasn't delivered.
Apple has been iterating on a yearly basis, A17, A18, A19 and A20 are all different core design. Some of them just so happens to come with new node. On the other hand Zen 5 to Zen 6 took 24 to 30 months.
With CPUs becoming wider and wider, it's almost like they run VLIW instructions, only more flexible.
Huh no, wide vector units or lots of execution units per core is nothing like wide instruction words? Superscalar out-of-order execution is specifically a way to get instruction level parallelism without the wide instruction words
I was thinking that for this new generation, "Superscalar" doesn't cover it anymore, except as a generic term for the design technique. "Ultrascalar"? "Superscalar Pro"?
I mean superscalar just is a different name for using wide cores and runtime dependency tracking and out-of-order execution for instruction level parallelism without VLIW.
To my eyes, your message looks kinda like, "The term 'bicycle' doesn't cover it anymore, things have changed so much since they were introduced. We should call them something else, maybe 'bicycle pro'? 'super bicycle'?" No it's just a bike, and no, modern CPUs are just CPUs which look scalar at the ISA level but have ILP
Sure, but in my mind it feels like the CPU is slurping up individual instructions until it has enough to execute a whole bunch of them at once in one step (which would be equivalent to a big VLIW).
It's not really like that, it's more that it slurps up individual instructions, executes some of them just in the front-end via register renaming and dispatches the rest to available execution units, then retires them as the execution units are done. Sometimes it fuses together a couple of instruction (macro-op fusion) but my understanding is that that's much more limited, like fusing "compare then conditionally branch" sequences into single compare-and-branch μops
I was recently able to get a workload of mine optimized enough on Zen 5 to hit a sustained 6.0 IPC/core (3.0 / thread) at 5.1 GHz. Seeing the a > 99.8% branch prediction rate and a > 99.99% L2 cache hit rate retiring > 1T instructions every 3 seconds feels amazing.
Zen5 is incredible when you're able to make the most of it. I’m super excited about Zen6.
zen5 is really the first CPU of the avx512 era (since it's the first time normal compiler devs have them to play with)
And it’s full width for the first time.
We had Ice Lake and Rocket Lake too, they seem to have disappeared from the collective memory
Their successors had more tangible performance improvements in everyday applications, so yes they were quickly forgotten. Also, it was a reduction in max core count from 10 to 8, intel changes sockets too often, and 14nm+++++.
Can you share any more details about the workload? Always interesting to hear of something like that which isn't a useless microbenchmark.
Does it use vector? What can you hit with SMT disabled?
It’s a backtracking search program looking for integer solutions to a specific problem. I’ve tuned a series of bloom filters to fill my L2 cache so that I rarely have to touch main memory (this alone took my IPC from 0.1-0.3 to 3.0 per thread). Without SMT it’s 4.6 IPC/core.
I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.
coded entirely or partially in assembly, i presume?
Entirely in C. Compilers are pretty good these days. Almost every time I try to outsmart them I make things worse.
I work on game servers. I can only dream of such an IPC, but it doesn’t stop me trying, heh.
Array-oriented programming. Add all velocities to all positions, etc
For someone who has never touched workload optimization, can you share a little about how you measure this? Does IPC mean inter-process communication in this context?
I know nothing about workload optimization, and I'd like to know more - right now I feel like the Good Burger gif. "Yeah, I know some of these words".
Instructions Per Cycle
I've never done workload optimisation myself, but have looked into it just-in-case. It's a really interesting field and unfortunately, you need to know about both your compiler and target CPU/GPU to get the big gains. I believe most modern CPUs have registers for cache use. You have to do some guesswork to get IPC numbers, since core frequency can vary.
Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).
Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).
This field is super deep, it's very fun to learn about!
Though you should always ask what your compiler/optimizer is doing for you. Research into how to turn your ifs into arithmetic is ongoing. If the compiler can optimize your simple ifs into the complex arithmetic then you should go with the simple if and let the compiler do that. (unfortunately in many cases the compiler cannot be sure because of some special case - even though odds are you don't care about it)
Thank you. Does workload optimisation usually refer to the CPU/GPU alone, or would it include disk or network IO? Or would that be more accurately called something else, like profiling?
I'm not too familiar with the CPU world of performance measurement, but in GPU land, and I suspect for CPUs too, there are a slew of hardware "performance counters" which are registers that increment every time the event they measure happens.
So there could be a "instructions completed" counter and a "cycle" counter, and before starting the benchmark, you record the current value of both counters, then after the benchmark, record the final values, and compute the Instructions Per Cycle (IPC) as the difference in instructions completed over the difference in cycle count.
As for optimization, at a high level it's about minimizing the amount of time any part of the CPU is waiting for other parts of the CPU. The specifics require a lot of background knowledge about how modern CPUs work, more than can fit in a post, but if you're interested, topics to read about include:
### CPU cache hierarchy
CPUs store copies of data from RAM in smaller, faster memory physically closer to where the computation happens, so it's available more quickly. The CPU decides what values to store in the cache, and gives only limited control to the program, so an optimal program needs to be careful to not make the CPU make bad caching decisions (including for synchronizing the cache between multiple threads of execution).
### Instruction pipelining and out-of-order execution (a.k.a. "superscalar" execution)
Modern CPUs operate like an assembly line. A new instruction can start executing before the previous instruction(s) finishes. CPUs also have redundant hardware, so multiple instructions can be in progress at the same step of the pipeline. But there are limitations; sometimes the input to one instruction depends on the output of the previous instruction, so the whole pipeline stalls until the result is ready. An optimal program orders its operations to avoid these stalls as much as possible.
### Branch predicition
When the code to execute depends on the result of a computation, like in an `if` statement, we call it a branch. Branches can stall the pipeline, because the CPU doesn't know what instructions to execute next until the current result is ready. However, to mitigate this, modern CPUs predict which code-path will be taken when a branch is reached and begin executing the associated instructions immediately. If the prediction is right, the pipeline stall is avoided, but if it's wrong the pipeline state has to be restored to what it was before the wrong branch started executing, which is even more expensive than a stall. Usually, the CPU predicts branches correctly, so branch prediction is a net gain. Optimal programs need to understand how the CPU makes these predictions and make their branches as predictable as possible.
### Single Instruction, Multiple Data (SIMD)
Some programs do the same operations to each item of a set of data. CPUs have so-called SIMD instructions to accelerate this by, unsurprisingly, performing the same operation to multiple items at once. For example, they could add 4, 8, or even up to 64 pairs of numbers at once, depending on the instruction set and the range of the inputs. Suitable programs are optimized by arranging their inputs and operations so that SIMD instructions can be used — either directly or by being written in a way that a compiler can translate individual operations to SIMD operations.
Thanks. Is there any convenient tool you'd recommend for watching those registers?
I've said in another comment that any optimisation I'd be doing would be looking at the entire stack, and CPU/GPU optimisation would only be as a learning exercise. I'm a tiny bit familiar with caching and instruction pipelines thanks to college, and branch prediction thanks to Spectre-class bugs. I want to find some time to dive deeper, now.
I’m doing all my CPU measurements with ‘perf stat’ on Linux.
If you're on x86-64, the Agner Fog manuals are kind of the bible. Free and he's been keeping them up to date for years. ("Copyright © 2004 - 2026. Last updated 2026-09-10.") I don't know how he does it. And even if you're not on x86-64, you can probably still learn a lot.
https://www.agner.org/optimize/
in the context of CPU performance / architechture `IPC` pretty much always means "Instructions per clock". That's a measure of the internal parallelism a given CPU core is achieving. Most of the time code is not able to get very close to the theoretical maximums a core can achieve for very long, so the numbers the OP is quoting are incredibly impressive.
Because nobody cares about CPUs anymore and they thought now was the time to flex those numbers that were always possible? Because all eyes are on NVIDIA GPUs?
The article compares 2022 CPU against a 2024 one.
CPUs did get 50% faster in two years. It wasn't the previous 2 years though.
All I get is a 403 forbidden when I try to access this page.
I've got a 9950x3D2 and it's pretty amazing. Most reviews are kind of ill-guided, as the CPUs in consumer class get graded by their bang-for-buck in gaming, though.
I have a 9900X (Zen 5) and I am able to clock 5,000,000 WebSocket localhost message sends and response receipts per second (single socket) on it from my WebSocket library written in TypeScript.
Lol at the obvious Intel CPUs used on the headline
Mainly iterative IPC improvements combined with maturing process nodes. That initial TSMC 7nm was solid but had room to grow.
And it shrank.
No mention of Moore's Law...
From https://en.wikipedia.org/wiki/Moore%27s_law
Moore's law is mostly dead these days. Process scaling still happens, but nowhere near at the rate predicted, and there are tradeoffs everywhere now. The Zen 5 core is in fact bigger in terms of die size too.
Is it? Without Googling to backup my intuition, I thought that Moore's Law is still alive, but people just get Moore's Law wrong. A lot of people think it states that PERFORMANCE doubles, but it's number of transistors.
And while in the past, the doubling of transistors DID double performance, it's no longer true.
Moore's Law is getting accelerated