- 23comments
- 18comments
- 81comments
- 50comments
- 44comments
- 7comments
- 29comments
- 143comments
- 196comments
- 802comments
- 12comments
- 114comments
- 1034comments
- 82comments
- —discuss
- 57comments
- 3comments
- 68comments
- 417comments
- 271comments
- 6comments
- 232comments
- 584comments
- 175comments
- 64comments
- 537comments
- 186comments
- 160comments
- 37comments
- 32comments
Yeah okay bud, anyone checked in with the state of consumer hardware recently? Not the author, evidently.
Yeah I'm sure Samsung, Nvidia and sk hynix will all be very calm with lower volumes and lower margins.
RAM prices will crash when demand drops even a little. They'll probably crash to a lower (inflation adjusted) level than before. This has happened before.
Industrial scaling in general often looks like a sawtooth: price spike, capacity investment, crash, repeat.
Part of what's keeping prices high a little longer is that everyone knows this and is a little reluctant to plow resources into chip fabs for fear of having the bottom fall out before they recoup or sell that to someone else to hold that bag.
Graph the average compute and RAM in a mid-high end laptop at an inflation adjusted price point for the past 40 years. It's very exponential and hasn't slowed down much.
except cxmt who is plowing resources in like crazy
That has more to do with geopolitics than it does the current price of memory.
No. Prices will crash when supply side expands to meet the increased demand. Because demand won't go down to pre-bubble times any time soon. Unfortunately the supply side has been very slow in increasing production, partly because most steps of the production chain are all maxed out.
On a long enough scale you are right that prices will likely normalize to a better level, but before 2030? That would mean the factories are built quickly once they begin.
The entire reason why hardware prices are so absurd right now is because manufacturers across the board are doing everything to prevent that crash.
They all collectively chose NOT to increase supply with increased demand. So if the bubble pops, they just go back to previous prices without oversupply driving the prices to rock bottom.
The article observes that the cost of frontier intelligence from 2025 has fallen 100x in the last year. It also notes that the energy to run models is also collapsing. Consumer hardware is borked right now because these new algorithms are revolutionizing the utility of a computer. Computing is technology who's cost has been collapsing for 90 years, and its a safe prediction that it will decrease again.
Seems like 3-6 years is a pretty conservative timeline to me. Memory and inference chip production could ramp up massively in that amount of time.
This is the core of my belief that data center construction is a huge bubble.
AI is not a bubble, IMO, though we may see a retrench and some companies with sky-high valuations will crash to more reasonable ones. But data center demand is probably a bubble, and the main driver will be reduction in the actual amount of power and data center space required to serve escalating demand.
I think hardware and model improvements will pace or maybe outrun demand and then when demand starts to saturate will keep going and leave a lot of orphaned data centers.
Jevon's paradox says that if data centers can serve a lot more tokens per dollar or watt there will be increased demand for data centers.
Jevon's paradox isn't a physical law, it doesn't magically apply to everything. Millions more copies of Atari's ET game didn't cause everyone to pickup a cheap copy, and cause extra demand for a garbage video game. Some times (actually, usually, I'd argue) things are made that will sell for less than the cost of construction because of irrationality, and they don't induce extra demand and they don't change the negative profit margins.
You can't simply wave Jevon's paradox at things. Thousands of miles of canals were dug in the UK that couldn't be sustained and were abandoned. Thousands of miles of railways were laid that could be sustained and were abandoned. And those are potentially durable investments, unlike cheap walls, pillars and roofs laid over a levelled concrete slab full of fast depreciating IT equipment.
It's true that Jevon's paradox doesn't always apply, although this does seem like a classic case.
But yes, if sold for a negative margin Jevon eventually stops because the decreasing supply will drive up prices.
Price is set at the marginal cost. Capital costs aren't in marginal costs.
You'll need a better counter-example than UK railways which suffered from Parliament price-fixing.
I agree, but I would say what the parent is saying is more akin to Say's law: https://en.wikipedia.org/wiki/Supply_creates_its_own_demand
I'm so glad the tide here is turning on this talking point, brought on by exactly the same people beating us over the head with it for months while no progress is made towards it materializing.
Many, many people who post here are capable neither of real analysis nor distinguishing real analysis from memes. They aren't hackers, they are adherents of a cult that happens to focus on the same subject matter as hackers.
Jevon's Paradox, RSI, revealed preference
Jev, ASI, RLHF, Jev, Jev, water usage
Jevon's paradox only applies to products with near infinite demand. Energy being the most famous example. I dont think its difficult to argue that compute/intelligence is also a base input into the economy and theres almost no limit to the amount of intelligence the world will want.
When people say "AI is a bubble", they mean economically as a whole, which includes data centers.
Perhaps we need better terminology for "product useful; numbers nonsensical"
I disagree. I own over 1TB of vram at home. I can tell you that it's not a bubble. From my builds, I would rather have cloud, cloud is easier. From running small models like Qwen3.8-27B to large models like Qwen3.8-2.4T. I can tell you that small models will never be enough or match up. Everyone will want the smartest model, not just a good enough model.
I think Nvidia is under the same pressure as Anthropic/OpenAI. Nvidia will dominate research and probably keep dominating training, but the real volume is in inference. And for inference Nvidia's lead is only a few months, similar to the lead frontier labs have over open source. Nvidia will sell a lot of Rubin CPX's, but their margin on that will be a lot smaller than B200 because there is so much more competition in that space.
for the first chart (sourced from https://epoch.ai/data/machine-learning-hardware?view=graph&y...), what is the audience supposed to think about that trend line? there's a step after you slap a regression on some points where you evaluate whether there's a real trend or noise, right? i don't see that either in the article or the linked source.
It is hard to see how the environmental side effects of this aren't going to be somewhere between bad and disastrous.
No it'll be fine as long as you do your part and not drive a car, or have AC, or eat meat, or have children, or live in detached housing, or...
I think this is a case where just drawing a "line goes up" extrapolation is incredibly misleading because there is _tremendous_ economic pressure to get costs down, and costs are very tightly tied to energy use. All of these systems are incredibly inefficient right now and have a lot of room to go down in energy use. I'd guess that the absolute _floor_ is burning model weights directly to silicon and that's like a 90+% reduction in energy use.
I found the OP insightful and worth a read. Thank you for sharing it on HN.
The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."
The OP glosses over questions of business model viability with a brief qualitative discussion and very little hard data. For example, to earn an annual return > 10% on every trillion dollars of capital sunk into infrastructure, the owners of that infrastructure must earn free cash flow (operating profit less investment) in excess of $100 billion per year in perpetuity. Is that feasible? Why? How?
The OP does not really consider such questions.
Exactly. "Cost-to-distill" is a critical parameter. Right now usage of frontier models for all tasks is both subsidized and irrationally popular even at the subsidized price. Deepseek would solve most tasks faster and 10x cheaper. I agree with the author that just as Deloitte exists, frontier labs will exist. But not because their products are proprietary technical marvels or gods, but rather because of branding.
DS wouldn't be 10x cheaper than the subsidized subscription plans from openai/anthropic. Although it is of course much cheaper than the enterprier/API pricing- I think if you're on the subscription plans, you can't beat that on performance per price.
They are already turning profits and inference has shown to be a cash cow. And they've already secured compute for the next several years.
Some frontier labs are reporting positive "adjusted EBITDA" (earnings before interest, taxes, depreciation, and amortization, with extra adjustments to make the figure positive).
Free cash flow (operating profit less investment), actual cash coming in, is deeply in the red.
EBITDA can be a sensible measure of profitability when there isn't much need for additional investment. That doesn't seem to be the case with these operators. They need to invest aggressively to avoid losing customers to competitors. All of these operators have made multi-year commitments to invest more in infrastructure. In addition, they have guaranteed quite a bit of debt to fund it.
Maybe it all will work out fine (and I sure hope it does!), but I didn't see any hard data from the OP, or from you, supporting that view.
EBITDA might make sense for the resellers who package up open weight models and sell inference. It is not appropriate for the labs who have billions in debt for RAM, new data centers, gobbling up competitors, etc.
Those real debt obligations are going to want to be paid back.
Who is the "they" that are turning profits?
Definitely. The question is: Is it enough to recoup the enormous capital costs and justify the level of investment they've received. I think there's a decent chance that it will be. But maybe not. And the longer they keep focusing on training new models more so than on inference, the more uncertain I become that it's all going to work out.
Labs are playing money games with EBITDA, which is not uncommon, but also hides the extent to which they are in the red (deeply, deeply, in the red, and projected by them to get worse).
I think the article's analysis is basically right in a vacuum. That is, I think it's clear that inference is a viable business model. But what isn't clear is whether it will be such a profitable business model for any given company that it will justify the investment that company has taken. I kind of think the winners might be a follow-on generation of companies that focus on this commodity inference business model instead of the invent-machine-god-first "business model" and thus are wiser about their level of investment and capital costs.
You may be right. I'm not so sure. Inference looks like a viable business model for those operators that have SOTA infrastructure in place, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its infrastructure quickly becomes non-competitive, and customers will quickly leave for alternatives. SOTA infrastructure is a moving target.
Maybe... I'm not enough of an expert on the financials to say, but it seems to me that inference should be able to recoup the cost of SOTA infrastructure, unless you then also use a large portion of that infrastructure to train new models. And I also think the race to remain SOTA itself is also largely a function of training, because my understanding is that training benefits more from the leading edge of hardware.
But yeah, I definitely don't have high confidence in any of this!
A phrase comes to mind: "Your margin is my opportunity."
The problem source is that the "cost" of tokens are taken at face value from business that are losing money at record speeds. E.g. https://artificialanalysis.ai says "doing task A costed us $10 using OpenAI", and that is the "cost" the OP used as basis for "tokens are cheap". Meanwhile OpenAI is losing $19 for each $1 in revenue... So right now OpenAI should be charging around $200 to do task A just to break even, but that would mean their use base would collapse.
The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)
Right. And some hardware improvements will speed up both grep and Luna, which won't close the gap.
not necessarily, one may be easier to parallelize while the other suffers some serial computation bottleneck.
Grep is trivially parallelizable if you care to do it though.
https://iepathos.github.io/ripgrep/performance/#work-stealin...
It might never beat out grep, but it could beat some more expensive to call tools, similar to how heuristics will often be faster than exact answers. Rust Analyzer can be slow at times, I could see an AI tool taking over a subset of its work.
LLM is spicy memoizing, so it can potentially be faster than a tool call. But people will spend a month tweaking and testing to ensure they have the level of determinism they need, which means it's more expensive, and that they should have used actual memoization in the first place.
Yeah I bumped on that too. If it's possible to make llms cheaper than current grep, then it is also almost certainly possible to make grep cheaper.
You can burn anything* into an ASIC to make it cheaper per-call.
non-backreferencing grep is not very difficult to implement in an ASIC either. But it's probably not worth it because of how relatively rarely you use it and of the data transfer costs.
LLMs are great candidates for ASIC-burning because they're slow compared even to network speeds and run all the time. The issue is that you don't want to burn a specific model or architecture that then becomes obsolete.
So you've got two possible futures, and both guarantee large price drops: (a) LLMs keep getting better and better and better, so ability/$ keeps rising; or (b) LLMs plateau in ability, in which they will start getting ASIC'd.
This whole story really reminds me of crypto coins. Like.. going from mining one coin, or lets say token, to millions of fractions like 0.00000000001 bitcoin a week.
There is a reason it took 4 and a half billion years for human intelligence to develop.
???
Yes, because it was a largely random undirected process.
Even in directed processes the low hanging fruit is harvested first.
Then you have to climb to another branch to get more fruit. The biggest issue with most problem space discovery is you're doing it blindfolded.
They don't need to plateu for that to happen. There are companies already building AI on ASIC, and IIRC they were approach 12 months lead time. A 12 months old frontier model (Sonnet 4.5, GPT-5, Kimi K2) for 1% of the price is still a rather good value proposition.
To be clear, I agree with the overall premise of the article!
But I would probably take a long horizon bet that the grep implementation on my machine will remain cheaper than an equivalent ai task, even though I think those ai tasks will become far cheaper over time.
I just think the original comment's model of asymptotic approach is probably more likely to be accurate than the model of the line blowing through this grep-like cost level.
depends on what you are grepping ... greapping a large file might be more expensive one day than generating n-th token with LLM that works fully in hardware
you could make hardware implementation of grep and store the file itself next to it in some ROM but that's not a very useful grep ... while hardware LLM is exactly as useful as software LLM only orders of magnitude faster
Yeah it's a pretty poorly specified problem. It needs to be some kind of "equivalent task", but it's not clear how to define that.
From a computational standpoint this is obviously nonsense, but from an attentional one I'm not so sure. It may already be more attentionally expensive to use grep in some cases, such the moment you need to remember a non standard arg. And if this applies for performing a simple http operations, then it certainly applies going up the complexity chain.
"Did you know that disco record sales were up 400% for the year ending 1976, if these trends continue...AY!"
AY-3-8913?
Won't that also help grep and then move the asymptote down more?
At some future point where LLM hardware is cheaper than simply running grep, then grep equivalent would benefit from those selfsame hardware improvements and be cheaper to run as well, probably still by the same ratio.
There seems to be a mistake in the cost comparison between 2025 and 2026. The 2025 chart axis is the cost to run the entire "intelligence index", and the 2026 version is a weighted average cost per task.
I don't disagree with the thesis here, I just don't think costs are coming down quite that quickly.
So this week's coordinated shock and awe campaign is about "open models" and "prices are falling rapidly". Also on the front page:
https://news.ycombinator.com/item?id=49815526
GPU case seems very weak. The graph is impossible for me to reason about at least. You could draw basically any trend line through that GPU graph and it would look equally plausible to me. The main takeaway I get is that the NVidia H100 from four whole years ago is barely different in efficiency from the state of the art, which is surprising to me, and seems to indicate the exact opposite of what the article says.
A much deeper analysis on the falling price per task was published yesterday by Epoch AI [1]. It's a real statistical analysis and comes to more defensible and grounded conclusions. The headline takeaway is:
The cost of a given level of performance often falls fastest right after that level is first achieved, that is, when it is state of the art (SOTA). We see this pattern on three of our five main benchmarks of AI capability. Averaging across all five, cost falls 66% per quarter (75× per year) for performance that has just debuted as SOTA. Two years later, prices fall half as fast, at 32% per quarter (4.7× per year).
but the analysis itself has more nuance and is a quite interesting read.
[1] https://epoch.ai/publications/the-plunging-price-of-thought
Too Cheap to Meter reminds me of the promise of Nuclear Power in 1954
"It is not too much to expect that our children will enjoy in their homes electrical energy too cheap to meter,..." Lewis Strauss
https://en.wikipedia.org/wiki/Too_cheap_to_meter#Origins
Oddly enough my power bill was metered and big.
It's really quite unfortunate that the promise was not delivered, mostly for political reasons. I hope that a new wave of reactors and the dire need for clean energy restarts the nuclear race.
The economics are not there. Solar + batteries are much cheaper per watt today and still improving. Even better, the solar can come online instantly and expand while US nuclear takes twenty years to start generating any energy.
I strongly agree. We _will_ get there with solar, wind, and batteries.
Eh, it looks like micro reactors are coming on line much faster these days.
Nuclear was never going to give us too cheap to meter regardless of politics. Uranium just isn’t that cheap. Fuel costs are lower then coal or gas, but not so low that operators would just not bother to charge for it
It's a deliberate reference/meme that is basically used to acknowledge the precedent of overly exuberant predictions of cost in an emerging technology but argue "however, this time it's true".
Of course, perilous territory for future irony depending on how your prediction plays out.
On the other hand, this did work out in other areas. I pay a flat monthly rate for all-I-care-to-eat internet access, for example. My email provider has limits on storage space but I don’t get charged per email sent or received. It’s not a crazy concept on its face, nuclear power just didn’t work out as well as was hoped.
We have people suggesting that ai is so costly to run that all labs are secretly subsidising tokens and we can expect a reprice soon.
Then we have these articles that say tokens will get so cheap that labs won’t know how to make profit.
Who is correct?
They're not secretly subsidizing, they're openly subsidizing.
Token pricing was a small minority of customers up until this year, when all the labs started trying to force customers onto token-based billing. Within the last week, Anthropic repriced my team's plan from a temporary "50% extra tokens" to 25%: https://support.claude.com/en/articles/15910845-claude-code-...
The fact that all this is ongoing within such a short timeframe should make you suspicious of any analysis that claims to be observing "statistical trends" like they've discovered a new Moore's Law out of 6 months of pricing data from 2 companies.
your repricing has nothing to do with subsidising which means selling at a loss. Within this year, the real prices have gone down more than 10x on average which is way more than the teeny 25% you are fighting for.
The labs themselves when they openly say that they’re subsidizing tokens, I’d imagine.
where? I mean the API costs
Not really related to the central point, by but I couldn't help but get caught up by
That is such an interesting set of models to use as examples here. One being essentially obsolete on release a month ago, and the other being completely ancient in LLM time. I really wonder how they landed on those two.
It's true that LLMs "want" to be be local, but they won't shift broadly to being local until there's a sufficiently large supply of VRAM or (at least) "unified" memory from the manufacturers. (I'm also assuming here that radical regulatory changes like government bans of local models aren't going to happen.) So (AFAICS—I am no expert) the future of LLMs over the next few years comes down primarily to the nitty-gritty of how much memory fab capacity will be added and when, and to a lesser extent of what happens to future demand from LLM SaaS services (& maybe their existing stock of hardware if they get in trouble). (I'm also assuming no roughly-AGI-sized leap forward which makes the frontier models of the near future vastly more valuable than the near-fontier models of today.) For the incumbent manufacturers the high-margin business is selling to LLM SaaS providers who use VRAM efficiently, but the high-volume business is getting chips into millions of laptops which will use VRAM very inefficiently. I assume that they will want to move from high margins to high volumes as they build they physical capacity to ship higher volumes, but they seem to prefer to do it at a stately pace. Hopefully some jostling from Chinese competitors, and maybe a dropoff in demand from data centres, will speed things along.
I agree with OP that we will continue to see improvements, but there are also some serious bottlenecks ahead of us:
- Energy is not infinite, neither energy efficiency is. - Datacentres neither. - Benchmarks are an abstraction of real world problems!
On top, there is an overall "economic" aspect that most of the people miss: every change carries a certain degree of risk (lose money, reputation, customers, death of people, ecc) that very few want to take and a lot of changes(e.g. rewrite some piece of SW in another Lang) don't produce a positive economic impact.
Of course, for collecting better telemetry using local AI for analyzing video from camera and audio from a microphone.
jgrep is not meant to replace grep. It’s supposed to extend it to new use cases. Keyword search will always have a place!
Improvements that affect local AI - Mamba...
I just stopped reading at that, for anyone else, Please find a better source and take everything in here with a grain of salt.
IMO
The number one improvement that mattered for local AI was llama.cpp, partial offloading to system cpu/ram. The next was quants, being able to take fp16 and turn it to q8, q4 etc. The next IMHO is unsloth dynamic quant, that have been able to do mixed precision so we have UDq1/q2 that is actually pretty damn coherent. Allowing individuals to drive K3 locally even if it's at Q1/Q2. Then MoE changed everything for everyone, cloud and local. The other is integrated GPU, Apple, Strix Halo, DGX Spark. Then all the extra improvements like MTP, DSpark, etc. Of course there's many other additional things that have mattered too