I'll believe it the day SpaceX - the AI company that is mostly making money selling datacenter compute using gas turbines for energy - takes a single day 90% or greater drop in stock price.
Currently nobody knows when the first big financial crisis is fully locked in. For example, if OpenAI can't close another round, and they default on their contracts with Oracle, there's your sign. Until then it looks like everybody is enjoying the communal hallucination.
I mean the first AI financial crisis at an AI company. I did provide an example. First, the investor cash has to dry up, then the AI data centers have to start getting itchy about what those multi billion dollar contracts are actually worth.
They mean the roughly 25 year tech sector run that's now culminating in the irrational exuberance of AI. There have been some hits already and the result of those is market cap consolidation of the largest companies. Not sure what the number is but the the 10 largest companies make up a huge percent of the entire S&P and most have extreme exposure to the same risks. The sector has been boosted by the NVIDIA circular financing as well but at some point there is a limit. Although they are angling for a pre-bailout with all the AI is going to kill humanity fear mongering. The only savior for the sector will be the government, the question is if the government steps in before or after an organic collapse.
They could potentially do another private bridge round, but for a company that was gearing up for the largest IPO in history a couple months ago, the reversal is a pretty bad sign. For investors that are looking for a fire sale, there's already smoke in the air.
Probably the next nvidia chip drop because the old chips where priced at their highest level at a depreciation of 5 years instead of the 2 year norm. What happens when the next chip is way better(hearing 67x better, not just the 10x from initial claims)? Well all those chips need to be dumped fast and depreciate that second and the datacenter gets bankrupted and parted out
We're already seeing repercussions from an economy that has been retooled not to actually produce anything of value, but to produce more air to fill up the largest economic bubble in the history of the world.
National debt through the roof, inflation through the roof, PHD and research programs gutted, non-ai startups dead and unfunded for the last 4 years. These are just a few things that have been sacrificed on the altar of this bubble - there's far more I haven't recounted.
We're already in a widespread long term economic collapse, but the delusion just hasn't broken yet.
Any data center that can remain profitable selling open source tokens at commodity prices will be fine. Any data center that relies on OpenAI/Anthropic level token prices and margins might be in trouble. After clearing their debts through bankruptcy, they'd likely be quite profitable selling open source tokens at commodity prices.
I think there's still some low-hanging fruit with thin clients and colocated-to-AI applications. I don't think the math for beefy personal computers is going to hold up, for most use cases.
If it costs you more to generate the tokens that the market is willing to pay for those tokens, then not even bankruptcy will save any of the costs invested in one of these datacenters.
If the cutting edge OpenAI token prices are $80 per 1M token, and the open source tokens are $1 per 1M token, that's a huge gap of "this will never be able to make money under any scenario if the bubble bursts" that will catch a lot of these new datacenters. No one will run a datacenter that costs $5 per 1M token to sell at $1 per 1M token even if the debts are cleared.
The point being made here is that most of those costs are amortized capital costs, which get wiped in bankruptcy.
That $5 per 1M token doesn't literally cost $5 per 1M token. It's more like they had to build a datacenter for $500M that can service 100T tokens over its lifetime. They did this by borrowing money on the capital markets, and now they have to pay interest to those bondholders, interest that they can recoup with their $80/1MT prices. But if it turns out they can't charge $80 and have to charge $1, they won't be able to make those interest payments. They enter bankruptcy, the court wipes the debt clean, and now they don't have to pay interest, only the actual operating costs, which may be more like 50c/1MT. The company gets recapitalized with the new owners being largely the bondholders, the existing equity holders get wiped out, and they can compete with the commodity producers now.
You have land taxes and or rent, building upkeep, staffing costs, electricity, water, hardware replacement costs.
And new build DCs have blown all these costs through the roof justifying the decision because the price of compute is so high. When the prices come crashing down, the expenses will remain fixed where they are now.
After bankruptcy they likely have no debt, so can out-compete those that didn't go through bankruptcy. It's the bankruptcy that makes them profitable -- it's a common pattern in nascent commodity industries.
By the time the bankruptcies come, they'll all be owned by index funds and main street. The price will crash at the break of the bankruptcy, and wall street will swoop in to buy up the "distressed assets" at bargain bin prices.
Bad news for wall Street though if they have to go bankrupt first?
Don't worry, the already stretched taxpayer will be on the hook for everything just like in 2008!
This time the relief mechanism is already baked into the system (capital does learn from its past mistakes, even if it may not be the lessons you'd hope for!)
assuming demand remains elevated and growing, maybe. But spend on AI is pretty stratospheric right now... companies are already starting to clamp down on spend. This makes you really wonder if there will be sufficient demand at current commodity prices for eg; OSS models to justify all these data centers.
Yes, I believe so. Using a sibling commenter's number of frontier models being 80X the price of commodity models, I think companies who switch will spend a minority of their savings to increase their token usage and only pocket the majority of the savings.
Eventually these models will get commoditized (we're already at the "good enough" stage for real work), then they will get turned into custom hardware and get 1000x faster, then that hardware will get commoditized (like DSPs) and they'll be everywhere and cost $1.
Is there anyone serious who thinks that the future is local models anyway? All computers used to be the size of rooms like these data centers and then they got smaller and faster until the home computer came. Is that not a possibility down the line as we improve efficiency of the models and increase compute?
I have no strong opinion on whether the endgame of AI services is local or in-cloud, but I think your historical analogy is pretty suspect: it's true that computers got smaller and faster, but it's also true that most people have shifted most of their workloads from local and on-prem to datacenters since the turn of the century. Why would AI be an exception?
AI budgets and AI pricing have too much squish in them currently. Frontier models are being sold at a loss, and AI budgets are experimental. And there is still a whiff of FOMO in the air.
Also, Google and Facebook are still spending like drunken sailors. Nobody has stubbed their toe on hard limitations yet. So yes of course people will figure out how to optimize the cost of AI in their products. Just probably not this year.
That's the usual response, along with "you can't compete with free". But, look at how much money the frontier models are printing, it is obvious that the bell curve of usefulness is still centered around them.
Let's also not forget that even the open models are not getting smaller, they are getting larger. Of course, you can distill them down into something that will fit on smaller compute, but at the end of the day, the data centers of compute, still play a huge role.
An easy counterargument is that the frontier models are swallowing all of the attention and money because they currently don't have a constraint of demonstrating that their value exceeds their cost. Take that away and it might turn out that cheaper models that are "good enough" and can be profitable are the more attractive route.
they currently don't have a constraint of demonstrating that their value exceeds their cost
There is no loyalty in AI. I can switch to another model at near zero cost. They absolutely do have to demonstrate value. The concept of "good enough" is a misnomer because we're talking about putting these products into the hands of people who need to generate value from them.
Open models keep getting bigger, but smaller open models also keep getting smarter. What I can do with an 8b used to require a 32b.
Also the decision makers who are signing off on things like ChatGPT Enterprise are at least 18 months behind the curve of what you can actually do with these things and how cheap they can be. They're still trying to figure out how to actually adopt the tech out of a sense of fomo, nevermind making nuanced decisions about hosting an open weights model. I see this firsthand in my own job.
I'm talking about adding facts to a model by modifying engrams or trying to bolster guardrails with J-washing, meanwhile they're still trying to figure out how to best prompt Copilot.
Give it a few years for everyone else to catch up, I'm barely able to catch my breath before there's some new development in the open source/weights space
Open models keep getting bigger, but smaller open models also keep getting smarter.
Not only bigger, but smarter and more capable. From what I can tell, smaller are only getting smarter in very specific areas. There is a subtle difference there, that is extremely important.
What I can do with an 8b used to require a 32b.
What exactly do you do with an 8b? I usually ask this question and either get no response or it is something that doesn't generate anything of value. So, please surprise me.
Smaller models have improved in every regard, not just domain specific challenges. You can see this in the benchmarks or by just testing one yourself side by side with an older release.
I'm using it at work in multiple data classification and redaction pipelines. It's replaced tedious manual labor and opened that staff up to focus on the parts of the work that requires their human intuition, rather than spending time on tedium.
Presumably part of the reason you don't get a response is your hostile approach to asking.
They have not improved in the one way that I use them today, which is coding. It is like talking to a halfwit, and I always end back up with codex.
Asking a direct question is not hostile. I'm glad to hear you've found a use for a smaller model that generates value. It gives me hope for the future, that said, I think we are still a long ways away from needing HPC in DC's.
I have a shoebox sized computer (Framework Desktop) running Qwen 3.8 Flash Next. It has completely replaced my use of proprietary models in my personal life. 6 months ago I would have told you this was impossible. Based on the current trajectory I expect 6 months from now I'll have a Mythos class model at home. The best part is not having to concern myself with token cost has unlocked all kinds of experimentation and use cases. I have been pushing over a billion tokens per week for multiple weeks now, all for the $52/year it costs to keep this machine running 24/7
What does Omarchy or Spotify have to do with anything I just said?
I am a software engineer, and using this software in my personal and professional work has lead me to a very different conclusion, but everyone is entitled to their opinion.
There will be people who want to host things on-device. At some point, you could probably do most day-to-day tasks with a Siri-like agent, so you don't necessarily need it to be on a datacenter rack somewhere.
More complex tasks being run quickly opens up a choice: insanely beefy individual devices, on-prem hosting, or cloud hosting, whether that be some data center running FOSS models, or ones from people like Anthropic or OpenAI.
Beefy hardware for individual users? Not cost-effective. Could have people share that hardware by putting it in a data center. Do you want to operate that data center? For proven business cases, sure, why not? If you're still working out what your scale will be, maybe you ask the Googles, Amazons, or Microsofts of the world to rent you the hardware so you don't have wasted or too little capacity.
The real question is, how much value is there in a few companies that talk about how their eventual goal is to create AGI as opposed to just giving you enough intelligence to augment your current workers?
The answer is "probably not enough to justify more than one company having a valuation of over a trillion dollars, and that's generous".
They scale with compute so even if today's frontier model equivalents work on future desktop hardware then the big servers will still have bigger and better models.
The value of those models doesn't necessarily though. To extend the personal computer analogy, a server rack has always been more powerful than the average desktop computer but a personal computer got "good enough" at enough tasks that people just used those instead.
Yes but at some point there is diminishing returns in model quality. Most people don't need a model that can solve Navier-Stokes to write emails for them.
One of these companies wanted an IPO valuing them at over $50B. To put in perspective what kind of crazy number that is, the entirety of Visa raised a valuation of $34B; and this company's parent, Softbank, had an IPO valuation of $64B.
I have generally been bullish on data centers independent of how AI demand/load evolves. People will find a way to use that compute, even if it isn't the precise use we expect. As scale increases and the price per computation comes down, we will be able to brute force solutions to problems that otherwise would be intractable or prohibitively expensive to solve. And there are effectively an infinite set of those problems.
This tracks the evolution of how we use cloud computing and GPUs over the last 20 years. Cloud computing was originally just about saving companies from needing to maintain their own server racks, but has unlocked previously unserviced demand by allowing people to build an app or service and scale it to meet rapidly rising use without needing to invest a ton upfront in hardware. Suddenly a hobbyist could spin something up in their spare time that previously required many thousands of dollars of investment. Or when someone wants to run a single large scale computation, they can do it without needing to waste capital on maintaining idling servers to meet an occasional demand spike, like when a company I used to work at moved from running atmospheric calculations on a server in a closet to the cloud and were able to achieve double digit accuracy increases with the increased scale, while spending less overall on batch computing jobs.
GPUs, as the name implies, were created for graphics, and primarily for gaming graphics, but then more or less accidentally ended up enabling the present AI boom, which depends on a scale of computation that would have been impossible with older CPU architectures. Maybe someone at some point predicted this, but I think for the vast majority of people, it was extremely surprising that a niche gaming product would enable an industrial revolution level technological leap forward.
LLMs are just one way that increased compute scale unlocks seemingly magical results, but they are far from the only example and I have no doubt that there are many unknown examples remaining to be discovered yet.
Fair, but my point is that a data center is not pets.com. They are a much more flexible asset, and are even more flexible than AI itself. Pets.com failed because the physical delivery infrastructure and consumer buying habits that support Chewy today did not yet exist and took years to create. Data centers can be repurposed from supporting the current generation of LLM-based AI models to an infinite variety of use cases. Pets.com could only ever hope to deliver pet food.
And Amazon.com could only ever hope to deliver physical books?
Anyway, it's clear an AI data center has utility for crunching AI inference, and that there is and will continue to be demand for AI inference. The trillion dollar question is whether you can make money doing this, and so far the answer is no.
You are not disproving but rather reinforcing my point. Amazon survived the dotcom bubble precisely because they could hope to do more than deliver physical books. It was a more flexible, generally applicable business model than pets.com, and they started out targeting a market more susceptible to ecommerce conversion than pet food delivery proved to be. Similarly, a company selling a product that can only be used for the limited use cases we have found for chatbot-style text generating LLMs will be more likely to fail than a company selling something that is more generally useful. Data centers are more generally useful than LLMs, and the companies building them are more likely to survive than those whose fortunes are pinned entirely to unrealistically high expectations for replacing white collar workers with LLMs.
Bullshit. Even if you don't think LLMs alone are enabling that level of advancement, GPUs in general very much are. They are the critical enabling technology underlying self-driving cars, autonomous drones, camera-based robotics controllers, faster/better drug discovery, faster/better modeling protein structure and design, ML weather forecasting, materials discovery, and many other applications.
I think you're conflating the impact of a small bubble on a narrow market space with our current situation. The current bubble is larger than anything ever seen before, and is now encompassing nearly the entire economy.
New datacenter builds are so far along the curve of diminishing returns it's absurd. No one is going to want to pay to run a datacenter that costs 10x as much to run for the same compute.
And the problem is with all these "freed" resources, the entire pipeline will be affected. No one will want to buy any new silicon if they can buy a B200 for $1000. We could potentially see a decade or more of stagnation in the chip sector, or even significant regressions in capabilities as foundries are shut down due to lack of demand.
The impact of what is coming scares me to my core. I don't think we're going to bounce back from this any time soon.
I think actually it's quite the opposite and you're the one conflating the impacts on a narrow market space with the larger economy. Right now, AI dominates growth in the stock market and demand for chips, but it is by no means encompassing the whole economy. Neither the stock market nor the chip sector are the entire economy. My grocery store isn't going to go out of business if Anthropic does.
A lot of investors may lose money as a result of the bubble bursting, but that does not mean the underlying asset will be forever worthless, just that it didn't provide sufficient returns sufficiently quickly to justify the upfront investment for the investors that funded it, at the time they made that investment. A different investor who could afford to ride out a period of reduced demand might have an entirely different experience.
Imagine you take out a five year loan to buy a truck to deliver packages, but then for the first two years you operate it, gas prices are elevated so you struggle to make the payments on your loan and end up not making as much money as you had hoped to originally, perhaps even to the point you need to declare bankruptcy and sell of the vehicle. But if not and then gas prices drop back down and demand shoots up for the last three years of the loan, you could then make up the difference. If the first two years drive you into bankruptcy, that is difficult for you, but amortized over the entire five year loan period, the truck may actually have been a profitable investment for someone who could have afforded to ride out the first two years. And just because you go bankrupt doesn't mean the truck stops being a valuable asset, it's just that you don't end up benefiting personally from that value because your timing was bad. From a macro perspective, the overall economy doesn't suffer, except to the extent that it might have been more efficient to invest the capital that went into procuring the truck elsewhere during those first two years. But only possibly and only on the margins, because the truck remains a profitable investment over the course of its entire lifetime.
Don't confuse the success or failure of individual investors or businesses with the success or failure of the overall economy. Current data center build outs premised on fanciful projections of demand for LLMs may end up not being profitable in the short term while still being profitable over the entire productive lifetime of the asset, if sufficient demand is found elsewhere or if demand for LLMs picks up later. Similarly, I anticipate at worst we will see chip prices plateau for a while if there is a pullback in LLM demand, but we won't see them fall and they will continue to rise over the longer term as more demand is generated elsewhere.
As one small example: we have barely begun to scratch the surface of what we can achieve with robotics. Think about the demand for video processing if you have tens of thousands of robots stocking shelves in supermarkets generating video all day long. On board processing will of course be the obvious primary demand for chips, which doesn't benefit data centers, but central processing of video to extract useful data from the entire fleet will generate demand for data centers. As will large scale training jobs. Now multiply that thinking across the entire scope of industries where robotics may be useful for replacing human labor, and you're talking about an extremely significant amount of valuable computational work.
Yeah but general purpose data centers that allow hobbyists to spin up an app or a service easily is exactly that: general purpose. AI data centers are built for the very specific kind of math that LLM's require. Even the GPU's can't even be used for something like cloud gaming, which isn't very popular anyways.
That math is very much general purpose. It's just linear algebra under the hood, and linear algebra is a wildly useful toolkit with an unbounded set of potential applications, including many applications other than LLMs. That's how something developed for gaming evolved into something used for text processing and generation. For example, we have barely begun to scratch the surface of what we can achieve with robotics, and I have no doubt that further advances in robotics will generate an enormous amount of computer vision work that GPUs are extremely well designed to handle.
There are two different clocks here. Financials could change in months, but the electrical infrastructure needs to be planned with years.
Gas turbines are all practically sold out till 2030, delivery time changed from 2-3 years to 5-7yrs.
Many announced data centers where destined to delay, independently from financial markets, due to the available infrastructure.
If the AI/data centers fomo dissapear, the effect on electrical production will take years in take effect, because capacity is already reserved and paid.
In oposition to software, energy infrastructure is not flexible: you can't speed up a turbine projected for 2032 neither cancel the order without a considerable cost.
The AI Bubble should more accurately be called the Data Center Bubble. Private credit sold everyone on the idea that AI needs massive data centers right away, because they needed a new narrative to lure investors into funds that are sitting on rapidly deteriorating assets from years of bad loans. As it becomes more and more obvious that these massive data centers are neither wanted nor needed, the Data Center Bubble will burst. AI is going to be just fine, but private credit and the hyperscalers, pension funds, and insurance trusts that they suckered in will be left holding the bag.
https://archive.today/1SJkP
I'll believe it the day SpaceX - the AI company that is mostly making money selling datacenter compute using gas turbines for energy - takes a single day 90% or greater drop in stock price.
Currently nobody knows when the first big financial crisis is fully locked in. For example, if OpenAI can't close another round, and they default on their contracts with Oracle, there's your sign. Until then it looks like everybody is enjoying the communal hallucination.
What do you mean the "first" big? 1929? 2001? 2008?
Do you mean 1929 wasn't a big financial crisis and that, this time, we'll have the first "real" big financial crisis?
I'm confused.
I mean the first AI financial crisis at an AI company. I did provide an example. First, the investor cash has to dry up, then the AI data centers have to start getting itchy about what those multi billion dollar contracts are actually worth.
They mean the roughly 25 year tech sector run that's now culminating in the irrational exuberance of AI. There have been some hits already and the result of those is market cap consolidation of the largest companies. Not sure what the number is but the the 10 largest companies make up a huge percent of the entire S&P and most have extreme exposure to the same risks. The sector has been boosted by the NVIDIA circular financing as well but at some point there is a limit. Although they are angling for a pre-bailout with all the AI is going to kill humanity fear mongering. The only savior for the sector will be the government, the question is if the government steps in before or after an organic collapse.
OpenAI's already announced they're not going public in 2026:
https://www.reuters.com/legal/litigation/openai-ipo-will-not...
They could potentially do another private bridge round, but for a company that was gearing up for the largest IPO in history a couple months ago, the reversal is a pretty bad sign. For investors that are looking for a fire sale, there's already smoke in the air.
Probably the next nvidia chip drop because the old chips where priced at their highest level at a depreciation of 5 years instead of the 2 year norm. What happens when the next chip is way better(hearing 67x better, not just the 10x from initial claims)? Well all those chips need to be dumped fast and depreciate that second and the datacenter gets bankrupted and parted out
Be careful what you wish for. The repercussions might be titanical.
We're already seeing repercussions from an economy that has been retooled not to actually produce anything of value, but to produce more air to fill up the largest economic bubble in the history of the world.
National debt through the roof, inflation through the roof, PHD and research programs gutted, non-ai startups dead and unfunded for the last 4 years. These are just a few things that have been sacrificed on the altar of this bubble - there's far more I haven't recounted.
We're already in a widespread long term economic collapse, but the delusion just hasn't broken yet.
You are in the warmup phase for one of the greatest take offs in human history. Enjoy the ride.
Oh, so our jobs and financial well-being are the rocket fuel being burnt, then?
Any data center that can remain profitable selling open source tokens at commodity prices will be fine. Any data center that relies on OpenAI/Anthropic level token prices and margins might be in trouble. After clearing their debts through bankruptcy, they'd likely be quite profitable selling open source tokens at commodity prices.
I think there's still some low-hanging fruit with thin clients and colocated-to-AI applications. I don't think the math for beefy personal computers is going to hold up, for most use cases.
If it costs you more to generate the tokens that the market is willing to pay for those tokens, then not even bankruptcy will save any of the costs invested in one of these datacenters.
If the cutting edge OpenAI token prices are $80 per 1M token, and the open source tokens are $1 per 1M token, that's a huge gap of "this will never be able to make money under any scenario if the bubble bursts" that will catch a lot of these new datacenters. No one will run a datacenter that costs $5 per 1M token to sell at $1 per 1M token even if the debts are cleared.
The point being made here is that most of those costs are amortized capital costs, which get wiped in bankruptcy.
That $5 per 1M token doesn't literally cost $5 per 1M token. It's more like they had to build a datacenter for $500M that can service 100T tokens over its lifetime. They did this by borrowing money on the capital markets, and now they have to pay interest to those bondholders, interest that they can recoup with their $80/1MT prices. But if it turns out they can't charge $80 and have to charge $1, they won't be able to make those interest payments. They enter bankruptcy, the court wipes the debt clean, and now they don't have to pay interest, only the actual operating costs, which may be more like 50c/1MT. The company gets recapitalized with the new owners being largely the bondholders, the existing equity holders get wiped out, and they can compete with the commodity producers now.
Datacenters aren't free to run.
You have land taxes and or rent, building upkeep, staffing costs, electricity, water, hardware replacement costs.
And new build DCs have blown all these costs through the roof justifying the decision because the price of compute is so high. When the prices come crashing down, the expenses will remain fixed where they are now.
Quite profitable at commodity prices? I don’t buy that. And all of them are building with debt / equity that expects high token prices?
After bankruptcy they likely have no debt, so can out-compete those that didn't go through bankruptcy. It's the bankruptcy that makes them profitable -- it's a common pattern in nascent commodity industries.
Bad news for wall Street though if they have to go bankrupt first?
I mean the physical hardware will be fine but the owners and investors should be worried then right
By the time the bankruptcies come, they'll all be owned by index funds and main street. The price will crash at the break of the bankruptcy, and wall street will swoop in to buy up the "distressed assets" at bargain bin prices.
Don't worry, the already stretched taxpayer will be on the hook for everything just like in 2008!
This time the relief mechanism is already baked into the system (capital does learn from its past mistakes, even if it may not be the lessons you'd hope for!)
https://prospect.org/2026/08/03/ai-bailout-could-be-baked-in...
Guess this is the logic for SPCX bond trading at junk valuations?
assuming demand remains elevated and growing, maybe. But spend on AI is pretty stratospheric right now... companies are already starting to clamp down on spend. This makes you really wonder if there will be sufficient demand at current commodity prices for eg; OSS models to justify all these data centers.
Yes, I believe so. Using a sibling commenter's number of frontier models being 80X the price of commodity models, I think companies who switch will spend a minority of their savings to increase their token usage and only pocket the majority of the savings.
Eventually these models will get commoditized (we're already at the "good enough" stage for real work), then they will get turned into custom hardware and get 1000x faster, then that hardware will get commoditized (like DSPs) and they'll be everywhere and cost $1.
Is there anyone serious who thinks that the future is local models anyway? All computers used to be the size of rooms like these data centers and then they got smaller and faster until the home computer came. Is that not a possibility down the line as we improve efficiency of the models and increase compute?
I have no strong opinion on whether the endgame of AI services is local or in-cloud, but I think your historical analogy is pretty suspect: it's true that computers got smaller and faster, but it's also true that most people have shifted most of their workloads from local and on-prem to datacenters since the turn of the century. Why would AI be an exception?
Do you mean is there anyone who thinks that the future is NOT local models anyway? Since that seems to be the gist of your other points
AI budgets and AI pricing have too much squish in them currently. Frontier models are being sold at a loss, and AI budgets are experimental. And there is still a whiff of FOMO in the air.
Also, Google and Facebook are still spending like drunken sailors. Nobody has stubbed their toe on hard limitations yet. So yes of course people will figure out how to optimize the cost of AI in their products. Just probably not this year.
Google and Facebook have profitable businesses.
Home computers are still nowhere close to as powerful as a $500k+ 10kW server full of 1.5TB of HBM GPU compute.
You don't need a data center to get real work done. Not every task requires "PhD level intelligence"
That's the usual response, along with "you can't compete with free". But, look at how much money the frontier models are printing, it is obvious that the bell curve of usefulness is still centered around them.
Let's also not forget that even the open models are not getting smaller, they are getting larger. Of course, you can distill them down into something that will fit on smaller compute, but at the end of the day, the data centers of compute, still play a huge role.
An easy counterargument is that the frontier models are swallowing all of the attention and money because they currently don't have a constraint of demonstrating that their value exceeds their cost. Take that away and it might turn out that cheaper models that are "good enough" and can be profitable are the more attractive route.
Well said, they're competing on marketing rather than merit. Everyone defaults to the frontier instead of exploring more optimal options
There is no loyalty in AI. I can switch to another model at near zero cost. They absolutely do have to demonstrate value. The concept of "good enough" is a misnomer because we're talking about putting these products into the hands of people who need to generate value from them.
Open models keep getting bigger, but smaller open models also keep getting smarter. What I can do with an 8b used to require a 32b.
Also the decision makers who are signing off on things like ChatGPT Enterprise are at least 18 months behind the curve of what you can actually do with these things and how cheap they can be. They're still trying to figure out how to actually adopt the tech out of a sense of fomo, nevermind making nuanced decisions about hosting an open weights model. I see this firsthand in my own job.
I'm talking about adding facts to a model by modifying engrams or trying to bolster guardrails with J-washing, meanwhile they're still trying to figure out how to best prompt Copilot.
Give it a few years for everyone else to catch up, I'm barely able to catch my breath before there's some new development in the open source/weights space
Not only bigger, but smarter and more capable. From what I can tell, smaller are only getting smarter in very specific areas. There is a subtle difference there, that is extremely important.
What exactly do you do with an 8b? I usually ask this question and either get no response or it is something that doesn't generate anything of value. So, please surprise me.
Smaller models have improved in every regard, not just domain specific challenges. You can see this in the benchmarks or by just testing one yourself side by side with an older release.
I'm using it at work in multiple data classification and redaction pipelines. It's replaced tedious manual labor and opened that staff up to focus on the parts of the work that requires their human intuition, rather than spending time on tedium.
Presumably part of the reason you don't get a response is your hostile approach to asking.
They have not improved in the one way that I use them today, which is coding. It is like talking to a halfwit, and I always end back up with codex.
Asking a direct question is not hostile. I'm glad to hear you've found a use for a smaller model that generates value. It gives me hope for the future, that said, I think we are still a long ways away from needing HPC in DC's.
I have a shoebox sized computer (Framework Desktop) running Qwen 3.8 Flash Next. It has completely replaced my use of proprietary models in my personal life. 6 months ago I would have told you this was impossible. Based on the current trajectory I expect 6 months from now I'll have a Mythos class model at home. The best part is not having to concern myself with token cost has unlocked all kinds of experimentation and use cases. I have been pushing over a billion tokens per week for multiple weeks now, all for the $52/year it costs to keep this machine running 24/7
Great news for Omarchy and the merchants at Spotify! Democratize now! Replace all software engineers!
Seriously, do you believe the garbage you wrote?
What does Omarchy or Spotify have to do with anything I just said?
I am a software engineer, and using this software in my personal and professional work has lead me to a very different conclusion, but everyone is entitled to their opinion.
What's the output rate and config?
Asked my agent to dump this to pastebin from my phone since I won't be home for a while
https://pastebin.com/bBXbpDyu
Here's the chat template I use:
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
Like everything else, "it depends".
There will be people who want to host things on-device. At some point, you could probably do most day-to-day tasks with a Siri-like agent, so you don't necessarily need it to be on a datacenter rack somewhere.
More complex tasks being run quickly opens up a choice: insanely beefy individual devices, on-prem hosting, or cloud hosting, whether that be some data center running FOSS models, or ones from people like Anthropic or OpenAI.
Beefy hardware for individual users? Not cost-effective. Could have people share that hardware by putting it in a data center. Do you want to operate that data center? For proven business cases, sure, why not? If you're still working out what your scale will be, maybe you ask the Googles, Amazons, or Microsofts of the world to rent you the hardware so you don't have wasted or too little capacity.
The real question is, how much value is there in a few companies that talk about how their eventual goal is to create AGI as opposed to just giving you enough intelligence to augment your current workers?
The answer is "probably not enough to justify more than one company having a valuation of over a trillion dollars, and that's generous".
They scale with compute so even if today's frontier model equivalents work on future desktop hardware then the big servers will still have bigger and better models.
The value of those models doesn't necessarily though. To extend the personal computer analogy, a server rack has always been more powerful than the average desktop computer but a personal computer got "good enough" at enough tasks that people just used those instead.
Yes but at some point there is diminishing returns in model quality. Most people don't need a model that can solve Navier-Stokes to write emails for them.
I cannot get into the article or the archive link, from the title seems even Wall Street is seeing AI as a funding black hole.
It's just the scam is reaching its end.
One of these companies wanted an IPO valuing them at over $50B. To put in perspective what kind of crazy number that is, the entirety of Visa raised a valuation of $34B; and this company's parent, Softbank, had an IPO valuation of $64B.
This just doesn't pass the smell test.
I have generally been bullish on data centers independent of how AI demand/load evolves. People will find a way to use that compute, even if it isn't the precise use we expect. As scale increases and the price per computation comes down, we will be able to brute force solutions to problems that otherwise would be intractable or prohibitively expensive to solve. And there are effectively an infinite set of those problems.
This tracks the evolution of how we use cloud computing and GPUs over the last 20 years. Cloud computing was originally just about saving companies from needing to maintain their own server racks, but has unlocked previously unserviced demand by allowing people to build an app or service and scale it to meet rapidly rising use without needing to invest a ton upfront in hardware. Suddenly a hobbyist could spin something up in their spare time that previously required many thousands of dollars of investment. Or when someone wants to run a single large scale computation, they can do it without needing to waste capital on maintaining idling servers to meet an occasional demand spike, like when a company I used to work at moved from running atmospheric calculations on a server in a closet to the cloud and were able to achieve double digit accuracy increases with the increased scale, while spending less overall on batch computing jobs.
GPUs, as the name implies, were created for graphics, and primarily for gaming graphics, but then more or less accidentally ended up enabling the present AI boom, which depends on a scale of computation that would have been impossible with older CPU architectures. Maybe someone at some point predicted this, but I think for the vast majority of people, it was extremely surprising that a niche gaming product would enable an industrial revolution level technological leap forward.
LLMs are just one way that increased compute scale unlocks seemingly magical results, but they are far from the only example and I have no doubt that there are many unknown examples remaining to be discovered yet.
The eventual utility of the Internet did not stop a lot of people losing a lot of money in the dotcom crash.
Fair, but my point is that a data center is not pets.com. They are a much more flexible asset, and are even more flexible than AI itself. Pets.com failed because the physical delivery infrastructure and consumer buying habits that support Chewy today did not yet exist and took years to create. Data centers can be repurposed from supporting the current generation of LLM-based AI models to an infinite variety of use cases. Pets.com could only ever hope to deliver pet food.
And Amazon.com could only ever hope to deliver physical books?
Anyway, it's clear an AI data center has utility for crunching AI inference, and that there is and will continue to be demand for AI inference. The trillion dollar question is whether you can make money doing this, and so far the answer is no.
https://isaiprofitable.com/
You are not disproving but rather reinforcing my point. Amazon survived the dotcom bubble precisely because they could hope to do more than deliver physical books. It was a more flexible, generally applicable business model than pets.com, and they started out targeting a market more susceptible to ecommerce conversion than pet food delivery proved to be. Similarly, a company selling a product that can only be used for the limited use cases we have found for chatbot-style text generating LLMs will be more likely to fail than a company selling something that is more generally useful. Data centers are more generally useful than LLMs, and the companies building them are more likely to survive than those whose fortunes are pinned entirely to unrealistically high expectations for replacing white collar workers with LLMs.
"a niche gaming product would enable an industrial revolution level technological leap forward."
something has to true to be surprising. your statement is false and nonsese
Bullshit. Even if you don't think LLMs alone are enabling that level of advancement, GPUs in general very much are. They are the critical enabling technology underlying self-driving cars, autonomous drones, camera-based robotics controllers, faster/better drug discovery, faster/better modeling protein structure and design, ML weather forecasting, materials discovery, and many other applications.
I think you're conflating the impact of a small bubble on a narrow market space with our current situation. The current bubble is larger than anything ever seen before, and is now encompassing nearly the entire economy.
New datacenter builds are so far along the curve of diminishing returns it's absurd. No one is going to want to pay to run a datacenter that costs 10x as much to run for the same compute.
And the problem is with all these "freed" resources, the entire pipeline will be affected. No one will want to buy any new silicon if they can buy a B200 for $1000. We could potentially see a decade or more of stagnation in the chip sector, or even significant regressions in capabilities as foundries are shut down due to lack of demand.
The impact of what is coming scares me to my core. I don't think we're going to bounce back from this any time soon.
I think actually it's quite the opposite and you're the one conflating the impacts on a narrow market space with the larger economy. Right now, AI dominates growth in the stock market and demand for chips, but it is by no means encompassing the whole economy. Neither the stock market nor the chip sector are the entire economy. My grocery store isn't going to go out of business if Anthropic does.
A lot of investors may lose money as a result of the bubble bursting, but that does not mean the underlying asset will be forever worthless, just that it didn't provide sufficient returns sufficiently quickly to justify the upfront investment for the investors that funded it, at the time they made that investment. A different investor who could afford to ride out a period of reduced demand might have an entirely different experience.
Imagine you take out a five year loan to buy a truck to deliver packages, but then for the first two years you operate it, gas prices are elevated so you struggle to make the payments on your loan and end up not making as much money as you had hoped to originally, perhaps even to the point you need to declare bankruptcy and sell of the vehicle. But if not and then gas prices drop back down and demand shoots up for the last three years of the loan, you could then make up the difference. If the first two years drive you into bankruptcy, that is difficult for you, but amortized over the entire five year loan period, the truck may actually have been a profitable investment for someone who could have afforded to ride out the first two years. And just because you go bankrupt doesn't mean the truck stops being a valuable asset, it's just that you don't end up benefiting personally from that value because your timing was bad. From a macro perspective, the overall economy doesn't suffer, except to the extent that it might have been more efficient to invest the capital that went into procuring the truck elsewhere during those first two years. But only possibly and only on the margins, because the truck remains a profitable investment over the course of its entire lifetime.
Don't confuse the success or failure of individual investors or businesses with the success or failure of the overall economy. Current data center build outs premised on fanciful projections of demand for LLMs may end up not being profitable in the short term while still being profitable over the entire productive lifetime of the asset, if sufficient demand is found elsewhere or if demand for LLMs picks up later. Similarly, I anticipate at worst we will see chip prices plateau for a while if there is a pullback in LLM demand, but we won't see them fall and they will continue to rise over the longer term as more demand is generated elsewhere.
As one small example: we have barely begun to scratch the surface of what we can achieve with robotics. Think about the demand for video processing if you have tens of thousands of robots stocking shelves in supermarkets generating video all day long. On board processing will of course be the obvious primary demand for chips, which doesn't benefit data centers, but central processing of video to extract useful data from the entire fleet will generate demand for data centers. As will large scale training jobs. Now multiply that thinking across the entire scope of industries where robotics may be useful for replacing human labor, and you're talking about an extremely significant amount of valuable computational work.
Yeah but general purpose data centers that allow hobbyists to spin up an app or a service easily is exactly that: general purpose. AI data centers are built for the very specific kind of math that LLM's require. Even the GPU's can't even be used for something like cloud gaming, which isn't very popular anyways.
That math is very much general purpose. It's just linear algebra under the hood, and linear algebra is a wildly useful toolkit with an unbounded set of potential applications, including many applications other than LLMs. That's how something developed for gaming evolved into something used for text processing and generation. For example, we have barely begun to scratch the surface of what we can achieve with robotics, and I have no doubt that further advances in robotics will generate an enormous amount of computer vision work that GPUs are extremely well designed to handle.
There are two different clocks here. Financials could change in months, but the electrical infrastructure needs to be planned with years.
Gas turbines are all practically sold out till 2030, delivery time changed from 2-3 years to 5-7yrs.
Many announced data centers where destined to delay, independently from financial markets, due to the available infrastructure.
If the AI/data centers fomo dissapear, the effect on electrical production will take years in take effect, because capacity is already reserved and paid.
In oposition to software, energy infrastructure is not flexible: you can't speed up a turbine projected for 2032 neither cancel the order without a considerable cost.
The AI Bubble should more accurately be called the Data Center Bubble. Private credit sold everyone on the idea that AI needs massive data centers right away, because they needed a new narrative to lure investors into funds that are sitting on rapidly deteriorating assets from years of bad loans. As it becomes more and more obvious that these massive data centers are neither wanted nor needed, the Data Center Bubble will burst. AI is going to be just fine, but private credit and the hyperscalers, pension funds, and insurance trusts that they suckered in will be left holding the bag.