Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Code reads AGENTS.md only when telemetry is on(szypowi.cz)
    105comments
  2. The GitHub wiki is an anti-pattern(michaelheap.com)
    31comments
  3. Jev in 25 Lines of Python(nobodywho.ai)
    128comments
  4. I Don't Want the Details(michaelheap.com)
    56comments
  5. Z80 REPL(abagames.github.io)
    9comments
  6. GPT-6 Sol and Luna(openai.com)
    788comments
  7. OpenAI is enlisting an influencer army to make it look 'good for the world'(businessinsider.com)
    88comments
  8. Samsung accidentally freezes its smart fridges with a software update(androidauthority.com)
    78comments
  9. Claude Opus 5.5(anthropic.com)
    1005comments
  10. Tokens Too Cheap to Meter(jyn.dev)
    27comments
  11. QuestDB (YC S20) Is Hiring a Sales Engineer(questdb.com)
    discuss
  12. Two Git ignore files nobody told me about(dinculescu.dev)
    18comments
  13. Transit rewards(waymo.com)
    243comments
  14. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    411comments
  15. Show HN: Jevper – the Jev interface on top of any OpenAI-compatible model(github.com/zhulinchng)
    1comments
  16. The Price of Intelligence Is Falling Rapidly(marginalrevolution.com)
    13comments
  17. Show HN: Ive Sent It – online courier for files, with signed proof of delivery(ivesentit.com)
    8comments
  18. Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived(foxscript.org)
    218comments
  19. What California is learning from solar panels built over irrigation canals(kqed.org)
    519comments
  20. ReBarUEFI: Resizable BAR for almost any UEFI system(github.com/xcuri0)
    60comments
  21. How did AMD Ryzen get 50% faster in two years?(lemire.me)
    166comments
  22. Comma's hands-off driving tech under investigation after 2 fatal crashes(techcrunch.com)
    discuss
  23. 'We hacked the FBI:' Hackers say they have data on all FBI employees(404media.co)
    512comments
  24. SAML: A fractal of bad design(trailofbits.com)
    153comments
  25. Show HN: RxFilm Studio–Create and edit your product videos with AI agent(rxlab.app)
    6comments
  26. Data-only attacks are easier than you think (2024)(usenix.org)
    29comments
  27. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    112comments
  28. Pentagon says overreliance on AI contributed to missile strike on Iran school(bloomberg.com)
    391comments
  29. Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)(artificialanalysis.ai)
    99comments
  30. Show HN: Npunlock – Run custom C kernels for Intel NPUs(github.com/hsfzxjy)
    12comments

Tokens Too Cheap to Meter

49 pointsby 5h agojyn.dev
27 comments
1h agoHN ↗

We are likely to see LLMs running locally at current frontier-quality on commodity hardware in the next 3-6 years

Yeah okay bud, anyone checked in with the state of consumer hardware recently? Not the author, evidently.

oh in 3-6 years this will all be over

Yeah I'm sure Samsung, Nvidia and sk hynix will all be very calm with lower volumes and lower margins.

1h agoHN ↗

Yeah okay bud, anyone checked in with the state of consumer hardware recently? Not the author, evidently.

RAM prices will crash when demand drops even a little. They'll probably crash to a lower (inflation adjusted) level than before. This has happened before.

Industrial scaling in general often looks like a sawtooth: price spike, capacity investment, crash, repeat.

Part of what's keeping prices high a little longer is that everyone knows this and is a little reluctant to plow resources into chip fabs for fear of having the bottom fall out before they recoup or sell that to someone else to hold that bag.

Graph the average compute and RAM in a mid-high end laptop at an inflation adjusted price point for the past 40 years. It's very exponential and hasn't slowed down much.

1h agoHN ↗

everyone knows this and is a little reluctant to plow resources into chip fabs

except cxmt who is plowing resources in like crazy

19m agoHN ↗

That has more to do with geopolitics than it does the current price of memory.

55m agoHN ↗

No. Prices will crash when supply side expands to meet the increased demand. Because demand won't go down to pre-bubble times any time soon. Unfortunately the supply side has been very slow in increasing production, partly because most steps of the production chain are all maxed out.

On a long enough scale you are right that prices will likely normalize to a better level, but before 2030? That would mean the factories are built quickly once they begin.

49m agoHN ↗

The article observes that the cost of frontier intelligence from 2025 has fallen 100x in the last year. It also notes that the energy to run models is also collapsing. Consumer hardware is borked right now because these new algorithms are revolutionizing the utility of a computer. Computing is technology who's cost has been collapsing for 90 years, and its a safe prediction that it will decrease again.

1h agoHN ↗

This is the core of my belief that data center construction is a huge bubble.

AI is not a bubble, IMO, though we may see a retrench and some companies with sky-high valuations will crash to more reasonable ones. But data center demand is probably a bubble, and the main driver will be reduction in the actual amount of power and data center space required to serve escalating demand.

I think hardware and model improvements will pace or maybe outrun demand and then when demand starts to saturate will keep going and leave a lot of orphaned data centers.

1h agoHN ↗

Jevon's paradox says that if data centers can serve a lot more tokens per dollar or watt there will be increased demand for data centers.

37m agoHN ↗

Jevon's paradox isn't a physical law, it doesn't magically apply to everything. Millions more copies of Atari's ET game didn't cause everyone to pickup a cheap copy, and cause extra demand for a garbage video game. Some times (actually, usually, I'd argue) things are made that will sell for less than the cost of construction because of irrationality, and they don't induce extra demand and they don't change the negative profit margins.

You can't simply wave Jevon's paradox at things. Thousands of miles of canals were dug in the UK that couldn't be sustained and were abandoned. Thousands of miles of railways were laid that could be sustained and were abandoned. And those are potentially durable investments, unlike cheap walls, pillars and roofs laid over a levelled concrete slab full of fast depreciating IT equipment.

26m agoHN ↗

It's true that Jevon's paradox doesn't always apply, although this does seem like a classic case.

But yes, if sold for a negative margin Jevon eventually stops because the decreasing supply will drive up prices.

things are made that will sell for less than the cost of construction

Price is set at the marginal cost. Capital costs aren't in marginal costs.

You'll need a better counter-example than UK railways which suffered from Parliament price-fixing.

1h agoHN ↗

NVIDIA will still boom

I think Nvidia is under the same pressure as Anthropic/OpenAI. Nvidia will dominate research and probably keep dominating training, but the real volume is in inference. And for inference Nvidia's lead is only a few months, similar to the lead frontier labs have over open source. Nvidia will sell a lot of Rubin CPX's, but their margin on that will be a lot smaller than B200 because there is so much more competition in that space.

1h agoHN ↗

It is hard to see how the environmental side effects of this aren't going to be somewhere between bad and disastrous.

1h agoHN ↗

No it'll be fine as long as you do your part and not drive a car, or have AC, or eat meat, or have children, or live in detached housing, or...

44m agoHN ↗

I think this is a case where just drawing a "line goes up" extrapolation is incredibly misleading because there is _tremendous_ economic pressure to get costs down, and costs are very tightly tied to energy use. All of these systems are incredibly inefficient right now and have a lot of room to go down in energy use. I'd guess that the absolute _floor_ is burning model weights directly to silicon and that's like a 90+% reduction in energy use.

55m agoHN ↗

I found the OP insightful and worth a read. Thank you for sharing it on HN.

The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."

The OP glosses over questions of business model viability with a brief qualitative discussion and very little hard data. For example, to earn an annual return > 10% on every trillion dollars of capital sunk into infrastructure, the owners of that infrastructure must earn free cash flow (operating profit less investment) in excess of $100 billion per year in perpetuity. Is that feasible? Why? How?

The OP does not really consider such questions.

51m agoHN ↗

Exactly. "Cost-to-distill" is a critical parameter. Right now usage of frontier models for all tasks is both subsidized and irrationally popular even at the subsidized price. Deepseek would solve most tasks faster and 10x cheaper. I agree with the author that just as Deloitte exists, frontier labs will exist. But not because their products are proprietary technical marvels or gods, but rather because of branding.

23m agoHN ↗

They are already turning profits and inference has shown to be a cash cow. And they've already secured compute for the next several years.

15m agoHN ↗

Some frontier labs are reporting positive "adjusted EBITDA" (earnings before interest, taxes, depreciation, and amortization, with extra adjustments to make the figure positive).

Free cash flow (operating profit less investment), actual cash coming in, is deeply in the red.

EBITDA can be a sensible measure of profitability when there isn't much need for additional investment. That doesn't seem to be the case with these operators. They need to invest aggressively to avoid losing customers to competitors. All of these operators have made multi-year commitments to invest more in infrastructure. In addition, they have guaranteed quite a bit of debt to fund it.

Maybe it all will work out fine, but I didn't see any hard data from the OP, or from you, supporting that view.

11m agoHN ↗

Who is the "they" that are turning profits?

43m agoHN ↗

Tokens become cheaper than tool calls

The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)

32m agoHN ↗

Right. And some hardware improvements will speed up both grep and Luna, which won't close the gap.

5m agoHN ↗

not necessarily, one may be easier to parallelize while the other suffers some serial computation bottleneck.

24m agoHN ↗

It might never beat out grep, but it could beat some more expensive to call tools, similar to how heuristics will often be faster than exact answers. Rust Analyzer can be slow at times, I could see an AI tool taking over a subset of its work.

6m agoHN ↗

LLM is spicy memoizing, so it can potentially be faster than a tool call. But people will spend a month tweaking and testing to ensure they have the level of determinism they need, which means it's more expensive, and that they should have used actual memoization in the first place.

38m agoHN ↗

There seems to be a mistake in the cost comparison between 2025 and 2026. The 2025 chart axis is the cost to run the entire "intelligence index", and the 2026 version is a weighted average cost per task.

I don't disagree with the thesis here, I just don't think costs are coming down quite that quickly.