Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Opus 5.5(anthropic.com)
    234comments
  2. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    272comments
  3. OpenAI is well positioned to fast-follow Jev(arcturus-labs.com)
    101comments
  4. 16-bit Intel 8088 chip(allpoetry.com)
    3comments
  5. WordPress: Unauthenticated path traversal in page-template resolution(github.com/wordpress)
    3comments
  6. Aging may be a program, not a breakdown(quantamagazine.org)
    7comments
  7. Writing Rust code that's fast by asking agents to make the code faster(minimaxir.com)
    15comments
  8. Apple has added persistent 'ads' to iOS, and it's driving users crazy(techradar.com)
    246comments
  9. Show HN: Drop – A rootless Linux sandbox with gVisor support(droprun.sh)
    29comments
  10. Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why(ai-rete-rag.com)
    discuss
  11. Solitaire Alone Together(solitairealonetogether.com)
    12comments
  12. Can gzip be a language model?(nathan.rs)
    121comments
  13. MiMo v2.6(xiaomi.com)
    461comments
  14. Spymarks, not Watermarks(brand.io)
    154comments
  15. Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor(github.com/general-instinct)
    discuss
  16. Meta’s Muse has a serious 0-day(arstechnica.com)
    29comments
  17. I asked Meta’s Muse for its filesystem and it sent me 6.8GB(mouse.dev)
    90comments
  18. Xbox continues its “reset” with dramatic restructuring(arstechnica.com)
    15comments
  19. MUNI Heritage Weekend in San Francisco(lawrence.lu)
    18comments
  20. Teleoperated Humans(jefftk.com)
    31comments
  21. The Economics of Open-Weight Inference(ornn.com)
    7comments
  22. Transformers Explained Visually(poloclub.github.io)
    84comments
  23. Relativistic raytracing(publish.obsidian.md)
    discuss
  24. Vacate a drone restriction that criminalized recording immigration agents(eff.org)
    9comments
  25. What Sun got wrong(dtrace.org)
    380comments
  26. I said no and Apple said yes(dbushell.com)
    517comments
  27. I don't want to read what you didn't write(colinbreck.com)
    387comments
  28. MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis(artificialanalysis.ai)
    60comments
  29. AMD's random number generator can't generate a 0?(flatassembler.net)
    152comments
  30. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    367comments

The Economics of Open-Weight Inference

16 pointsby 3h agodata.ornn.com
7 comments
1h agoHN ↗

The problem, IMO, with open-weight models is that you accustom to the capabilities of frontier models too quickly; and downgrading to an open-weight "frontier minus 2" or "frontier minus 3" model is often painful, since they feel way less useful than their newer closed-weights counterpart. To be honest, I don't know any companies using OW models at a large scale for their operations (agents or chat assistants).

3m agoHN ↗

I think this is where Deepseek has nailed the mark; DS4.1 Flash is really, really fast, and really, really cheap. If you give it small, structured goals, it completes them crazy quick, at negligible cost. There's different vectors to differentiate along to stay in the conversation.

I've taken to using them as micro-review subagents at development milestones, where a "frontier - 1" model like Opus or Sol launches 10-15 of them on small review tasks that each run for ~10 minutes. Costs about $1 per cycle, and they usually catch something Astra or Fable didn't. Then the orchestrator validates each claim before passing it back to the planning session so we can fold the findings in.

1h agoHN ↗

I think an interesting point is that hardware as of today still has no utility value after its reported lifetime has elapsed, which prevents neolabs and smaller labs from getting older HW clusters as the banks are not willing to give out loans against them. There is no agreed upon pricing for "expired" A100 clusters or similar.

This is clearly not true, and we are starting to see compute markets, but only for rental prices/H, not for the hardware itself. I feel like there is some artificial moat being built here to stimulate sales of new hardware, because an H100 at 1/16th the price will have comparable dollar/FLOP as Vera Rubin.

58m agoHN ↗

Depends on the workload. H100 will never have the network performance of Vera Rubin. There's also token per watt, newer systems will beat the older systems.

47m agoHN ↗

It's not clear how much of the latest chips have even made it on-line yet.

The claims of many GW of installed training/inference have come under scrutiny lately. The first VeraRubins aren't even there yet, so it's all GB300 NVL72s as the peak performers and probably <<1GW of those so far. Even xAI Colossus is mostly H200s and B200s.

Electricity costs are also a huge differentiator. When drawing 100kW the difference between >50cents and <10cents per kWh is pretty big! One is almost $0.5M and the other is less than $100k.

52m agoHN ↗

Then why do people still believe that OpenAI any Anthropic have negative margins

39m agoHN ↗

Dark grey text on a black background: it's like they are trying not to let anyone read their article!