Top stories

Live mirror
30 storiesupdated 0s agoView source snapshot
  1. Introducing System One Models and Jev(typesafe.ai ↗)
    193comments
  2. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    168comments
  3. An Update on Wayback Machine Access(blog.archive.org ↗)
    173comments
  4. German Rheinmetall open-sources its Battlesuite connected weapon system protcol(rheinmetall.github.io ↗)
    22comments
  5. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
    171comments
  6. Jean-Pierre Serre is 100 years old today(st-andrews.ac.uk ↗)
    11comments
  7. We got admin access to Baseten's production GitHub in 25 minutes(strix.ai ↗)
    94comments
  8. Building a Linux GPU Driver for the M4 Mac Mini in One Month(codyho.dev ↗)
    55comments
  9. Why I'm still bearish on LLMs after Navier-Stokes(dank.systems ↗)
    18comments
  10. Chopping up books when they're physically too big(mattkirkland.com ↗)
    94comments
  11. Learning to solve hard problems in RL for LLMs by never giving up(mnoukhov.github.io ↗)
    discuss
  12. WangNet – 1.8 MB, zero-dependency Numberwang adjudication in 11 languages(github.com/graafhenk ↗)
    32comments
  13. I can't stop thinking about Papua New Guinea(notnottalmud.substack.com ↗)
    415comments
  14. Show HN: Capsule – Single-file web apps that save their data into SQLite(withcapsule.app ↗)
    113comments
  15. Data races and the limits of ThreadSanitizer in C and Go(theconsensus.dev ↗)
    1comments
  16. GEFS on OpenBSD: A Early Preview(marc.info ↗)
    53comments
  17. Let's make quality the norm again(forbrukerradet.no ↗)
    287comments
  18. Suspected sabotage causes major Netherlands rail disruption(bbc.com ↗)
    379comments
  19. Jiga (YC W21) Is Hiring Product Engineer (Remote/US)(jiga.io ↗)
    discuss
  20. Show HN: Pizza Bot – An inbox for AI agents that work in the background(github.com/pizza-bot-app ↗)
    3comments
  21. Vibe Coding is the new Internet Dating?(joecmarshall.com ↗)
    41comments
  22. A single firm is behind OpenAI, Anthropic, and Meta hacking scandals(effort.news ↗)
    140comments
  23. The CSS Zen Garden dream, finally shipped(josprague.com ↗)
    57comments
  24. Show HN: Hacking a $20 4G wireless hotspot into a texting device(bkovac.github.io ↗)
    30comments
  25. Cartesian – AI 3D Modeling for Design(formas.ai ↗)
    69comments
  26. XLS: Accelerated HW Synthesis(google.github.io ↗)
    3comments
  27. US confirms for first time it has deployed space weapons(bbc.com ↗)
    294comments
  28. Most people prefer traditional architecture(worksinprogress.news ↗)
    188comments
  29. The Inference Hardware Revolution of 2026(ieee.org ↗)
    9comments
  30. Giving up on smart rings(notesbylex.com ↗)
    137comments

Why I'm still bearish on LLMs after Navier-Stokes

54 pointsby 5h agodank.systems
17 comments
41m agoHN ↗

I really appreciate seeing a tempered take that's not literally denialist about current capabilities.

32m agoHN ↗

Thanks :) I do enjoy and use these things every day and the current capabilities are indeed amazing, just ludicrously overpriced at the frontier.

10m agoHN ↗

I can see current limitations, but how do you expect capabilities to change in the next few years? A repeat of the gain that happened in the last two years feels like it would be significant, even if it took a little more than two years this time around.

22m agoHN ↗

current frontier models need laborious oversight and guardrails on even the simplest tasks

It is literally denialist about current capabilities

19m agoHN ↗

why don't anthropic and openai ship yolo mode by default?

6m agoHN ↗

They do…? Well, “auto” mode has been default in Claude Code for a couple months now. It’s effectively “safer yolo:” tool calls are inspected by a separate classification system (another smaller LLM, I believe) to approve or deny. And you can always layer on additional sandboxing mechanisms to limit the blast radius deterministically.

3m agoHN ↗

Anthropic basically does at this point with Auto Mode being default. Or was that the point you were making?

17m agoHN ↗

I don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are claiming.

The lack of temperament is very skewed towards the bulls who have been saying AGI is here, software engineering is solved, mathematics is solved, it’s going to destroy the white collar job market, and it’s going to kill us all for like 5 years now.

12m agoHN ↗

Gary Marcus is an especially puzzling addition. If I recall correctly, he has made statements along the lines that superintelligence this century is more likely than not. If you’re AGI-pilled that might read as bearish, but that is still extremely rapid progress in the grand scheme of things.

41m agoHN ↗

the best alternative to rigorous specification is human review. human review doesn't scale well to the volumes of output produced by language models. to make matters worse

when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns

40m agoHN ↗

Specifically: bearish on LLMs generally, not bearish on LLMs for pure math.

31m agoHN ↗

yes, huge for pure math and activities that look like it.

15m agoHN ↗

I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.

Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures.

In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better.

The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.

8m agoHN ↗

the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on, and even then with severe caveats. the frontier labs have developed a general recipe to teach models almost any specific task enjoying clearly defined levels of task performance; many tasks are covered in the training data

Is this really any different to how humans learn, it takes a lot of training on one specific task to make a human expert as well?

5m agoHN ↗

Is this really any different to how humans learn

yes.

5m agoHN ↗

This April 2026 paper is a fun and related read.

https://arxiv.org/html/2509.24239v4

Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.

The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.