Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    22comments
  2. AI companies in race to demonstrate their model most threatening to humanity (thecivilian.co.nz)
    53comments
  3. Owed a billion dollars in Nvidia stock (colo.to)
    335comments
  4. Footguns with Postgres "at time zone 'UTC'" (bookofrevenue.com)
    31comments
  5. SpaceX's Starship launching to orbit for first time ever today (space.com)
    25comments
  6. 37,500 border drawings: a map of the world as people remember it (habibicode.org)
    12comments
  7. Thinking fast and slow in AI: The role of metacognition (2021) (arxiv.org)
    40comments
  8. Ember-1 (fireworks.ai)
    224comments
  9. When did Google get so weird? (sancho.bearblog.dev)
    765comments
  10. Nissan's third generation e-POWER powertrain (nissan-global.com)
    115comments
  11. Functional Mechanical Sympathy [video] (youtube.com)
    3comments
  12. Malleable software: Restoring user agency in a world of locked-down apps (2025) (inkandswitch.com)
    51comments
  13. Prompting Claude Opus 5.5 (claude.com)
    130comments
  14. Made by Mechanical Means (felixrieseberg.com)
    9comments
  15. Alan Kay's answer to “Did the ENIAC have a BIOS”? (quora.com)
    42comments
  16. Self-Hosting on the Dark Web (alvarezrosa.com)
    90comments
  17. Guitar amp and effects pedal built on the Waveshare ESP32-S3-Touch-AMOLED-2.06 (github.com/dashersw)
    52comments
  18. Lunar Terminator Paradox (secretsauce.net)
    56comments
  19. Three Days in August: What a DDoS Attack Exposed in Our Network (nine.ch)
    6comments
  20. Don't couple your Go code to GitHub (iain.rocks)
    125comments
  21. The state of SIMD in Rust in 2026 (shnatsel.github.io)
    40comments
  22. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    113comments
  23. In an $80 motel room, a discovery to shed light on the origins of life (nytimes.com)
    96comments
  24. Deterministic Concurrency [video] (youtube.com)
    5comments
  25. Reading’s Bayeux Tapestry (diamondgeezer.blogspot.com)
    22comments
  26. Imp is a full port of DSPy to the BEAM (github.com/deepfates)
    7comments
  27. Replacing the old battery on rechargeable bike lights (jvns.ca)
    100comments
  28. What I did at Recurse Center (thill.me)
    37comments
  29. Previously unheard recordings of John Coltrane, captured by Frank Tiberi (jazzwise.com)
    34comments
  30. Writing Efficient C++ Code (2013) (asawicki.info)
    127comments

Thinking fast and slow in AI: The role of metacognition (2021)

123 pointsby 8h agoarxiv.org
40 comments
7h agoHN ↗

submitted Oct 5 2021

(In case people miss that before discussion)

7h agoHN ↗

This is still a great paper, but it's missing the second axis of the quadric -- if the only two options are thinking fast or thinking about thinking, that leaves no room for thinking slow yet deliberately, AKA selfconsciousness. See https://www.gutenberg.org/cache/epub/4280/pg4280-images.html for details

I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...

2h agoHN ↗

Isn't system 2 by itself semantically means thinking by thinking

6h agoHN ↗

It looks like a lot like how data bases query optimizers work, with the exception that in the paper there is also a learning/memory component that conditions the evaluation of the answer provided by the first model.

2h agoHN ↗

Given how LLMs access compressed knowledge from their model weights, the similarities make sense

6h agoHN ↗

It reminds me of "At the time we drew boxes labeled 'perception', 'cognition' with arrows between them." An imprecise quote that I can't place.

I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.

4h agoHN ↗

McDermott's A Critique of Pure Reason pretty much captured all of the misgivings I had about "Good Old-Fashioned Artificial Intelligence", which was slightly unfortunate as I was trying to complete a PhD in that very area at the time (around 1990 or so...)

6h agoHN ↗

2021. Please remember the rule of HN to add the year if it’s not actual.

5h agoHN ↗

This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?

edit: why is this downvoted?

4h agoHN ↗

This has already been solved by GPT 5 Adaptive reasoning. A single model that knows when to reason or not based on a thinking parameter we provide (like xhigh). What’s the relevancy to post it today?

Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.

4h agoHN ↗

Asking earnestly, I don’t know what you mean by this reply. I know what system 1 and 2 is. But this has already been solved using same model.

4h agoHN ↗

I know what system 1 and 2 is. But this has already been solved using same model.

No, it hasn't. Maybe you have a different definition of System 1 and System 2. I last read the book well over a decade ago (2011, maybe? 2012?), but System 1 and System 2 are different systems. IOW, System 2 is not a more computational version of System 1.

The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.

In computery terms, System 1 runs in O(1) time, System 2 runs in O(log n) (or maybe just O(n)) time.

This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer"). We don't have LLMs that do that. We have System 2 - run in O(log n) time and produce an answer.

System 1 is completely bereft of thought.

4h agoHN ↗

The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.

No, system 2 is the emergent capability to reason and increase the space of places to find the answer. Forget the paper's proposal, and look at the problem it is trying to solve. Ability to give quick answers, ability to give thought out answers, and the ability to know when to choose what. Adaptive reasoning does all three.

This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer").

No, I don't think we humans use o(1) to for understanding 1000 tokens or 2 tokens. I simply don't think that's the case. There's a new model called "Jev" and it is literally named System 1 (from the book) and even it is billed per input token.

4h agoHN ↗

It's being downvoted, I think, for a few reasons:

* The person who posted it likely posted it not as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on what you're calling its out-of-dateness.

* You say "this has been solved" without defining what "this" is.

* Your description of the solution -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model and past experience. "Effort" isn't so much the issue as types of effort using different systems, modeled specifically after Kahneman's idea of fast and slow thinking.

* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.

In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.

4h agoHN ↗

it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model.

do you even know how adaptive reasoning works?

4h agoHN ↗

Godspeed to you in your efforts to contribute productively to Hacker News threads.

4h agoHN ↗

For the first time, GPT‑5.1 Instant can use adaptive reasoning to decide when to think before responding to more challenging questions, resulting in more thorough and accurate answers, while still responding quickly. This is reflected in significant improvements on math and coding evaluations like AIME 2025 and Codeforces.

https://openai.com/index/gpt-5-1/

It says literally the thing you wanted from system 2. Its almost exactly that.

This is what you said btw:

"it's that the system itself decides how to reason based on the nature of the problem it faces"

4h agoHN ↗

How relevant is this fast/slow thinking thing with regards to current frontier models?

I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.

4h agoHN ↗

You can ask a model for output directly and stop, or you can recursively ask it to keep refining the output.

That seems to fit the fast vs slow model of human thought reasonably well.

2h agoHN ↗

You can ask a model for output directly and stop

That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here

2h agoHN ↗

That’s a function of the amount of processing power involved not the underlying architecture of decision making.

1h agoHN ↗

In terms of making an LLM faster but not in terms of meta-cognition. System 1 thinking as defined by Kahneman doesn’t have 100000x more compute than System 2, it is actually the opposite. That completely contradicts your claim

2h agoHN ↗

That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language. It can be said that language is a tool for the serialization (writing) and deserialization (reading) of human ideas. It is also an incredible useful and powerful tool by itself. This last sentence has been proved true by LLMs themselves. However, since it is working on the serialized version of ideas, I agree with you in that's not the optimal way to think and something not serialized (maybe world models) can be invented that's better for thinking. All this in no way diminishes the usefulness of language and of automated language generation.

56m agoHN ↗

That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language

Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.

[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)

2h agoHN ↗

Structurally speaking we learn nothing like AI, we don't use vast amounts of information to pick up completely new skills. We also make decisions by using prior knowledge and emotions.The latter part is important, Thinking fast and slow cannot operate in a world of AIs as they stand today unless we are willing to grant them rights — because you have to teach them to make decisions based on all kinds of emotions — which is tricky at best.

1h agoHN ↗

we don't use vast amounts of information to pick up completely new skills.

Except, we do.

1h agoHN ↗

Seems like a terrible idea in the first place to build an entire organization around a single pop-sci book, but that’s just me.

1h agoHN ↗

It's indeed a terrible idea - as in, it's great. You get benefits of cross-marketing: you ride on a popularity of a well-known book, and as you also drive more sales of it, even if you don't have a deal and don't benefit from that directly, you strengthen the loop and solidify your brand.

This choice doesn't really constrain what the organization can do, either. Pop-sci books have plenty of wiggle room in interpretation, and afford a lot of "you're holding it wrong" dismissals of criticism, that with a bit of clever copywriting, the organization can do absolutely anything and still claim it's embodying the framework/theory of the book.

59m agoHN ↗

No clue what’s the consensus on this but my internal mental model is absolutely that LLM AI is pure fast mode, no slow mode. The “reasoning” loops are an attempt to mimic the slow mode but ultimately it doesn’t really work. I’m curious about the recent maths advances though, they seem to possibly challenge this.

4h agoHN ↗

Trying to get LLMs to 'think about their thinking' is my daily struggle. This paper nails why it's so critical.

3h agoHN ↗

If I recall correctly, all that fast and slow business has been debunked as yet more non-replicable pop psychology.

I shouldn't be surprised that it shows up in a screed on AI

1h agoHN ↗

There was this post a few days ago https://news.ycombinator.com/item?id=49797323

It had this to say in the linked post:

  This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:

  gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200

  MENENIUS:
  'Though all at once canq

  MARCIUS:
  Pray now, nocamest thou to a morsel.

  LARTIUS:
  Hence, and
  I' the end admire, where G
  again; and after it ag .

Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.

The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.

47m agoHN ↗

Interestingly that is not what we got, but maybe we should loop at architectures like this again? The JEPA loop is interesting, but might fail for the in-flexibility of the component ordering