Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Small Programming Tricks(will-keleher.com ↗)
    63comments
  2. Dream-RSI: Recursive Self-Improvement through Evolving Worlds(arxiv.org ↗)
    31comments
  3. Mistral X Mozilla: Private, Multilingual AI Browsing(mistral.ai ↗)
    128comments
  4. Introducing System One Models and Jev(typesafe.ai ↗)
    461comments
  5. Tell the speakers that you liked their talks(ohhelloana.blog ↗)
    32comments
  6. Claude Cowork and chat are now one Claude(claude.com ↗)
    73comments
  7. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    223comments
  8. How big are factorials?(thegreenplace.net ↗)
    13comments
  9. Apple Reference Image: A New Approach for Verified Photography(security.apple.com ↗)
    293comments
  10. Hackers Got Inside a Flock Camera(wired.com ↗)
    142comments
  11. The Google Play app review process now regularly takes longer than a week(gultsch.social ↗)
    255comments
  12. Can we stop with the uptime percentages?(jim-nielsen.com ↗)
    56comments
  13. This Code Is CRAP (2011)(googleblog.com ↗)
    39comments
  14. The DeepMind Institute(deepmind.com ↗)
    2comments
  15. Scaling Golang CI by Replacing actions/setup-go(cloudx.ai ↗)
    6comments
  16. Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models(stale.jock.pl ↗)
    25comments
  17. Kyber (YC W23) Is Hiring a Forward Deployed Engineer(ycombinator.com ↗)
    discuss
  18. Prisma's pgbouncer=true on Supabase made every query 4 round-trips (postmortem)(simbastack.com ↗)
    3comments
  19. An update on Wayback Machine access(blog.archive.org ↗)
    339comments
  20. Measuring Gauss-Seidel loop-carried dependency and fixing it via loop unrolling(loiseaujc.github.io ↗)
    3comments
  21. Original Sony PlayStation 2 security chip 'broken wide open' after 26 years(tomshardware.com ↗)
    53comments
  22. The Siberian Ice Maiden and the Scythian World(patrickwyman.substack.com ↗)
    discuss
  23. Salesforce Global Outage(salesforce.com ↗)
    135comments
  24. Show HN: I made a flight simulator, except you're just a passenger(inflightsimulator.com ↗)
    189comments
  25. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
    314comments
  26. Anatomy of a Texture(agentlien.github.io ↗)
    4comments
  27. Doing Everyone Else's Job(yosefk.com ↗)
    84comments
  28. Why I'm still bearish on LLMs after Navier-Stokes(dank.systems ↗)
    507comments
  29. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI(arxiv.org ↗)
    45comments
  30. DeepSeek v4.1 Flash Is Now Our Best Hacking Model(enclave.ai ↗)
    42comments

Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models

37 pointsby 4h agostale.jock.pl
25 comments
3h agoHN ↗

Do people prefer the new flat style LLMs are producing? I don’t mind it as much as the gradient theme they were pumping out previously.

3h agoHN ↗

It still matters, but in the age of good reasoning, tool use, and web search, this is much less of a problem than it used to be.

3h agoHN ↗

all the reasoning still comes from pretraining data

2h agoHN ↗

I think they mean reasoning their way to the need for a web search.

2h agoHN ↗

Or other kind of search / knowledge acquisition / computer use etc to get the information needed

2h agoHN ↗

Says who? Models can also use results from tool calls in their reasoning loops.

1h agoHN ↗

in my chat with gemini it could not differentiate between current events and fiction.

if you point it to the web it got the point, but started treating everything like fiction. so it simply started making up possible scenarios and playing them off as real answers when asked for factual information.

i could not tell what the issue was or how to fix it because the reasoning is encrypted. the obfuscation model spat out something like: 'the user is asking for details about a fictional scenario in which the usa has assassinated the leader of iran'

i really don't like the way big ai companies are going. encrypted thinking, guardrails, adversarial personality, moralizing. it is creating something anti-human.

3h agoHN ↗

After Trump's last inauguration, ChatGPT would still tell me that Biden was President of the US. I understand that the training cutoff was before Biden dropped out. But it knew, or should have known, the current date and that there had been an election since its last update, but it didn't qualify the answer. When I asked it to search the web, it got it right. The moral I took away was to always ask for the search whenever I ask about current events. I do that so routinely that I wouldn't know if this problem has been fixed. I suppose that failing to update my priors per individual model release is a form of bigotry against a widely hated class.

3h agoHN ↗

ChatGPT recently started web searching for for basically every general knowledge question, which I found quite odd. Maybe an overcorrection to the issue you were having?

38m agoHN ↗

I forgot which was it, ChatGPT or Gemini, but one of them insisted on calling Trump "former president" even when discussing decisions he just announced as president. Lol

3h agoHN ↗

I remember when the US captured Venezuelan president Maduro, and when I posed a prompt related to this, the model said that’s pure fiction. I told it to double check. Still didn’t want to entertain the idea. It only acquiesced when I specifically directed it to check Reuters. I haven’t noticed this problem in months. Model cutoff seems to be less of a problem these days.

2h agoHN ↗

It's a "problem" of compute, I think. If you query without an account on ChatGPT you will see the model look up less stuff and research less, than when you have a paid account and choose "medium" or "high" in the effort slider.

Which makes sense, because of you have looked into search and crawlers you notice that search is actual quite expensive (which is why e.g. Kagi charges a few bucks for search every month).

2h agoHN ↗

It's not strictly compute, because this has noticeably improved in open-weight models too, such as Gemma and Qwen. I suspect they noticed this issue and adjusted their training to be better about it over time.

41m agoHN ↗

I built a toy news-summarizing agent with Gemma 4, and it was so frustrating, actually, because of the cut-off date.

The model wasted over half the token budget, each time, on internal debates over the current date.

When generating a World Cup summary, for example, it refused to believe qualification rounds were over and refused to even call the web searching tool to collect the data.

I injected the current datetime at the very beginning of the system prompt, but Gemma refused to believe it!

The m-effer insisted the timestamp was fake and hypothesized it was being evaluated in a synthetic lab test with simulated future dates!

No amount of system prompting could convince it to trust the clock.

That was the most frustrating and bizarre "bug" I ever faced!

2h agoHN ↗

Came here to say the same thing. Models used to rely heavily on world knowledge from their training data. They are now much better at tool use and deciding when to research a topic, rather than just answering from memory.

I wonder how much that extends to using LLMs for programming. I assume most knowledge of programming language syntax still comes from training data.

1h agoHN ↗

the model said that’s pure fiction.

Were you expecting your model to be updated on current events? Why?

Also the specific event you are referring to is a statistically very improbable event, prior to its actually happening.

It only acquiesced when I specifically directed it to check Reuters.

Do all models do this? They check in with Reuters? Why would a model think that you asking about an extremely improbable event warranted reaching out to Reuters?

1h agoHN ↗

He asked it to double check. It's reasonable to expect the LLM to handle that trivial task.

41m agoHN ↗

I was not expecting model weights to be updated on current events.

It’s clearly warranted because a model that trusts its weights on current events will give an outdated answer. Extremely improbable events happen all the time.

10m agoHN ↗

I think the models are trying to optimistically avoid doing web searches, because they're surprisingly a lot harder to do well than you'd think.

3h agoHN ↗

Depending on the use case certain models very well remain as or more reliable for certain tasks.

3h agoHN ↗

Pre-AI internet data is like pre-war steel

The slop would multiply if we keep feeding it to new models in a loop

1h agoHN ↗

This is one of the problems that eventually solves itself, somehow

2h agoHN ↗

I remember running the docker container for ollama and its knowledge cutoff is somewhere in 2023 still. That's unacceptable.

2h agoHN ↗

ollama is just an inference engine - it just runs models.

it must ship with some default old model if you didn't need to explicitly download one