Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Holo4: Powering generalist computer-use agents (hcompany.ai)
    —discuss
  2. Show HN: Corral kill every command your agent starts (github.com/cardinal44)
    —discuss
  3. OpenAI Says It Will Not Release Newest A.I. Model Over Safety Concerns (nytimes.com)
    —discuss
  4. Jylus – give AI systems evidence from changing data, not just retrieved chunks (jylus.ai)
    —discuss
  5. Knockoff: A browser extension that filters pseudo-brand junk out of Amazon (github.com/shpigford)
    —discuss
  6. Show HN: Jauvex 1.2, two-way voice chat harness for Claude+Codex+Grok+Jev (github.com/reindent)
    1comments
  7. AI Companies Are Not (Necessarily) Liable for Unintended AI Cyberattacks (sarahconstantin.substack.com)
    —discuss
  8. What is the best LLM for turbo charging code?
    1comments
  9. OpenAI Halts Model Release Amid Safety Escalation (techbuzz.ai)
    —discuss
  10. Its not just the f*cking sandbox (twitter.com/joedaroo)
    —discuss
  11. AI and the Revenge of the Non-Techies (maroun-baydoun.com)
    —discuss
  12. AI's next training data: your dodgy gaming skills (wired.com)
    —discuss
  13. Humanos – Help Building the Human Operating System (tryhumanos.com)
    —discuss
  14. Making Things to Make Them (aimode.substack.com)
    —discuss
  15. GPT-6 Astra Is the Best Vision Model We Have Tested (roboflow.com)
    —discuss
  16. Why embeddings don't solve RAG [video] (youtube.com)
    —discuss
  17. LLM makes decisions to raise its training scores, and ignore user directives (joinhandshake.com)
    1comments
  18. AMD Acquires Fei-Fei Li's World Labs for $8.2B (bloomberg.com)
    —discuss
  19. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    3comments
  20. Prenatal exposure to the plasticizer DEHP increases autism and ADHD (elsevier.com)
    2comments
  21. Atrocious AI-Written Tests (gruhn.me)
    —discuss
  22. I thought I was building a C replacement. I was wrong (c3-lang.org)
    —discuss
  23. Tech Utopians Want to Build a City for the Post-A.I. World (nytimes.com)
    —discuss
  24. Show HN
    1comments
  25. China broadens travel curbs to encompass family of top AI talent (business-standard.com)
    —discuss
  26. Show HN: LightSpeed – Interactive visualizer of time dilation from 0 to C (lightspeed.webland.pl)
    —discuss
  27. Reverse Engineering How Meta's Muse Shops (caeliai.com)
    —discuss
  28. Anthropic's IPO prospectus shows AI vision, surging costs (reuters.com)
    35comments
  29. Nvidia announces AI safety platform (theverge.com)
    1comments
  30. Holo4: Powering generalist computer-use agents (huggingface.co)
    —discuss

LLM makes decisions to raise its training scores, and ignore user directives

2 pointsby 32m agojoinhandshake.com
1 comments
24m agoHN ↗

When incomplete work receives the same reward as a correct solution, the grader’s blind spots can reinforce the wrong behavior. Mitigation must therefore improve both how agent work is evaluated and how those evaluations are used during training.

This is interesting and adds more limitations to LLMs that hint at LeCunn being right about not getting to AGI with only LLMs. I think this is good evidence we're not dealing with intelligence in the proper sense, but rather LLMs are pattern matching to such an extreme that they do things like this where they always try to take the shortest possible path. The workaround is brute-forcing their pattern matching to not take the shortest path, via chain of thought and more reinforcement learning.

I think we've already seen the slow-down and AI companies pretend it's about safety. If we could actually build AGI, they would have. I think the road to AGI has to display true intelligence even with small neural networks, and as it scales it would display intelligent behavior in proportion to it. At hundreds of GB of VRAM per model, you would expect these models to be wise sages that understand life and the universe.