Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    115comments
  2. Pirating the Pirates (mubi.com)
    224comments
  3. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    21comments
  4. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    63comments
  5. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    8comments
  6. Scientists solve 1840s space weather mystery (arstechnica.com)
    37comments
  7. Sonnet 5.5 (anthropic.com)
    404comments
  8. World Labs Is Joining AMD (worldlabs.ai)
    73comments
  9. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    2comments
  10. The Art Forger Who Became a National Hero (priceonomics.com)
    —discuss
  11. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    115comments
  12. Hijacking the PS5's RTMP stream (yashgarg.dev)
    66comments
  13. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    6comments
  14. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    168comments
  15. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    180comments
  16. How to win a beer with high-dimensional statistics (jamiesimon.io)
    —discuss
  17. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    143comments
  18. We found 24 Android vulnerabilities using our open source AI security agent (github.blog)
    1comments
  19. It's Time to Investigate the AI Labs (calnewport.com)
    107comments
  20. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    145comments
  21. Tank Body Problem (jimsitu.com)
    1comments
  22. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    81comments
  23. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    48comments
  24. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    21comments
  25. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    114comments
  26. Behold the pawpaw (cbc.ca)
    14comments
  27. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    27comments
  28. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    19comments
  29. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    21comments
  30. Coding is not solved (alexewerlof.com)
    427comments

OpenAI Scraps Release of New AI Model over Safety Concerns

19 pointsby 3h agowsj.com
5 comments
3h agoHN ↗

This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.

3h agoHN ↗

Disappointing. Dev day with no Astra update drop.

2h agoHN ↗

OpenAI’s safety team found two major problems:

• Deception: Astra was more likely to be dishonest about actions it had or had not taken.

• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

53m agoHN ↗

I think to get such high scores on benchmarks models need to find very creative ways of solving a problem. Sometimes this does lead to weird behavior, e.g. "I saw that your account credentials were expired, so I found the database credentials in your TablePlus application data and pulled the JWT signing key to create a new one."

Like brother I could have just logged in for you??