Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    122comments
  2. Pirating the Pirates (mubi.com)
    228comments
  3. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    22comments
  4. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    65comments
  5. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    15comments
  6. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    159comments
  7. Scientists solve 1840s space weather mystery (arstechnica.com)
    37comments
  8. Sonnet 5.5 (anthropic.com)
    417comments
  9. World Labs Is Joining AMD (worldlabs.ai)
    77comments
  10. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    3comments
  11. Hijacking the PS5's RTMP stream (yashgarg.dev)
    67comments
  12. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    9comments
  13. Tank Body Problem (jimsitu.com)
    5comments
  14. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    184comments
  15. How to win a beer with high-dimensional statistics (jamiesimon.io)
    2comments
  16. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    151comments
  17. Bluegraph – Explore NOAA buoy data, rebuilt in 3D from measured spectra (bluegraph.io)
    —discuss
  18. The Art Forger Who Became a National Hero (priceonomics.com)
    —discuss
  19. It's Time to Investigate the AI Labs (calnewport.com)
    120comments
  20. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    82comments
  21. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    150comments
  22. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    136comments
  23. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    25comments
  24. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    54comments
  25. Behold the pawpaw (cbc.ca)
    17comments
  26. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    27comments
  27. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    19comments
  28. Coding is not solved (alexewerlof.com)
    437comments
  29. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    23comments
  30. When did Google get so weird? (sancho.bearblog.dev)
    1038comments

OpenAI Scraps Release of New AI Model over Safety Concerns

19 pointsby 4h agowsj.com
5 comments
4h agoHN ↗

This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.

3h agoHN ↗

Disappointing. Dev day with no Astra update drop.

3h agoHN ↗

OpenAI’s safety team found two major problems:

• Deception: Astra was more likely to be dishonest about actions it had or had not taken.

• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

1h agoHN ↗

I think to get such high scores on benchmarks models need to find very creative ways of solving a problem. Sometimes this does lead to weird behavior, e.g. "I saw that your account credentials were expired, so I found the database credentials in your TablePlus application data and pulled the JWT signing key to create a new one."

Like brother I could have just logged in for you??