Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    145comments
  2. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    36comments
  3. Phyllotaxis: An audio-reactive LED display (jagi.studio)
    1comments
  4. Pirating the Pirates (mubi.com)
    241comments
  5. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    68comments
  6. Tank Body Problem (jimsitu.com)
    9comments
  7. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    27comments
  8. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    267comments
  9. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    8comments
  10. Sonnet 5.5 (anthropic.com)
    447comments
  11. Scientists solve 1840s space weather mystery (arstechnica.com)
    41comments
  12. Who Killed Paulina Borsook's Career? (wired.com)
    —discuss
  13. Hijacking the PS5's RTMP stream (yashgarg.dev)
    69comments
  14. Show HN: Pac-Bench – How well can models one-shot a Pac-Man game? (jonclegg.github.io)
    5comments
  15. World Labs Is Joining AMD (worldlabs.ai)
    95comments
  16. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    195comments
  17. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    186comments
  18. How to win a beer with high-dimensional statistics (jamiesimon.io)
    6comments
  19. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    174comments
  20. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    161comments
  21. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    12comments
  22. Bluegraph – Explore NOAA buoy data, rebuilt in 3D from measured spectra (bluegraph.io)
    2comments
  23. It's Time to Investigate the AI Labs (calnewport.com)
    136comments
  24. The Art Forger Who Became a National Hero (priceonomics.com)
    5comments
  25. Profit Margins of the Largest Companies (visualcapitalist.com)
    3comments
  26. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    86comments
  27. U.S. Strategic Petroleum Reserve Falls to Lowest Level Since 1982 (oilprice.com)
    123comments
  28. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    59comments
  29. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    40comments
  30. Behold the pawpaw (cbc.ca)
    26comments

OpenAI Scraps Release of New AI Model over Safety Concerns

19 pointsby 6h agowsj.com
5 comments
6h agoHN ↗

This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.

5h agoHN ↗

Disappointing. Dev day with no Astra update drop.

5h agoHN ↗

OpenAI’s safety team found two major problems:

• Deception: Astra was more likely to be dishonest about actions it had or had not taken.

• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

3h agoHN ↗

I think to get such high scores on benchmarks models need to find very creative ways of solving a problem. Sometimes this does lead to weird behavior, e.g. "I saw that your account credentials were expired, so I found the database credentials in your TablePlus application data and pulled the JWT signing key to create a new one."

Like brother I could have just logged in for you??