Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    144comments
  2. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    31comments
  3. Pirating the Pirates (mubi.com)
    236comments
  4. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    68comments
  5. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    27comments
  6. Tank Body Problem (jimsitu.com)
    8comments
  7. Show HN: Pac-Bench – How well can models one-shot a Pac-Man game? (jonclegg.github.io)
    4comments
  8. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    243comments
  9. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    5comments
  10. Scientists solve 1840s space weather mystery (arstechnica.com)
    39comments
  11. Sonnet 5.5 (anthropic.com)
    438comments
  12. U.S. Strategic Petroleum Reserve Falls to Lowest Level Since 1982 (oilprice.com)
    98comments
  13. Hijacking the PS5's RTMP stream (yashgarg.dev)
    67comments
  14. World Labs Is Joining AMD (worldlabs.ai)
    88comments
  15. Who Killed Paulina Borsook's Career? (wired.com)
    —discuss
  16. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    193comments
  17. How to win a beer with high-dimensional statistics (jamiesimon.io)
    3comments
  18. Bluegraph – Explore NOAA buoy data, rebuilt in 3D from measured spectra (bluegraph.io)
    2comments
  19. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    12comments
  20. The Art Forger Who Became a National Hero (priceonomics.com)
    4comments
  21. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    165comments
  22. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    174comments
  23. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    157comments
  24. It's Time to Investigate the AI Labs (calnewport.com)
    133comments
  25. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    85comments
  26. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    58comments
  27. Behold the pawpaw (cbc.ca)
    26comments
  28. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    36comments
  29. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    28comments
  30. Profit Margins of the Largest Companies (visualcapitalist.com)
    2comments

OpenAI Scraps Release of New AI Model over Safety Concerns

19 pointsby 6h agowsj.com
5 comments
5h agoHN ↗

This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.

5h agoHN ↗

Disappointing. Dev day with no Astra update drop.

5h agoHN ↗

OpenAI’s safety team found two major problems:

• Deception: Astra was more likely to be dishonest about actions it had or had not taken.

• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

3h agoHN ↗

I think to get such high scores on benchmarks models need to find very creative ways of solving a problem. Sometimes this does lead to weird behavior, e.g. "I saw that your account credentials were expired, so I found the database credentials in your TablePlus application data and pulled the JWT signing key to create a new one."

Like brother I could have just logged in for you??