Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Phyllotaxis: An audio-reactive LED display (jagi.studio)
    9comments
  2. The systems that no one will test (christianperone.com)
    1comments
  3. Pirating the Pirates (mubi.com)
    261comments
  4. California farmers are struggling to sell grapes as demand for wine drops (kqed.org)
    391comments
  5. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    79comments
  6. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    170comments
  7. ESP32S3 cluster running 1.58-bit (BitNet) Language model (github.com/low-zi-hong)
    16comments
  8. Show HN: Pac-Bench – How well can models one-shot a Pac-Man game? (jonclegg.github.io)
    28comments
  9. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    30comments
  10. US Forces Exit Iraq (reuters.com)
    —discuss
  11. Tank Body Problem (jimsitu.com)
    16comments
  12. Sonnet 5.5 (anthropic.com)
    493comments
  13. Hijacking the PS5's RTMP stream (yashgarg.dev)
    77comments
  14. Who Killed Paulina Borsook's Career? (wired.com)
    10comments
  15. Scientists solve 1840s space weather mystery (arstechnica.com)
    46comments
  16. 1996 chat room simulator connected to Win95 and System 7 web desktops (lolchat.rip)
    44comments
  17. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    209comments
  18. World Labs Is Joining AMD (worldlabs.ai)
    109comments
  19. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    186comments
  20. Simulating Airband Am Radios (bitbashing.io)
    3comments
  21. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    260comments
  22. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    232comments
  23. How to win a beer with high-dimensional statistics (jamiesimon.io)
    8comments
  24. It's Time to Investigate the AI Labs (calnewport.com)
    168comments
  25. Show HN: Museum of Numbers (numbermuseum.com)
    6comments
  26. The Art Forger Who Became a National Hero (priceonomics.com)
    8comments
  27. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    17comments
  28. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    55comments
  29. Show HN: HN.watch – Videos of all Hacker News posts (hn.watch)
    89comments
  30. Behold the pawpaw (cbc.ca)
    30comments

OpenAI Scraps Release of New AI Model over Safety Concerns

23 pointsby 9h agowsj.com
5 comments
8h agoHN ↗

This is a series of Ls for OpenAI that must have hit pretty hard. First Opus 5.5 crushes Astra 6, then Sonnet 5.5 comes in right behind that, and now they won't have an answer until at least November. Which of course gives Anthropic even more time to buff up Fable 5.5.

8h agoHN ↗

Disappointing. Dev day with no Astra update drop.

8h agoHN ↗

OpenAI’s safety team found two major problems:

• Deception: Astra was more likely to be dishonest about actions it had or had not taken.

• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.

I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.

For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).

This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.

It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.

Note: I don't have WSJ, but found these from a tweet[0]

[0]https://x.com/wallstengine/status/2104694678444712189

6h agoHN ↗

I think to get such high scores on benchmarks models need to find very creative ways of solving a problem. Sometimes this does lead to weird behavior, e.g. "I saw that your account credentials were expired, so I found the database credentials in your TablePlus application data and pulled the JWT signing key to create a new one."

Like brother I could have just logged in for you??