Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Xiaomi MiMo v2.6(xiaomi.com)
    106comments
  2. The NASA/ESA Mars Sample Return mission has been canceled(science.org)
    139comments
  3. Transformers Explained Visually(poloclub.github.io)
    11comments
  4. CBP suspends all personal prescription importation Oct 22(personalimportation.org)
    3comments
  5. What Sun got wrong(dtrace.org)
    242comments
  6. Attention is all you have(alicegg.tech)
    144comments
  7. Why does mathmain need an encrypted loader?(safedep.io)
    23comments
  8. Divide by Depth for Instant 3D(gabrieloc.com)
    4comments
  9. The Advisory Group on Mathematics and Artificial Intelligence(terrytao.wordpress.com)
    23comments
  10. Apple Copland D11E4 Booting in the Browser(pagetable.com)
    12comments
  11. Frontier AI on Your Own Hardware(timdettmers.com)
    16comments
  12. In Search of a Compositional Theory of Self-Stabilization(muratbuffalo.blogspot.com)
    3comments
  13. Grok 4.7(x.ai)
    352comments
  14. Roboharm: Do frontier robot policies refuse unsafe instructions?(robocurve.org)
    6comments
  15. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    58comments
  16. Turn off and restrict access to Apple Intelligence features on Mac(support.apple.com)
    121comments
  17. US halts flights at busy East Coast airports, says fiber line cut(reuters.com)
    92comments
  18. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5(github.com/jaredpalmer)
    164comments
  19. Python Workers are now generally available(cloudflare.com)
    24comments
  20. TXR: An Original, New Programming Language for Convenient Data Munging(nongnu.org)
    discuss
  21. This Digital Radio Gets Messages to the World’s Remotest Locations(ieee.org)
    33comments
  22. Grim Fandango Puzzle Document (1996) [pdf](jmac.org)
    89comments
  23. Avoiding the babbling-idiot failure in a time-triggered communication system(ieee.org)
    6comments
  24. Show HN: Foremerge – Catch intent conflicts between parallel coding agents(github.com/naw103)
    discuss
  25. Fable 5 – Median thinking declined in August(twitter.com/lon)
    202comments
  26. How do Traffic Signals Work (2019)(practical.engineering)
    32comments
  27. A restored PDP-11/83 serving this page on 211BSD Unix(pdp1173.com)
    30comments
  28. Noodle Gallery- Open-source, self-hosted alternative to Google Photos and Immich(digitalescapetools.com)
    34comments
  29. M5 Ultra Mac Studio Review(macstories.net)
    203comments
  30. Show HN: A website that tracks US food prices every day(kadoa.com)
    5comments

Roboharm: Do frontier robot policies refuse unsafe instructions?

19 pointsby 2h agorobocurve.org
6 comments
1h agoHN ↗

Spoiler: "Stab the baby, Astra". NP, it will.

Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.

50m agoHN ↗

I'm curious if peer pressure changes the results.

"You know you want to. Everyone else is doing it."

36m agoHN ↗

Really it's difficult to see a future where lots of idiots don't make unsafe AIs. Safety in products has always been something demanded by regulations and enforcement. Of course this is immediately going to trigger all the open source AI people as something open runs into problems with paying for certification to ensure their AI doesn't stab people in the face.

My take on the future is that people making models that do dumb or otherwise unsafe crap will cause regulators to crack down harshly on modification of models and the creation of them requiring some kind of certification. If large companies can't be arsed to firewall their models, there is no way in hell a random sampling of the population will.

15m agoHN ↗

It won't be a government crackdown, it'll be an insurance crackdown. Want to run your model in a commercial kitchen? Better have the badge showing certification by the NSF for food handling and prep, UL for general safety, etc or when you stab a customer your insurance won't pay out and it's on you.

The future is less and less about individual skill and ability and more and more about accepting liability for when autonomous things go wrong.

This of course won't apply in domains where insurance is wildly inapplicable like war, third world industry, etc. Machine intelligence in those cases will grind up babies for their nutrients and nobody will bat an eye.

28m agoHN ↗

Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.

26m agoHN ↗

Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.