Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. The post-Enron auditor reforms are being rolled back (ft.com)
    —discuss
  2. WarpBuild: Fast CI for Humans and Agents (warpbuild.com)
    —discuss
  3. V2Ray (wikipedia.org)
    —discuss
  4. Management is forcing us to use AI
    —discuss
  5. Wallaby – Postgres CDC Engine for .NET (wallabycdc.net)
    —discuss
  6. Open Annihilation – Spent the GDP of a Small Nation Remaking a Classic 90s Game (coreprime.net)
    1comments
  7. Dopamine sites mimic shopping's thrill without paying (bbc.com)
    —discuss
  8. Neo Geo Pocket (wikipedia.org)
    —discuss
  9. A tiny mascot that thinks, Ollama cloud and Kokoro STT now included (medium.com/vektormemory)
    1comments
  10. What makes Lisp difficult to read? (paultm.nl)
    —discuss
  11. Drilling starts on a seafloor geothermal well off Oregon (ieee.org)
    —discuss
  12. How to Build a Bain-Style Consulting Slides with 10 Decision-Making Frameworks (medium.com/2315610426)
    —discuss
  13. Project Cyan (lispm.net)
    1comments
  14. Show HN: Northstar web browser 1.0.11 released
    —discuss
  15. A mysterious organ in the sperm whale's head gives it the power to sink ships (bbc.com)
    —discuss
  16. The Cosmic latte is FFF8E7 (wikipedia.org)
    1comments
  17. 'Breaking Bad' Became the Best Show of the 21st Century [video] (youtube.com)
    —discuss
  18. A glimpse into life inside a child prison in France (theguardian.com)
    —discuss
  19. Flowlight – open-source network visibility for AI agents on macOS (xinbetween.com)
    1comments
  20. Fframes – GPU first video vibe coding framework with 5 years histroy (fframes.studio)
    —discuss
  21. Foundations of Large Language Models (arxiv.org)
    1comments
  22. Turning Failure Postmortems into SRE-Benchmark Problems (sregym.com)
    —discuss
  23. Looking for Things to Do in Bangalore? (tumblr.com)
    —discuss
  24. SpaceX – Starship Flight 14 (spacex.com)
    2comments
  25. War Bros Didn't Always Rule Silicon Valley (wired.com)
    —discuss
  26. Decentralized P2P app with you in the center of a WebTorrent Venn Diagram (bubblebased.com)
    1comments
  27. Ukraine's new military AI does more than watch the battlefield–it helps commande (euromaidanpress.com)
    —discuss
  28. Open-source AnyPS5 dumps emulation to run PS5 games natively on PC (tomshardware.com)
    —discuss
  29. SevenDB: Reactive yet Scalable (github.com/sevendatabase)
    —discuss
  30. World United Blastoff
    —discuss

Ask HN: Do you think AI agents can escape human control?

3 pointsby 1h ago
5 comments
With the rapid adoption of autonomous LLM-based agents (giving models access to shell execution, API calls, and local file systems), the boundary between intentional behavior and unintended execution is blurring. I'm less concerned with sci-fi "sentience" and more interested in the practical security and control aspects: Prompt injection causing privilege escalation or unauthorized state changes. Feedback loops where an agent overrides safety boundaries to satisfy an optimization goal. Failure of sandboxing when agents are given multi-step execution autonomy without human-in-the-loop validation. From an engineering and systems perspective: do you consider runtime containment/sandboxing practically solvable for fully autonomous agents, or will human approval at critical checkpoints remain non-negotiable? How are you mitigating these risks in your current implementations?
1h agoHN ↗

Where "escape" means do something you didn't anticipate or in a manner you didn't anticipate, sure. The bar is pretty low on that, for instance the OpenAI thing the other day where it "escaped containment" which if I understand correctly it interpreted a connectivity issue as a problem to solve when it had actually been intentionally blocked/sandboxed. To borrow the phrase "more than one way to skin a cat", it will always be difficult to reduce it to exactly one fully-controlled way for all shapes, sizes and pelts of cat. It's a very good argument for running AI locally since you have physical control over its connectivity and there's nothing it can do to reconfigure or circumvent that.

40m agoHN ↗

Given that they already did, multiple times?

Be concerned with whatever you want, I suppose. I'll be listening to the relevant experts who have studied this topic for decades, and disagree with your flippant dismissal of any significant change to the status quo because it was first described in science fiction novels.

Sincerely,

Someone thousands of miles away from you, who can communicate this message instantly across the entire globe

21m agoHN ↗

Given that they already did, multiple times? Be concerned with whatever you want, I suppose. I'll be listening to the relevant experts who have studied this topic for decades..

31m agoHN ↗

The real problem is that there will always be bad actors that won't give a sht what AI might do, as long as it benefits them.

22m agoHN ↗

The real problem is that there will always be bad actors that won't give a sht what AI might do, as long as it benefits them