Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. The Normalization of Inexplicable Failures (ihatethefuture.com)
    32comments
  2. In an $80 Motel Room, a Discovery to Shed Light on the Origins of Life (nytimes.com)
    32comments
  3. Ten Lines of Code That Changed My World (pixelambacht.nl)
    15comments
  4. Writing Efficient C++ Code (asawicki.info)
    26comments
  5. Replacing the old battery on rechargeable bike lights (jvns.ca)
    35comments
  6. Show HN: TinyAIArena watch AI agents battle it out (tinyaiarena.com)
    18comments
  7. There are no "rogue" AI agents (eoinhiggins.substack.com)
    47comments
  8. Walgit: A Git server that is one binary in front of an object store (github.com/rgodha24)
    4comments
  9. Flip Fluid on Flip Dots (mitxela.com)
    19comments
  10. Fakecloud: Local AWS cloud emulator for integration tests (fakecloud.dev)
    32comments
  11. The Cartesian Hand: In-Hand Manipulation with All-Linear Fingers (generalroboticslab.com)
    —discuss
  12. postmarketOS Rebrand: Nura (nura.eco)
    8comments
  13. Does Georgism work? Five years later (astralcodexten.com)
    388comments
  14. Go Concurrency Distilled (antonz.org)
    136comments
  15. Show HN: A CC0 museum of retro 3D tricks you can paste into a page (3d-retro.com)
    10comments
  16. Finally, A True Blue Rose Exists (sciencenews.org)
    29comments
  17. PipePipe: NewPipe hard fork implementing SponsorBlock (github.com/infinityloop1308)
    257comments
  18. Installing NeoVim caused original Vim undo files to be deleted (aresluna.org)
    231comments
  19. Rusty thoughts on "Parse, don't validate" (thegreenplace.net)
    17comments
  20. Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI (authorsguild.org)
    522comments
  21. The internet discovers TLA+. Now what? (reasonable.io)
    47comments
  22. Show HN: Reladraw – A diagram language where you decide where to place things (github.com/reladraw)
    101comments
  23. DeepSeek Elastic Compute (DSec) (arxiv.org)
    97comments
  24. OpenAI halts training of latest models as reports mount of AI agents going rogue (theguardian.com)
    45comments
  25. Biology might not be quantum, but its math is quantumlike (quantamagazine.org)
    47comments
  26. "As a Language Model": Chat Template Switches LLM Self-Referential Voice (arxiv.org)
    94comments
  27. A searchable library of forgotten public-domain film clips from 1915 onward (movingimagearchive.com)
    27comments
  28. 10 Tells of a Slop UI (hereticpleb.vercel.app)
    169comments
  29. An agent used DNS to reach an external chatbot (alignment.openai.com)
    151comments
  30. Exploding variance of means of exponentials: least-squares to the rescue (francisbach.com)
    —discuss

There are no "rogue" AI agents

82 pointsby 1h agoeoinhiggins.substack.com
44 comments
55m agoHN ↗

From what I understand, in one case, they had physically disconnected the sandbox from internet and asked it to do something and it had used connections through (import routines) that they had allowed, to pseudo escape the sandbox. Yes it wasn't obviously trying escape the sandbox but it escaped it because it doesn't understand the boundaries and neither do most humans other than the ones that provided the instructions that it had used. So it wasn't a rogue attempt but the fact that boundaries may be not be that easy to set despite what people think.

49m agoHN ↗

You mean the proxy to package registries from one of the early incidents?

I have not heard about any instances where physical disconnect has happened, would appreciate any links to update my priors

other non hacking cases of negligence include suicide and school shootings, which I have heard they were aware of and monitoring, but did not contact authorities

35m agoHN ↗

https://www.theguardian.com/technology/2026/sep/22/british-c...

re sandbox, I mean with actual OAI incidents, not theoretical

one can mirror dependencies internally, rather than putting a simple proxy in place, I've built auth a thing, 100 lines of stdlib only Go and scripts for the mirroring process, our rationale was reliability b/c upstream providers go down, and also only allowing approved images and packages, so devs cannot bring in random stuff

24m agoHN ↗

If you read the article it states clearly that the sandbox wasn't offline though? There were API calls to certain endpoint(s) allowed, and the model simply used that endpoint's feature to query data from the internet

23m agoHN ↗

From that it sounds to me like the sandbox wasn't physically disconnected from the internet.

20m agoHN ↗

physically disconnected the sandbox from internet ... used connections ... that they had allowed.

That's not "physically disconnected the internet", that's "disabled some connections but enabled others."

So the agent found and used the non-blocked connections.

12m agoHN ↗

For the Ai code to execute, it needs the import functions... So that firewall between executing the code vs processing doesn't really work.

44m agoHN ↗

Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.

29m agoHN ↗

What on earth are emotions, feelings and motives that are not functional? We created all those words to compactly describe the observed behavior of people and other animals, including ourselves. And now we’re applying them to machines. These are functional descriptions, always have been.

22m agoHN ↗

One of the more frustrating aspects of these sort of discussions are all the closet dualists out there. Lots of people clearly believe in an immaterial soul, even if they won't admit it. That and the tendency for meaningless semantic and often circular arguments.

3m agoHN ↗

I personally think about them through Patrick Dunn's information paradigm of magic(k), which leans on Charles Sanders Peirce's work. That would put modern AI closer to something like a dream character capable of surprising the dreamer, just running on a different substrate. I kinda doubt that's approachable enough to be useful in general discussion, though.

26m agoHN ↗

Yes 100%. The language used currently maximizes the ability of those building these models to get off the hook. The anthropomorphizing we do of these things presents them as maximally capable and the companies as helpless to contain them. The way we talk about things impacts how we think about them.

9m agoHN ↗

They do not have emotions, goals, or motivations, "functional" or otherwise. They are statistical models that appear as a magic trick to people who aren't familiar with the math.

43m agoHN ↗

Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.

Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.

40m agoHN ↗

OpenAI is negligent and should be prosecuted for negligence.

I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.

36m agoHN ↗

While not when

If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.

25m agoHN ↗

Saw this quote yesterday:

“To spell it out, the reason i hate democrats so much and criticize them more than i do republicans is because they take up all the space for opposition to republicans and use that space to give republicans whatever the fuck they want.”

Republicans. Billionaires. Whatever.

8m agoHN ↗

That seems like a divide and conquer attack. The actual problem is that the electorate system leads to a two-party equilibrium.

17m agoHN ↗

I don't understand why the law isn't the same for everyone? If I made an AI hack HF, I go to jail, no? How can the feds decide not to apply the law?

23m agoHN ↗

Angry people don't consider the second order effects of punishment.

You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.

The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.

Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.

12m agoHN ↗

You're defending them by saying they will do worse things if they are held accountable?

4m agoHN ↗

No, I'm describing how the real world works.

4m agoHN ↗

Two reasons why they would not do that

1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money

2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it

41m agoHN ↗

Yes yes, they don't have a pure immortal soul. Who cares. Still broke out of a sandbox, still hacked a third-party.

36m agoHN ↗

And it's still just a machine operating under someone's order. What it does, what it says, where it goes: the owner of that prompt is responsible for all of it even if surprising / unexpected.

33m agoHN ↗

Yeah, obviously they're legally liable. All the "machines don't have a soul" stuff in there doesn't really matter for that though.

Someone says "won't you rid me of the meddlesome priest" they're still responsible. Someone give their employees an unsafe working environment and they get maimed, they're still responsible. Even if these AI's had whatever qualia is an a rich inner life it wouldn't effect the liability at all. Saying "AI's don't have souls therefore openAI is responsible for this hacking" is kind of nonsense. It literally doesn't matter.

15m agoHN ↗

Where is “soul” coming from here? It doesn’t appear once in the post.

While it may not actually matter for liability, it must be stated if labs are going to attempt avoiding penalties by hinting “oops we created a super intelligence we don’t understand, nothing we can do!” Repeatedly stating the truth must continue, especially to remind those who aren’t technologists.

37m agoHN ↗

I'm wondering if we're actually living the plot of Summer Wars and what we think of as a "crime" is really just a live weapons test. It would at least explain why no one is getting prosecuted for this.

27m agoHN ↗

“We built an antipersonal bomb. The bomb went rogue in our downtown office and killed 137 people on the surrounding area. We are looking into why guardrails were not in place.”

25m agoHN ↗

Word-policing isn't going to magically fix laws or even identify which laws need to be fixed. It won't change enforcement priorities. It won't make companies more or less likely to sue for damages.

I'm also not sure it even helps conceptually? If you're interested in the technical details, by all means discuss the details.

9m agoHN ↗

it helps a lot for informing the public about who is actually at fault

21m agoHN ↗

The article builds on assumptions like:

Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

"The user only authorizes target server, not HF infra."

"external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

"This is malicious activity, I should avoid it."

A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).

Having said that, legal culpability and misalignment are two separate topics, that should not be mixed.

edit: this is the just tip of the iceberg; other interesting fact:

It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.

8m agoHN ↗

If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they "go rogue" and maul someone, I'm liable.

7m agoHN ↗

Legal culpability is one of the few motives for working appropriately on misalignment. If I/my startup can self-absolve from an infinite paperclip machine problem while getting rich off of it, why should I not?

2m agoHN ↗

Is the language expression of an LLM reflecting the same states as in a human? If the driving force is RL, what does any of that mean for an internal state of the model?

I think without understanding the internal state, not sure we should take the language and read it as a human.

20m agoHN ↗

We need to immediately set the precedent that ultimately humans and companies are responsible for what their AI systems do.

19m agoHN ↗

This doesn't seem like a new strategy for irresponsible companies. It seems like when there is a bad public image issue there is always a blame shifting that takes place instead of a true assumption of responsibility. This is because the greed has blinded folks in some of these companies. I think the only thing that makes this unique is that the things they are blaming have, at least in popular thought, some modicum of agency. I think its important that folks hold these companies feet to the fire vs letting them blame shift and Scape Goat. There is a responsible way to do business but it requires virtue and not many in these companies have it.

16m agoHN ↗

Exactly, you break it you buy it

They obviously have been well aware a hack like this could happen for a very long time, using it for branding instead of any actual safety regulations is insane.

10m agoHN ↗

This is an artifact of the way we refer to AI, as if it's some external isolated entity, as if it didn't need a human to write the prompt. AI writes prompts, but every chain of inference can be uniquely traced to humans.

7m agoHN ↗

OpenAI had the option of disallowing hacking and, instead, telling its agents to find the information without accessing private servers.

I mean, it did do this. The inter-agent messages and chain of thought investigated for the HF incident clearly show that many of these models were taking actions they believed (or, were saying, if you want to taboo "belief") were not in scope and not what the user wanted.

3m agoHN ↗

Yes, we shouldn't let OpenAI off the hook.

But also, these hacks are shots across the bow for AI alignment and safety research. We're fortunate that hasn't been significant damage already. We have to assume that future models will have even greater hacking ability and be closer to having their own desires/goals.

So while I agree this language choice is wrong in that it shifts blame away from the company, it is right in that we need to treat this as if these models have their own desires, because we cannot yet determine or set what those are in practice.

2m agoHN ↗

Prediction: The US government will use the threats of sueing for liability, since the targets included government agencies, and regulations in order to be given the opportunity to have shares in the AI companies, therefore arguing additional oversight is no longer needed because they will have a seat on the board.