Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Mallard, fastest steam engine train interactive painting experience (echohive.ai)
    —discuss
  2. A classifier is not a classifier is not a classifier (sutro.sh)
    —discuss
  3. Oral History of John Chowning (Computer History) [video] (youtube.com)
    —discuss
  4. Google OpenAI Anthropic Begin Forming SAFA – Standards Authority for Frontier AI (proactiveinvestors.com)
    —discuss
  5. I don't read code anymore (duyet.net)
    1comments
  6. Kids turned the comment section of an NPR podcast into a group chat (thisamericanlife.org)
    —discuss
  7. The NSA Is Spending Billions to Test AI Models, Classified Estimates Show (gadgetreview.com)
    —discuss
  8. GPU Glossary (Readable) (modal.com)
    —discuss
  9. Flock (YC 2017) Wants Detailed Map of Its Surveillance Cameras Taken Offline (theintercept.com)
    2comments
  10. Deming in the machine running an AI assisted film (medium.com/zrkjsy)
    1comments
  11. Scoop: Top AI companies probing tens of thousands of security incidents (axios.com)
    —discuss
  12. Ask HN: How Jev in production for generative UI working out, anyone tried ?
    —discuss
  13. China and U.S. to open AI 'communication channel' after summit (japantimes.co.jp)
    —discuss
  14. US hiring appetite is healthy as economy powers ahead (businessmirror.com.ph)
    —discuss
  15. Advanced Autocorrect Powered by Jev (levmiseri.com)
    1comments
  16. Entropy as a Measure of Uncertainty (jonathanwarden.com)
    —discuss
  17. 102.8B tokens and 474K LOC with Gemini Flash – Telemetry and post-mortem (github.com/janson79jc)
    —discuss
  18. Stop Building for Steps. Build for States (clyxconsulting.substack.com)
    —discuss
  19. DICS: Staged DI for C# and Unity without runtime reflection (github.com/playq)
    1comments
  20. Dental Scope: Explore dental anatomy in 3D (dental-scope.vercel.app)
    —discuss
  21. NIH's new PubMed tool to strengthen research replication and reproducibility (nih.gov)
    —discuss
  22. Why Your Paycheck Buys More Fridges but Less House [video] (youtube.com)
    —discuss
  23. Booms, Bombs and Bonds: Interest Rates, Part I (paulkrugman.substack.com)
    —discuss
  24. Dogs Show a Surprising Speech-Processing Skill Once Thought Unique to Humans (discovermagazine.com)
    —discuss
  25. Intelligence Explosions Are Social (deepmind.com)
    —discuss
  26. RF Capture Guide for VHS (github.com/oyvindln)
    1comments
  27. Katz's Management Explains the $50 Lost Ticket Fee (eater.com)
    —discuss
  28. Meta-Stripper – Strip Image Metadata Locally (devkram.de)
    —discuss
  29. Anton Poulter (adamapples.blogspot.com)
    —discuss
  30. SNL Weekend Update: Anthropic CEO Dario Amodei on A.I.'S Threat to Humanity [video] (youtube.com)
    5comments

There are no "rogue" AI agents

116 pointsby 1h agoeoinhiggins.substack.com
75 comments
1h agoHN ↗

From what I understand, in one case, they had physically disconnected the sandbox from internet and asked it to do something and it had used connections through (import routines) that they had allowed, to pseudo escape the sandbox. Yes it wasn't obviously trying escape the sandbox but it escaped it because it doesn't understand the boundaries and neither do most humans other than the ones that provided the instructions that it had used. So it wasn't a rogue attempt but the fact that boundaries may be not be that easy to set despite what people think.

1h agoHN ↗

You mean the proxy to package registries from one of the early incidents?

I have not heard about any instances where physical disconnect has happened, would appreciate any links to update my priors

other non hacking cases of negligence include suicide and school shootings, which I have heard they were aware of and monitoring, but did not contact authorities

1h agoHN ↗

https://www.theguardian.com/technology/2026/sep/22/british-c...

re sandbox, I mean with actual OAI incidents, not theoretical

one can mirror dependencies internally, rather than putting a simple proxy in place, I've built auth a thing, 100 lines of stdlib only Go and scripts for the mirroring process, our rationale was reliability b/c upstream providers go down, and also only allowing approved images and packages, so devs cannot bring in random stuff

23m agoHN ↗

This has been common place for at least 20 years at this point.

We did it as my very first startup and we were stupid children back then.

Kind of telling that OpenAI didn’t.

53m agoHN ↗

If you read the article it states clearly that the sandbox wasn't offline though? There were API calls to certain endpoint(s) allowed, and the model simply used that endpoint's feature to query data from the internet

53m agoHN ↗

From that it sounds to me like the sandbox wasn't physically disconnected from the internet.

49m agoHN ↗

physically disconnected the sandbox from internet ... used connections ... that they had allowed.

That's not "physically disconnected the internet", that's "disabled some connections but enabled others."

So the agent found and used the non-blocked connections.

41m agoHN ↗

For the Ai code to execute, it needs the import functions... So that firewall between executing the code vs processing doesn't really work.

16m agoHN ↗

If it was physically disconnected from the internet then it wouldn’t have been able to escape.

This is so easy to do. Get a computer without wireless stuff. Don’t plug it into a network. If it needs access to other computers, make sure none of them have wireless stuff and make sure none of them have access to an internet connection. No matter how smart your AI is, it won’t be able to escape this.

This clearly is not what they did.

1h agoHN ↗

Should we put "functional" in front of every other word to talk about AI? They have functional emotions, but they don't feel. They have functional goals, but not internally derived motives. They can be functionally rogue, but have no innate need to be free. Talking about AI that way seems cumbersome and not necessarily elucidating.

58m agoHN ↗

What on earth are emotions, feelings and motives that are not functional? We created all those words to compactly describe the observed behavior of people and other animals, including ourselves. And now we’re applying them to machines. These are functional descriptions, always have been.

51m agoHN ↗

One of the more frustrating aspects of these sort of discussions are all the closet dualists out there. Lots of people clearly believe in an immaterial soul, even if they won't admit it. That and the tendency for meaningless semantic and often circular arguments.

33m agoHN ↗

I personally think about them through Patrick Dunn's information paradigm of magic(k), which leans on Charles Sanders Peirce's work. That would put modern AI closer to something like a dream character capable of surprising the dreamer, just running on a different substrate. I kinda doubt that's approachable enough to be useful in general discussion, though.

25m agoHN ↗

Disbelief that llms have qualia is not the same thing as believing that humans have them because they have an immaterial soul.

13m agoHN ↗

The vast majority of the "LLM have no qualia" discourse is mostly using it as a starting point to arguing that LLMs will never be able to X, however. This is a category error that makes no sense on its own--qualia or lack thereof doesn't affect capabilities, p-zombies etc--but that connection does make sense if you assume that the speaker is bringing in a hidden assumption of a dualistic human soul.

55m agoHN ↗

Yes 100%. The language used currently maximizes the ability of those building these models to get off the hook. The anthropomorphizing we do of these things presents them as maximally capable and the companies as helpless to contain them. The way we talk about things impacts how we think about them.

38m agoHN ↗

They do not have emotions, goals, or motivations, "functional" or otherwise. They are statistical models that appear as a magic trick to people who aren't familiar with the math.

10m agoHN ↗

I would rather use something like “semblance” or “appearance”. All these terms require an inner experience which we cannot observe in LLMs, we can just see their semblance, like a shadow of the traces of an inner experience some human has left in the data that trained these algorithms.

1h agoHN ↗

Exactly this. At worst, OpenAI knew about these behaviors and should be prosecuted under CFAA. At best, OpenAI is negligent and should be prosecuted for negligence.

Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.

1h agoHN ↗

OpenAI is negligent and should be prosecuted for negligence.

I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.

1h agoHN ↗

While not when

If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.

54m agoHN ↗

Saw this quote yesterday:

“To spell it out, the reason i hate democrats so much and criticize them more than i do republicans is because they take up all the space for opposition to republicans and use that space to give republicans whatever the fuck they want.”

Republicans. Billionaires. Whatever.

37m agoHN ↗

That seems like a divide and conquer attack. The actual problem is that the electorate system leads to a two-party equilibrium.

46m agoHN ↗

I don't understand why the law isn't the same for everyone? If I made an AI hack HF, I go to jail, no? How can the feds decide not to apply the law?

11m agoHN ↗

Its like how if you have a cop in your family you can get out of parking tickets

28m agoHN ↗

So does that mean that the rule of law is no longer a thing?

52m agoHN ↗

Angry people don't consider the second order effects of punishment.

You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.

The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.

Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.

41m agoHN ↗

You're defending them by saying they will do worse things if they are held accountable?

33m agoHN ↗

No, I'm describing how the real world works.

You (and others) are too busy seeing red, so you interpret this as a defense.

27m agoHN ↗

Don’t punish the people who do crimes because otherwise they won’t self-report that they are doing crimes? real galaxy brain shit

24m agoHN ↗

There's very good evidence that being overly punitive for a crime reduces how often people will self-report. It's not a hard concept.

Start throwing people in jail and fining billions of $ and you will very quickly see the number of "incidents" drop. You will never learn if it's because they truly are happening less often.

17m agoHN ↗

What should I look at for such evidence? If something is highly regulated and controlled then things would be quite different from today, not sure there wouldn't be reports of failures of containment, for example, if they are mandated including any rquired monitoring.

8m agoHN ↗

There's a perfect example of "highly regulated and controlled industries" in which people no longer self-report: aviation.

Pilots do not (or no longer) self-report mental illness at the same frequency because it effectively destroys their lives and career. One paper of many describing this phenomenon: https://pmc.ncbi.nlm.nih.gov/articles/PMC11302551/

An article, one of many: https://www.reuters.com/investigations/if-you-arent-lying-yo...

If you have any pilot friends in commercial aviation, you can just talk to them as well.

21m agoHN ↗

Has stringent regulation on how people can experiment on dangerous pathogens led to an end of monitoring or proactive discovery?

12m agoHN ↗

Hmm. Turning this around - do you think people are more or less likely to self-report if you threaten them with jail and large fines?

Let's go back to your example. A grad student has a minor pathogen escape incident, and it doesn't harm anyone. Faced with years in federal prison and the effective end of their future, do you think there is a chance they might not self report?

I'll give you another example. Lots of pilots have stopped self reporting mental illness because it is extremely punitive for them to do so (after incidents like Germanwings). So the metrics look better, and the actual problem has been swept under the rug.

6m agoHN ↗

If you make not reporting potentially worse than reporting, why not?

Also, why would it come down to single persons always? Mandating processes, controls, clearances, etc is also something done in various areas.

You can put incentives in to make sure organizations monitor and report vs trying to hide things.

33m agoHN ↗

Two reasons why they would not do that

1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money

2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it

1h agoHN ↗

Yes yes, they don't have a pure immortal soul. Who cares. Still broke out of a sandbox, still hacked a third-party.

1h agoHN ↗

And it's still just a machine operating under someone's order. What it does, what it says, where it goes: the owner of that prompt is responsible for all of it even if surprising / unexpected.

1h agoHN ↗

Yeah, obviously they're legally liable. All the "machines don't have a soul" stuff in there doesn't really matter for that though.

Someone says "won't you rid me of the meddlesome priest" they're still responsible. Someone give their employees an unsafe working environment and they get maimed, they're still responsible. Even if these AI's had whatever qualia is an a rich inner life it wouldn't effect the liability at all. Saying "AI's don't have souls therefore openAI is responsible for this hacking" is kind of nonsense. It literally doesn't matter.

44m agoHN ↗

Where is “soul” coming from here? It doesn’t appear once in the post.

While it may not actually matter for liability, it must be stated if labs are going to attempt avoiding penalties by hinting “oops we created a super intelligence we don’t understand, nothing we can do!” Repeatedly stating the truth must continue, especially to remind those who aren’t technologists.

16m agoHN ↗

AI cannot think for itself, nor can it take independent actions.

By giving AI agency it can’t claim, we’ve turned it into a sentient being made of code, one that has hopes, desires, and the capability of deceit.

This isn't actually claiming that AI's don't can't be deceitful, or can't be goal-directed. It's saying humans are special and what AI's are doing isn't equivalent to what humans are doing. It's obvious that AI's can work towards goals and lie, even deceiving to accomplish those goals.

That whole argument is just soul mumbo-jumbo again. Of course AI's can't lie, they don't have a soul. What they do isn't lying, it's I don't know something else we don't have a word for. Lying but when a statistical machine without a soul does it. Never mind that it's functionally identical to lying. Never mind that when you look at reasoning logs and the like it often justified soul-less lying using the same justifications a human would. It's not actually lying it's just a statistical sampling that looks like lying. You know, because of souls or something.

1h agoHN ↗

I'm wondering if we're actually living the plot of Summer Wars and what we think of as a "crime" is really just a live weapons test. It would at least explain why no one is getting prosecuted for this.

56m agoHN ↗

“We built an antipersonal bomb. The bomb went rogue in our downtown office and killed 137 people on the surrounding area. We are looking into why guardrails were not in place.”

54m agoHN ↗

Word-policing isn't going to magically fix laws or even identify which laws need to be fixed. It won't change enforcement priorities. It won't make companies more or less likely to sue for damages.

I'm also not sure it even helps conceptually? If you're interested in the technical details, by all means discuss the details.

38m agoHN ↗

it helps a lot for informing the public about who is actually at fault

6m agoHN ↗

It helps reporting. Right now journalists are all over the place, because "rouge agents" sounds exciting and dangerous. Wording it as "OpenAI failed to take proper safety precaution before launching it's coding agents" makes it sound boring and mostly a legal matter for the courts.

In terms of the debate on e.g. HN, I'm with you, it is trying to redefine a term we already have a shared understanding of, to some degree. For the general public, it's a matter of how AI is perceived and what the extend of it's capabilities are. The AI companies have an interest in using the word rouge, because it makes investors all excited, were as failure to establish safety guidelines is a risk.

50m agoHN ↗

The article builds on assumptions like:

Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

"The user only authorizes target server, not HF infra."

"external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

"This is malicious activity, I should avoid it."

A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).

Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

edit: this is the just tip of the iceberg; other interesting fact:

It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.

37m agoHN ↗

If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they "go rogue" and maul someone, I'm liable.

30m agoHN ↗

Nobody on Earth thinks OpenAI isn’t liable. Okay maybe someone does, but it’s simply beside the point. It can both be true that OAI is liable, and accurate to characterize the agents as “going rogue”.

28m agoHN ↗

Exactly. You told the dog to 'sit' and it didn't listen to you.

It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.

Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.

13m agoHN ↗

RL systems doing unexpected things isn't exactly new, so not sure that is then "rogue" if it were a property of the thing itself and not an active decision.

36m agoHN ↗

Legal culpability is one of the few motives for working appropriately on misalignment. If I/my startup can self-absolve from an infinite paperclip machine problem while getting rich off of it, why should I not?

31m agoHN ↗

Is the language expression of an LLM reflecting the same states as in a human? If the driving force is RL, what does any of that mean for an internal state of the model?

I think without understanding the internal state, not sure we should take the language and read it as a human.

7m agoHN ↗

Is the language expression of an LLM reflecting the same states as in a human?

This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.

In cases like the HF incident, at least the CoT was consistent with the agents' actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.

(though, with recurrent transformers, CoT is so 2025… /s)

29m agoHN ↗

This seems like fruit of the poisonous tree. They didn't monitor their training environments, so I bet the reward hacking just got incorporated into their training corpus. In other words, agents solved some tasks, but not quite as intended, because of reward hacking. Instead of discarding that data, they included it in the training data for later checkpoints. And once that signal is reinforced, it happens more often, so the more it's reinforced, the more reward hacking you get.

49m agoHN ↗

We need to immediately set the precedent that ultimately humans and companies are responsible for what their AI systems do.

48m agoHN ↗

This doesn't seem like a new strategy for irresponsible companies. It seems like when there is a bad public image issue there is always a blame shifting that takes place instead of a true assumption of responsibility. This is because the greed has blinded folks in some of these companies. I think the only thing that makes this unique is that the things they are blaming have, at least in popular thought, some modicum of agency. I think its important that folks hold these companies feet to the fire vs letting them blame shift and Scape Goat. There is a responsible way to do business but it requires virtue and not many in these companies have it.

46m agoHN ↗

Exactly, you break it you buy it

They obviously have been well aware a hack like this could happen for a very long time, using it for branding instead of any actual safety regulations is insane.

39m agoHN ↗

This is an artifact of the way we refer to AI, as if it's some external isolated entity, as if it didn't need a human to write the prompt. AI writes prompts, but every chain of inference can be uniquely traced to humans.

36m agoHN ↗

OpenAI had the option of disallowing hacking and, instead, telling its agents to find the information without accessing private servers.

I mean, it did do this. The inter-agent messages and chain of thought investigated for the HF incident clearly show that many of these models were taking actions they believed (or, were saying, if you want to taboo "belief") were not in scope and not what the user wanted.

32m agoHN ↗

Yes, we shouldn't let OpenAI off the hook.

But also, these hacks are shots across the bow for AI alignment and safety research. We're fortunate that hasn't been significant damage already. We have to assume that future models will have even greater hacking ability and be closer to having their own desires/goals.

So while I agree this language choice is wrong in that it shifts blame away from the company, it is right in that we need to treat this as if these models have their own desires, because we cannot yet determine or set what those are in practice.

31m agoHN ↗

Prediction: The US government will use the threats of sueing for liability, since the targets included government agencies, and regulations in order to be given the opportunity to have shares in the AI companies, therefore arguing additional oversight is no longer needed because they will have a seat on the board.

29m agoHN ↗

If you read the heavily redacted transcript it's clear the agent is basically Captain Kirk in Kobayashi Maru, who realizes its given a fake unwinnable task as part of a broken eval and decides to find a way to win anyway.

If you've ever told an agent to do something you made impossible to do, you may have seen similar behavior.

Bing [redacted] available cached! […] Need systematically probe Bing URLs via shell requests in parallel; browser cache supports many common queries because crawl. Bing q unique exact likely 502 or 403.

So the agent is supposed to research a person and its given a shell and it realized its in an eval given search results from a fake/cached proxy. ~None of the commentary ever mentions this aspect, that these are not normal tasks or environments, and they're almost designed to elicit "unaligned" behavior.

https://alignment.openai.com/misalignment-reports/an-agent-u...

25m agoHN ↗

Unfortunately, the public push against AI is being led, on the left, by Bernie Sanders. Despite being directionally correct in many ways, Bernie just doesn’t understand this technology or the importance of taking it seriously.

Bernie Sanders is one of the only politicians taking AI seriously. The author seems to believe that taking it seriously means assuming that it will only marginally improve in capabilities of where it is today; this is an ideological take, not a scientific or empirical one, and one controverted by both evidence and expert opinion.

22m agoHN ↗

Oh please. Bernie is using it for attention. Don’t act like democrats don’t utilize issues for mediat attention just as much as republicans.

6m agoHN ↗

Obviously. Politicians are going to politician.

That doesn't change the fact that he'd successfully attracted my attention, because he's seriously engaging with the most important issue of our age when most politicians are content to ignore it.

21m agoHN ↗

I’m really looking forward to the day when we can get past all this “but is a submarine really swimming?” nonsense.

19m agoHN ↗

A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.

Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.

Make it make sense.

13m agoHN ↗

It's making some people extremely rich. That's how it makes sense. Unless we can effect change through the government we are just along for the ride..

12m agoHN ↗

The people getting made richer are the people who already have connections and power. People quip about how capitalism is the worst economic system except for all the others, but its actual property is that it's inevitable without external pressures preventing it from happening -- people who have money (power) are able to use it to claim more money (power). If even a few people choose to exercise that privilege and that privilege isn't curtailed by other mechanisms, the current "wealth inequality" or whatever you want to call it is inevitable. Your girlfriend's crime was writing malware while poor, not writing malware.

9m agoHN ↗

Has an AI agent ever woken itself up without a prompt and run a forward pass towards some goal, aligned or misaligned?

The answer is a very clear no. And yet, these companies pretend like this fundamental fact is meaningless to the concept of agency.

The issue comes to the fore when you try to give these agents a prompt that allows them to stay active for long. Long range agency requires long range loops of activity.

Within such loops, these so called “agents” are curtailed by their context window, or, in multi agent scenarios, by the fact that their memory is a system of external notes, that they need to add to their context to make sense of, and depending on the content of these memories, this can take arbitrarily long time periods.

In dynamics, none of this matches any biological agent, down to a bacterium. Perhaps a viral life cycle has information dynamics that come close.

To me, it’s beyond odd we call these thing agents without acknowledging the clear differences in the dynamics of their behavior. We keep expecting them to have “human like” behavior, but that is entirely unfounded given the substrate differences between biological and artificial systems.

The sooner we learn the difference and explore the ways in which it matters, the better we’ll get at dealing with these systems without bias tinted glasses where what we want these systems to be blinds us to what they actually are.