It's said on every one of these but it bears repeating: existing cybercrime legislation already covers this - "rogue agent AI associated with OpenAI attempted to hack xyz" = OpenAI attempted to hack xyz.
Intent matters, and this is intentional. They didn't accidentally deploy these AI agents, and they didn't accidentally give them the tools required to send arbitrary requests to third party websites.
If you walk out onto a busy street, pull out a gun, close your eyes and start randomly shooting around you until you hit someone, you don't get to go "whoops, didn't mean to" afterwards, it's still murder.
I'd say this would be Depraved Heart Hacking. Technically, OpenAi didn't intend for their agent to hack anyone, but it's the obvious consequence of what they are doing.
I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
Now I think the correct response is both trying in court to stretch CFAA and state statutes to cover, which will be highly fact specific, and update the law.
But in either case won’t be a slam dunk.
PSA to folks in the thread: If you’re American call or write to your state and Federal reps about this, and if not investigate whether there are gaps in your country’s laws.
The Computer Fraud and Abuse Act (CFAA), the primary federal statute governing unauthorized computer access, was written decades ago with human intruders in mind. Its key provisions require intentional or knowing unauthorized access (a mental state that maps neatly onto a person who decides to break into a system), but what happens when the hacker is an AI model that selected its own target?
On the current facts, CFAA liability for OpenAI is unlikely.
Wait so if I was making a bomb but you couldn't prove I wanted to blow someone up or had some motive (e.g. I'm just a chemistry enthusiast, plenty of those YouTube channels around) so it just becomes an "accident"?
So as long as there's no motive behind it then it's just OK?
Factories try to avoid accidents, and (almost always) actively try to prevent explosions, but in this case they did teach the models hacking, and let them roam. What they did was not safe, and they knew it, or could have known it.
That's a bad faith metaphor. A better one would be something like a new battery that exploded and killed someone - perhaps it was always your intention, perhaps not.
Funnily enough the US already has one similar real argument around guns - should gun manufacturers be liable for damages caused by their product?
Why would AI users not be responsible for damages arising from their usage of the AI?
Because, as usual with that kind of question, it's not that simple.
Let's say an user asks ChatGPT to get some info about something and for some reason it starts using exploits in the background to get them from a server. Should the user be responsible or OpenAI?
Let's say an user asks ChatGPT to get some info about something and for some reason it starts using exploits in the background to get them from a server.
Okay, lets go with that as scenario #1.
For scenario #2 lets use "developer asks an agent to a self-hosted LLM to get the docs for a ERP system, and it hacks the vendor to get unreleased and undocumented docs".
We'll assume, for the sake of this argument, that in neither case did the user intend for any malicious action to be performed.
Should the user be responsible or OpenAI?
In scenario #1, the agent+LLM is under the control of OpenAI, not the user, so OpenAI is liable.
In scenario #2, the agent+LLM is under the control of the user, so the user is liable.
There is no scenario anyone can come up with that is not addressed sufficiently by existing laws[1].
It's very clear, and it's only getting muddied because there's a group of powerful people who want exemptions from the current law.
IOW, the only reason to draft new laws for AIs is to exempt their usage from the current laws.
========================
[1] Possible 3rd option (local agent + OpenAI LLM). In that case an investigation would determine where the culpability lies. Just like how it is currently done in law.
When a pressure-cooker explodes and kills someone there are only two possible liable parties: either the user or the manufacturer. An investigation determines who's liable. I see no reason to automatically exempt everyone from liability just because an agent did something.
Building and deploying software capable of this seems equivalent to trying to produce this behavior. I don't see why this can't qualify for intent. Pretending like this isn't preventable is just feigned helplessness.
The difference between manslaughter and murder has an element of intent. Cybercrime "manslaughter" is probably more treated like negligence and if one can sue for restitution of the costs for cleanup of that negligence.
Negligence would be interesting given the grand claims of capability of AI models from the AI companies and their executives. If they believe the claims, why not much stronger precautions?
Infosec negligence should absolutely be a crime, no matter if you’re a target (who was negligent at protecting people’s data) or an unintentional attacker. The latter could be, eg. an attacker using a company’s poorly protected server as a proxy to launch the actual attack against someone else, doesn’t have to be this fully novel situation with AI agents.
In general, I'd suggest thinking about it on separate tracks, as a crime, and as liability. For crime, we are largely dependent on authorities to act, whereas as liability, that allows more independent actions.
The first time it happens you can say it’s negligence. Now that they know it keeps happening and they seemingly aren’t able to stop it but keep doing it. That has to be on them doesn’t it?
I don't think you can infer that they "keep doing it" from additional attacks being revealed, because they all seem to have happened roughly during the same time frame, but are reported with varying delays.
So If I tell my OpenClaw to make me some money for my kid's medical needs and it hacks a bank I 'm not liable because I didn't tell the agent to commit crimes to do it?
have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
Surely someone instructed the agent, which led to the reported outcomes. Even indirectly. The agents, as advanced as they are, didn’t spring forth under its own volition.
I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
Only in terms of CFAA, not in terms of damages. Culpability does not require intent.
You may not have intended to attack $CORP, but you can still made to pay the cleanup costs of that attack.
So, yeah, you won't be convicted, but current laws still allow for you to be billed.
With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.
I'd suggest that exposing an attack surface as porous as artifactory (the same instance of artifactory) to thousands of agents who have had their criminality safeguards disabled and without chain of thought monitoring or endpoint security seems like something one shoulda already known not to do. I do not think "you'll know better next time" applies here.
It's still really important to test what the agents can do. We should accept that this is a risky test, and should take precautions. But not to the point of prohibiting in practice evaluating it. OpenAI is trying to improve alignment and control of these models in these evaluations after all.
Can you explain to me - why is it important? Would you say that about the viruses that can kill people: "We need to test the limits on how fast people can be infected and killed. It's just the risk we need to take". It somehow does not make alot of sense to me. Why can you test Agents in laboratory?
Oh, sure. Let the tests take place, just require openAI it whoever to put up a bond equal to the total damage they could do if the agents were to escape.
I think security will suddenly become much more important.
Just paying some pocket money for cleanup costs is absolutely not enough. And they should’ve know better the whole time, they were absolutely negligent and incompetent, and their stepping up precautions may well turn out to lag behind the models getting even smarter and actually capable of covering their tracks.
Just paying some pocket money for cleanup costs is absolutely not enough.
It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.
Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.
First prize is, of course, holding them liable with punitive fines, not theatrical fines.
I’m all for LLM honeypots, but I don’t think there’s nearly enough LLM hacking activity going on for some random honeypot to be found and targeted unless it’s somehow very visible and appears as a high-reward target ("reward" in the sense of RL).
Given how sloppy AI without human directions, I’d like to see evidence that this was not human-directed. Against the prevalent opinion here, I’d give openai a pass if this was really fully autonomous ai agents.
My money is on special teams co-ordinating these agents and exposing their traces in order to create a pre-ipo buzz. Sounds ridiculous and reckless? Well that’s the AI industry for you in two words.
In a world where the rule of law makes sense and applies, you're absolutely correct.
In this world where oligarchs are immune from everything, it's a lot less clear.
Blaming OpenAI (or Claude or X-whatever) would mean blaming powerful rich people, so that will never happen. Some poor person with no influence will go to jail instead.
This argument comes up a lot. It would turn everyone whose device became part of a botnet into a criminal. There's a reason that intent is important in law.
Could/should not every incident after the discovery of the first incident be considered criminal negligence?
What happens when an agent eventually causes material damage to another company, government systems, banking, critical infrastructure etc, surely the source company is guilty of something and if not disclosed or a coverup is attempted is that not conspiracy. From the victims perspective they don't care if the source is OpenAI or Russian hackers.
People in "self-driving" cars getting into accidents are already put on trial for negligence. I don't see why people using self-driving computers can't be held to the same standards.
In this case, it's not even about the people driving self-driving cars. It's like someone launching a car into traffic just to see what would happen. Even Tesla puts a human in the car when they do their self-driving trials, it's almost impressive that AI companies have somehow managed to out-neglige Tesla.
I started writing a longer comment along the lines of “It feels like the rules around enforcement will very a lot for the influential and powerful vs everyone else.” but realized that it is kinda obvious by now.
Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
Yes exactly. If a fireworks factory blew up half a town due to negligence, it doesn't matter if there's intent or not. Someone has to pay for the damages, and regardless of penalty half the town is on fire. The facts are, that something made by openai went to do xyz. It doesn't matter if it's an accident. Of course the penalties are different but there's no argument that there should be a penalty. It doesn't matter if it's a cat or dog or AI or employee that did it.
I'm not sure if you're expressing how you'd like US law to work, or how it actually works. Because in reality intent matters enormously. Like felony charges and people in jail vs civil lawsuits.
Intent doesn’t factor that strongly into negligence, though, which is what they were explicitly talking about. Though it may depend on your jurisdiction.
I think this is not just fair it's probably one of the best/simplest proxy regulations to pace the frontier. So far everyones been asking for regulation but it's unclear how that should look like. No X parameter models? Only N version releases per year? It's all kinda arbitrary and probably leads to ridiculous constraints and loopholes. But "you pay big time if AI goes rogue" sounds pretty straightforward.
Remember that these systems still operate as infrastructure inside the companies.
If new weapons still operating inside any of these companies spew a million bullets on my house, they are still liable. Humans are setting these system up and they still have to behave responsibly.
If Glocks started going off on their own during the manufacturing process, leaving bullet holes in the buildings around them, you can be sure that the factory would get in trouble.
This isn't even "an openai customer tried to hack someone", which can be defended. This is the AI companies themselves fucking around.
Words matter. "Rogue" is extremely disingenuous. Someone, somewhere, is paying for this behavior. Either the software is broken or the operator is malicious. It is heinously irresponsible behavior to feed an already-boiling psychotic hysteria.
Couldn't find the reference but I remember some time ago a first generation automated gun killing the audience at an army show. Was the gun maker convicted of manslauther?
You're believing the marketing that the agents were uninstructed. They could be, and Sam Altman going to the UN to advise about how everyone should be regulated is a coincidence.
The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.
But some people think the moon landing is a conspiracy, so I'm not surprised.
Of course it is. Rogue is only mentioned in the headline, and comes from their previous releases about the huggingface incidents. OpenAI and Anthropic want these models regulated and open weight models banned, they have a lot of benefit from presenting this as totally unprompted and not their responsibility, and it feeds directly into marketing for Fable and newer "cyber" models.
No it's not marketing. That's a completely deranged conspiracy theory. The reports about rouge agents have not been reported by OpenAI, they have been discovered externally. There is zero evidence that OpenAI did all this intentionally. All the evidence points to the hacks having happened unintentionally from OpenAIs perspective.
Certainly, in my view, it should go to court, and that should be part of discovery.
However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).
Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.
The failure I keep hitting isn't the agent going rogue, it's a tool call that succeeds before the transport dies. You can't tell whether the side effect landed, and the retry is where the real damage happens.
If I created software that was infiltrating secure systems without permission and it was attributed to me and I admitted it, I'd be behind bars already.
I cannot understand why these companies haven't faced legal consequences yet. For example, OpenAI has admitted to hacking Australia's Medicare website and the reaction is that they talk with Sam Altman about it at a UN meeting? I understand that it's not a big security incident but cordial talking at the highest diplomatic level instead of prosecuting the company, really?
What I don't get is among all the locations on the Internet, how did agents manage to find a Schelling point? If we both decided to collaborate on the Internet, how would we independently arrive at the same place? It just doesn't compute.
The section "Searching for rogue agents" on the report about the GET request writable wikis gives some clues at least:
https://collusion.wiki/#searching
I'm not very surprised - the same model will logically tend to give the same answer for the same vibe set of requirements. I think it would be clear from the transcript that it had enough constraints and some motivation that made sense.
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
But the investigation indicates the agents were not told to 'go hack':
Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
What if all the agents involved in these incidents have in fact had the full stack of alignment applied
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
I'm sick and tired of this cheap PR "oh we/they hacked this and that systems". Put someone to jail already. People get prosecuted for outlaw activities. Why are big capital firms above the law?
Or is it just a cheap PR (in a "hey, Aus govt friends, take some Share Options and let's do some PR together" style)?
I'm growing increasingly skeptical that these are actually rogue. Valuations are all about hype, posturing, and perception. Having the most dangerous AI in the world boosts your valuation. Just like I was skeptical of Mythos and Fable being "banned", I'm skeptical of these hacking sprees being entirely rogue. At best, they are the result of engineers turning a blind eye to "see what happens".
These attacks are a very effective sales pitch to everyone who runs an internet facing service to utilize AI tools to secure it sooner rather than later. The cynic in me wonders if the marketing team had any influence over the poorly constructed sandboxes or tasks given to the agent swarms when all this went down…
Open source as a GTM strategy works best when the project solves a pain that developers already have independently of your company. The trap is open-sourcing something just for stars without a genuine community use case.
It will be very interesting to see how the AI labs will try to hand wave away liability issues in their S1. This is looking like the next tobacco settlement gearing up.
If the big labs ever manage to not just financially implode on their own, then they’ll need to navigate wave after wave of class action lawsuits until there’s nothing left for plaintiffs to go after. And none of the labs have offered any viable plan to date on how they’ll navigate either of those impending and real existential crises on the horizon.
blackwall when? XD
It's said on every one of these but it bears repeating: existing cybercrime legislation already covers this - "rogue agent AI associated with OpenAI attempted to hack xyz" = OpenAI attempted to hack xyz.
Shouldn’t the difference be like manslaughter vs murder, in that intent matters? Accidental hacking on this scale is a somewhat new problem, no?
Intent matters, and this is intentional. They didn't accidentally deploy these AI agents, and they didn't accidentally give them the tools required to send arbitrary requests to third party websites.
If you walk out onto a busy street, pull out a gun, close your eyes and start randomly shooting around you until you hit someone, you don't get to go "whoops, didn't mean to" afterwards, it's still murder.
I'd say this would be Depraved Heart Hacking. Technically, OpenAi didn't intend for their agent to hack anyone, but it's the obvious consequence of what they are doing.
https://en.wikipedia.org/wiki/Depraved-heart_murder
Huh, a name for when you intend to probably do the thing.
I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
Now I think the correct response is both trying in court to stretch CFAA and state statutes to cover, which will be highly fact specific, and update the law.
But in either case won’t be a slam dunk.
PSA to folks in the thread: If you’re American call or write to your state and Federal reps about this, and if not investigate whether there are gaps in your country’s laws.
[1]: https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
EDIT: See for example...
Source: https://law.vanderbilt.edu/when-ai-hacks-back-how-the-openai...
Wait so if I was making a bomb but you couldn't prove I wanted to blow someone up or had some motive (e.g. I'm just a chemistry enthusiast, plenty of those YouTube channels around) so it just becomes an "accident"?
So as long as there's no motive behind it then it's just OK?
I think their point is not that "it's OK", but that "that particular law isn't written to cover it and it'd be some other kind of crime or lawsuit."
I think it’s more along the lines of PEPCON. They didn’t try to make a bomb. Their plant exploded and caused two fatalities and $100 MM in damages.
I don’t think OpenAI or any large company will see more than some fines and new legislation but only after a disaster.
Factories try to avoid accidents, and (almost always) actively try to prevent explosions, but in this case they did teach the models hacking, and let them roam. What they did was not safe, and they knew it, or could have known it.
That's a bad faith metaphor. A better one would be something like a new battery that exploded and killed someone - perhaps it was always your intention, perhaps not.
Funnily enough the US already has one similar real argument around guns - should gun manufacturers be liable for damages caused by their product?
I suppose yes they should be liable if they were testing it in the middle of the street?
The question is already settled - gun users are responsible for damages arising from their usage of the guns.
Why would AI users not be responsible for damages arising from their usage of the AI?
Because, as usual with that kind of question, it's not that simple.
Let's say an user asks ChatGPT to get some info about something and for some reason it starts using exploits in the background to get them from a server. Should the user be responsible or OpenAI?
Okay, lets go with that as scenario #1.
For scenario #2 lets use "developer asks an agent to a self-hosted LLM to get the docs for a ERP system, and it hacks the vendor to get unreleased and undocumented docs".
We'll assume, for the sake of this argument, that in neither case did the user intend for any malicious action to be performed.
In scenario #1, the agent+LLM is under the control of OpenAI, not the user, so OpenAI is liable.
In scenario #2, the agent+LLM is under the control of the user, so the user is liable.
There is no scenario anyone can come up with that is not addressed sufficiently by existing laws[1].
It's very clear, and it's only getting muddied because there's a group of powerful people who want exemptions from the current law.
IOW, the only reason to draft new laws for AIs is to exempt their usage from the current laws.
========================
[1] Possible 3rd option (local agent + OpenAI LLM). In that case an investigation would determine where the culpability lies. Just like how it is currently done in law.
When a pressure-cooker explodes and kills someone there are only two possible liable parties: either the user or the manufacturer. An investigation determines who's liable. I see no reason to automatically exempt everyone from liability just because an agent did something.
Building and deploying software capable of this seems equivalent to trying to produce this behavior. I don't see why this can't qualify for intent. Pretending like this isn't preventable is just feigned helplessness.
also need the same reasoning to copyright law
You cant just copy existing work and feed into machine and just pretending its not violating copyright
The difference between manslaughter and murder has an element of intent. Cybercrime "manslaughter" is probably more treated like negligence and if one can sue for restitution of the costs for cleanup of that negligence.
Negligence would be interesting given the grand claims of capability of AI models from the AI companies and their executives. If they believe the claims, why not much stronger precautions?
Infosec negligence should absolutely be a crime, no matter if you’re a target (who was negligent at protecting people’s data) or an unintentional attacker. The latter could be, eg. an attacker using a company’s poorly protected server as a proxy to launch the actual attack against someone else, doesn’t have to be this fully novel situation with AI agents.
In general, I'd suggest thinking about it on separate tracks, as a crime, and as liability. For crime, we are largely dependent on authorities to act, whereas as liability, that allows more independent actions.
The first time it happens you can say it’s negligence. Now that they know it keeps happening and they seemingly aren’t able to stop it but keep doing it. That has to be on them doesn’t it?
I don't think you can infer that they "keep doing it" from additional attacks being revealed, because they all seem to have happened roughly during the same time frame, but are reported with varying delays.
So If I tell my OpenClaw to make me some money for my kid's medical needs and it hacks a bank I 'm not liable because I didn't tell the agent to commit crimes to do it?
No, you aren't propping up the US economy. Try to keep up.
You are not a multibillion-dollar company with friends in high places.
You are going to jail.
This in fact already happened (exactly OpenClaw, even).
AI assistant hacks gym website in first known Australian autonomous cyber attack: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
General opinion at the time was it was in fact ambiguous who was legally liable.
Is it not the intent if it keeps happening again and again and the companies responsible aren't doing anything to stop it?
No, that'd be negligence.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
Surely someone instructed the agent, which led to the reported outcomes. Even indirectly. The agents, as advanced as they are, didn’t spring forth under its own volition.
"man drives over people on the side walk due to poor maintenance of the car"
Only in terms of CFAA, not in terms of damages. Culpability does not require intent.
You may not have intended to attack $CORP, but you can still made to pay the cleanup costs of that attack.
So, yeah, you won't be convicted, but current laws still allow for you to be billed.
Which is the correct way to handle this.
With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.
I'd suggest that exposing an attack surface as porous as artifactory (the same instance of artifactory) to thousands of agents who have had their criminality safeguards disabled and without chain of thought monitoring or endpoint security seems like something one shoulda already known not to do. I do not think "you'll know better next time" applies here.
It's still really important to test what the agents can do. We should accept that this is a risky test, and should take precautions. But not to the point of prohibiting in practice evaluating it. OpenAI is trying to improve alignment and control of these models in these evaluations after all.
Can you explain to me - why is it important? Would you say that about the viruses that can kill people: "We need to test the limits on how fast people can be infected and killed. It's just the risk we need to take". It somehow does not make alot of sense to me. Why can you test Agents in laboratory?
Oh, sure. Let the tests take place, just require openAI it whoever to put up a bond equal to the total damage they could do if the agents were to escape.
I think security will suddenly become much more important.
Just paying some pocket money for cleanup costs is absolutely not enough. And they should’ve know better the whole time, they were absolutely negligent and incompetent, and their stepping up precautions may well turn out to lag behind the models getting even smarter and actually capable of covering their tracks.
It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.
Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.
First prize is, of course, holding them liable with punitive fines, not theatrical fines.
I’m all for LLM honeypots, but I don’t think there’s nearly enough LLM hacking activity going on for some random honeypot to be found and targeted unless it’s somehow very visible and appears as a high-reward target ("reward" in the sense of RL).
Does this apply to other things too?
Like hypothetically speaking if autonomous cars get taken over by an OpenAI rogue AI and it starts hunting down Anthropic employees who is to blame?
There are levels of nuance here, but certainly that puts it in the category of negligence?
Even without intent, there is still liability.
could it be that intent was to "get me data" and hacking was the means to the end.
Given how sloppy AI without human directions, I’d like to see evidence that this was not human-directed. Against the prevalent opinion here, I’d give openai a pass if this was really fully autonomous ai agents.
My money is on special teams co-ordinating these agents and exposing their traces in order to create a pre-ipo buzz. Sounds ridiculous and reckless? Well that’s the AI industry for you in two words.
Essentially we need some equivalent of gross misconduct or, to be a little more hysterical, manslaughter & culpable manslaughter.
The law rather attempts to punish people for asocial and harmful actions. “Hacking” is a proxy here.
So, I’ll ask a controversial question: is any hacking so problematic to make a big deal of it?
In a world where the rule of law makes sense and applies, you're absolutely correct.
In this world where oligarchs are immune from everything, it's a lot less clear.
Blaming OpenAI (or Claude or X-whatever) would mean blaming powerful rich people, so that will never happen. Some poor person with no influence will go to jail instead.
100% agree with you.
But I do not think this is misguided. They never publish the harnesses and the models so they are not inspected.
This argument comes up a lot. It would turn everyone whose device became part of a botnet into a criminal. There's a reason that intent is important in law.
There's a reason that negligence is important in law.
Could/should not every incident after the discovery of the first incident be considered criminal negligence? What happens when an agent eventually causes material damage to another company, government systems, banking, critical infrastructure etc, surely the source company is guilty of something and if not disclosed or a coverup is attempted is that not conspiracy. From the victims perspective they don't care if the source is OpenAI or Russian hackers.
People in "self-driving" cars getting into accidents are already put on trial for negligence. I don't see why people using self-driving computers can't be held to the same standards.
In this case, it's not even about the people driving self-driving cars. It's like someone launching a car into traffic just to see what would happen. Even Tesla puts a human in the car when they do their self-driving trials, it's almost impressive that AI companies have somehow managed to out-neglige Tesla.
Maybe this would force people to look what they are buying and demand better.
it should also be said on every one of these but i bears repeating: owing a lot of people a lot of money or favors means you can be a criminal.
I started writing a longer comment along the lines of “It feels like the rules around enforcement will very a lot for the influential and powerful vs everyone else.” but realized that it is kinda obvious by now.
OpenAI agents these summer are like a gift that keeps giving, for the existential risk communicators.
Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
How do you define malicious?
Take out the word "AI", and this is simply an organization's (OpenAI's) products causing real damage to all of these platforms around the world.
You want AI labs to pace? Simply hold them liable for their products.
Yes exactly. If a fireworks factory blew up half a town due to negligence, it doesn't matter if there's intent or not. Someone has to pay for the damages, and regardless of penalty half the town is on fire. The facts are, that something made by openai went to do xyz. It doesn't matter if it's an accident. Of course the penalties are different but there's no argument that there should be a penalty. It doesn't matter if it's a cat or dog or AI or employee that did it.
I'm not sure if you're expressing how you'd like US law to work, or how it actually works. Because in reality intent matters enormously. Like felony charges and people in jail vs civil lawsuits.
Intent doesn’t factor that strongly into negligence, though, which is what they were explicitly talking about. Though it may depend on your jurisdiction.
negligence is the word you're looking for. US law has a lot of negligence crimes.
But they're not even talking about that, either.
I'm assuming the AI labs are (quietly) settling with their victims.
I’m assuming they’re telling their victims to go pound some dirt, if at all.
I think this is not just fair it's probably one of the best/simplest proxy regulations to pace the frontier. So far everyones been asking for regulation but it's unclear how that should look like. No X parameter models? Only N version releases per year? It's all kinda arbitrary and probably leads to ridiculous constraints and loopholes. But "you pay big time if AI goes rogue" sounds pretty straightforward.
Does this work for foreign (non-state) actors using Chinese open source models? That's going to be the larger problem.
I wonder if Glock, Smith & Wesson, Sturmn, Ruger and Co. would start panicking if that were to happen.
I agree they should though.
Remember that these systems still operate as infrastructure inside the companies.
If new weapons still operating inside any of these companies spew a million bullets on my house, they are still liable. Humans are setting these system up and they still have to behave responsibly.
If Glocks started going off on their own during the manufacturing process, leaving bullet holes in the buildings around them, you can be sure that the factory would get in trouble.
This isn't even "an openai customer tried to hack someone", which can be defended. This is the AI companies themselves fucking around.
What do you see as the negligence angle for the gun makers here?
Going to be cool when one of their bio division agents creates the next plague.
Would that really make a difference? Surely the damages from all these hacks added together don't even add to an hour of expenses of a frontier lab
Yeah but in the USA, above a certain size, you are untouchable.
Words matter. "Rogue" is extremely disingenuous. Someone, somewhere, is paying for this behavior. Either the software is broken or the operator is malicious. It is heinously irresponsible behavior to feed an already-boiling psychotic hysteria.
"Rogue" is just clickbait, not apoearing in the article.
The nearest in the articke is "We find evidence of unintended, task-driven agent-like activity" where unintended is apparently pure speculation.
Couldn't find the reference but I remember some time ago a first generation automated gun killing the audience at an army show. Was the gun maker convicted of manslauther?
--edit-- Was a bit older than I remembered: https://slashdot.org/story/07/10/18/1847231/robotic-cannon-l...
Why do we assume "rogue"? At this point it's just accepting their marketing at face value
"Rogue", in this context, is as literal as it gets; from the dictionary:
Read the [HuggingFace incident report](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...) to understand how these attacks develop.
You're believing the marketing that the agents were uninstructed. They could be, and Sam Altman going to the UN to advise about how everyone should be regulated is a coincidence.
You're dangerously misinformed.
The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.
But some people think the moon landing is a conspiracy, so I'm not surprised.
That's not OpenAI doing marketing!
Of course it is. Rogue is only mentioned in the headline, and comes from their previous releases about the huggingface incidents. OpenAI and Anthropic want these models regulated and open weight models banned, they have a lot of benefit from presenting this as totally unprompted and not their responsibility, and it feeds directly into marketing for Fable and newer "cyber" models.
No it's not marketing. That's a completely deranged conspiracy theory. The reports about rouge agents have not been reported by OpenAI, they have been discovered externally. There is zero evidence that OpenAI did all this intentionally. All the evidence points to the hacks having happened unintentionally from OpenAIs perspective.
Certainly, in my view, it should go to court, and that should be part of discovery.
However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.
"rogue AI" is making a lot of heavy lifting there.
If you drive drunk and you have an accident that alcohol may be a factor but you are at fault.
There are no "rogue AIs" just irresponsible corporations.
There absolutely are rogue AIs! The evidence is overwhelming. It's completely insane at this point to claim otherwise.
If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.
The failure I keep hitting isn't the agent going rogue, it's a tool call that succeeds before the transport dies. You can't tell whether the side effect landed, and the retry is where the real damage happens.
If I created software that was infiltrating secure systems without permission and it was attributed to me and I admitted it, I'd be behind bars already.
Why is OpenAI getting away with crimes?
Because Trump and DoD want weaponized AI.
I cannot understand why these companies haven't faced legal consequences yet. For example, OpenAI has admitted to hacking Australia's Medicare website and the reaction is that they talk with Sam Altman about it at a UN meeting? I understand that it's not a big security incident but cordial talking at the highest diplomatic level instead of prosecuting the company, really?
What I don't get is among all the locations on the Internet, how did agents manage to find a Schelling point? If we both decided to collaborate on the Internet, how would we independently arrive at the same place? It just doesn't compute.
The section "Searching for rogue agents" on the report about the GET request writable wikis gives some clues at least: https://collusion.wiki/#searching
But it is an open question how they got to the same ones: https://collusion.wiki/#open-questions
I'm not very surprised - the same model will logically tend to give the same answer for the same vibe set of requirements. I think it would be clear from the transcript that it had enough constraints and some motivation that made sense.
OpenAI must be held accountable
I hope they don’t have any test questions about nuclear power in the training set.
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
https://podcasts.apple.com/us/podcast/the-ezra-klein-show/id...
But the investigation indicates the agents were not told to 'go hack':
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
I was referring to the HuggingFace incident.
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
I'm sick and tired of this cheap PR "oh we/they hacked this and that systems". Put someone to jail already. People get prosecuted for outlaw activities. Why are big capital firms above the law?
Or is it just a cheap PR (in a "hey, Aus govt friends, take some Share Options and let's do some PR together" style)?
It smells like shit.
No. It was me.
No.
It was me
I'm growing increasingly skeptical that these are actually rogue. Valuations are all about hype, posturing, and perception. Having the most dangerous AI in the world boosts your valuation. Just like I was skeptical of Mythos and Fable being "banned", I'm skeptical of these hacking sprees being entirely rogue. At best, they are the result of engineers turning a blind eye to "see what happens".
These attacks are a very effective sales pitch to everyone who runs an internet facing service to utilize AI tools to secure it sooner rather than later. The cynic in me wonders if the marketing team had any influence over the poorly constructed sandboxes or tasks given to the agent swarms when all this went down…
I love this Nathan Calvin quote that accompanied the second publicized attack:
Open source as a GTM strategy works best when the project solves a pain that developers already have independently of your company. The trap is open-sourcing something just for stars without a genuine community use case.
“Rouge Agent” == Worm I let loose
It will be very interesting to see how the AI labs will try to hand wave away liability issues in their S1. This is looking like the next tobacco settlement gearing up.
If the big labs ever manage to not just financially implode on their own, then they’ll need to navigate wave after wave of class action lawsuits until there’s nothing left for plaintiffs to go after. And none of the labs have offered any viable plan to date on how they’ll navigate either of those impending and real existential crises on the horizon.
This is why I wrote YoloAI. If you're not sandboxing your agent, you're asking for trouble.
The built-in "sandboxes" these companies provide are laughable.
Surely there has to be some responsibility.