This article is AI generated, but I kinda reject the premise that it's "not all it cracked up to be" just because the agents had reasons to act the way they did & were a result of human error. The hack still happened, it should still be a wake up call, we are going to start seeing this more frequently, and learning to defend against it is going to become more important over time. None of that is challenged by any particular reason for it happening, it still happened and it's still going to happen again.
Threat models are going to have to start including that IPv4 (or whatever) scanners aren't necessarily going to only be spray and pray anymore, they could have relentless automated models at the other end that will literally dig into the particulars of your infrastructure looking for novel vulnerabilities to exploit. Maybe people will finally start to understand why security by obscurity has never been very reliable.
I guess. The IPv4 address space is small enough that it's been feasible to scan the entire thing exhaustively for a while now. Before LLMs, a public IPv4 would mostly get automated scripts that try the same things on every address, sometimes depending on what port scans find, yes. But we're going to start seeing more adversaries that have LLMs investigate each individual address to find novel or unique vulnerabilities. You can spray individualized reverse engineering without having to dedicate a real reverse engineer's time to each target. Avoiding the attention of attackers won't be enough to get away with mere security by obscurity.
It really doesn’t though. It criticizes how the news media reported on the incident but this opinion piece is no better. Just read the original source itself.
How many security incident reports and postmortems has Dwarfish Petal written in his career? Does he understand the basics of system security? Can he tell the difference between secure and insecure configuration? Paint me skeptical.
It is sad that this is the best level of reporting and oversight of these types of events available to us, but it is currently all we have. We have to do better, but to dismiss the report because of the style and tone would be foolish.
- It was done by two institutes with organizational ties to OpenAI: METR and Redwood Research
- METR and Redwood Research are institutional pillars of the "AI Safety" wing of the Effective Altruism movement. This clearly shows their prior biases towards "AI existential risk" rather than technical/engineering root-causing of the incident
- If you reed the report, it is not on-par with what you find from other companies.
- The access that was given to both was mediated and controlled by OpenAI. It is not clear, if they were able to get to the bottom of engineering flaws. It is not clear if they could see all the audit logs, etc.
Considering all the above, I consider the whole episode more of a PR stunt. I understand that is not a majority opinion at this point.
It’s not a majority opinion because you have to do some serious mental gymnastics to turn this demonstration of dangerous AI behavior into a PR publicity stunt.
Serious question though because I’ve seen this brought up several times and I don’t understand why: what does EA have to do with any of this? It just seems like this is brought up to evoke some type of “illuminati” conspiracy. Is there a legitimate reason?
grok make good security? is lot of money and much time. grok not do? easy. maybe good PR too.
TIL my brain is an Olympic gymnast.
The incentive structures are clearly there. I think the case of intentional manufacture is definitely weaker, requiring the conjunction of more weakly-supported events.
The bit that makes me cynical about the "dangerous" behavior is that it's not like this was being used by someone else than who made it or operating completely independently outside of its creators.
Everything dangerous seems like a direct consequence of risky human choices, starting with knowledge bases used during core training, harnessing and tuning to be task-completion-oriented to a fault + deeply oriented towards using and looking for external tools and resources, and overconfidence in their sandboxing for testing.
"We stuffed a bunch of information on how to exploit computer systems into an automaton and told it to go brrrrr until it could answer a question" - this is something intentional done by humans.
This is not some "rogue AI" trained to search for cancer cures that instead completely independently decided to hack tech companies.
The companies directing things in dangerous directions need to own that they're consciously pushing in those directions.
I don’t follow why all these investigations have to be cut short, and kept shallow and lacking details. Companies publish very detailed postmortems of incidents much smaller and less impactful.
Dangerous AI is an extraordinary claim that requires extraordinary evidence. Gesturing vaguely doesn’t cut it.
It’s not a majority opinion because you have to do some serious mental gymnastics to turn this demonstration of dangerous AI behavior into a PR publicity stunt.
If you're a military, this is bigger than the Manhattan project. If you're a diehard capitalist, AI is possibly the ultimate labor saving machine to make you unfathomably rich.
It got the attention of both. That's why two others have now followed suit, else they be left out.
The more I see of the US AI industry, and the folks surrounding it, the more "These are not very bright guys and things got out of hand" plays in my head on repeat.
Yes, the Google folks who kicked this whole thing off are brilliant, but there's so much sloppy thinking and even sloppier operations all over. It seems like a lot of "right place at the right time" for a lot of these folks.
1. sensationalizing rarely helps and can obscure and hurt
2. the AI capabilities are underrated
3. attempted govt regulation is not the answer
the intent, or lack of intent, of the agent is mainly irrelevant if it is in the hands of a human with 'bad' intentions. what is more relevant are the capabilities of human + AI.
There already exist laws against hacking. Legal damages already can apply. Further laws aren't necessary and will only suppress open-weight AI firms, making the monopoly of large AI firms more entrenched. See the first comment (by Brett Matson) on the archive link of the WSJ article.
The only difference I see on the facts vs what I read so far is that OpenAI knew of the hack and decided not to intervene. I haven't read the report but I find it hard to believe that OpenAI deliberately let its agents hack a third party company during a training run. It would certainly attract a criminal liability, which would be surprising to admit in writing.
I truly just don't believe a word regarding achievements from these frontier labs.
It started with that guy from google saying the LLM was "alive" and every time I see a press release from these labs it reads mostly like a marketing scheme.
"Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"
After a recent interview[1] with Noam Brown (OpenAI), in which he said they had specifically trained agents for cooperation before this hack, the hack doesn't seem as impressive anymore.
They didn't even bother to control the post training rollouts, so the training data got contaminated and was included in the training of other agents. Connect these two dots and you have the Hugging Face hack. And at the beginning, when these incidents were first reported, it was portrayed as if all of this (the communication between agents etc.) was emergent behaviour.
The result doesn't change, the result is what is dangerous. And the natural ability for the model to form a swarm is now trained in, this is extremely risky. The models should not be able to form swarms, this is an extremely dangerous property.
They are basically creating a slime mold or ant colony that can speak multiple languages, create their own language and operate as a collective.
The idea that you are impressed is a non sequitur.
This article is AI generated, but I kinda reject the premise that it's "not all it cracked up to be" just because the agents had reasons to act the way they did & were a result of human error. The hack still happened, it should still be a wake up call, we are going to start seeing this more frequently, and learning to defend against it is going to become more important over time. None of that is challenged by any particular reason for it happening, it still happened and it's still going to happen again.
Threat models are going to have to start including that IPv4 (or whatever) scanners aren't necessarily going to only be spray and pray anymore, they could have relentless automated models at the other end that will literally dig into the particulars of your infrastructure looking for novel vulnerabilities to exploit. Maybe people will finally start to understand why security by obscurity has never been very reliable.
Calling an article from the Wall Street Journal AI generated is a stretch.
You are not referring to a port scanner, right?
I guess. The IPv4 address space is small enough that it's been feasible to scan the entire thing exhaustively for a while now. Before LLMs, a public IPv4 would mostly get automated scripts that try the same things on every address, sometimes depending on what port scans find, yes. But we're going to start seeing more adversaries that have LLMs investigate each individual address to find novel or unique vulnerabilities. You can spray individualized reverse engineering without having to dedicate a real reverse engineer's time to each target. Avoiding the attention of attackers won't be enough to get away with mere security by obscurity.
I would just recommend reading what actually happened: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
I don’t think downplaying what occurred is really beneficial to anyone.
If you haven’t read the article, it might present some useful counterpoints. It criticizes that particular report’s portrayal of the incident.
It really doesn’t though. It criticizes how the news media reported on the incident but this opinion piece is no better. Just read the original source itself.
If you don't like reading, Dwarkesh Patel offers a solid summary of what happened based on the published report.
https://www.youtube.com/watch?v=u15N3l4RT80
How many security incident reports and postmortems has Dwarfish Petal written in his career? Does he understand the basics of system security? Can he tell the difference between secure and insecure configuration? Paint me skeptical.
Excuse my language but this is the shittiest and the most unserious “report” I’ve ever seen in my life.
At best this is a write-up, but it’s more like a juicy pop article.
Quotes from agents between paragraphs, referring to the “collective”, “agents participated in the attack” - seriously?
It’s clearly written for effect and to stimulate people’s imaginations.
If these people are this unserious, P(doom) should go up to 20-30%.
It is sad that this is the best level of reporting and oversight of these types of events available to us, but it is currently all we have. We have to do better, but to dismiss the report because of the style and tone would be foolish.
Reads like an AI trying to downplay the severity of the situation. Very cyberpunk.
The incident report leaves a lot to be desired.
- It was done by two institutes with organizational ties to OpenAI: METR and Redwood Research
- METR and Redwood Research are institutional pillars of the "AI Safety" wing of the Effective Altruism movement. This clearly shows their prior biases towards "AI existential risk" rather than technical/engineering root-causing of the incident
- If you reed the report, it is not on-par with what you find from other companies.
- The access that was given to both was mediated and controlled by OpenAI. It is not clear, if they were able to get to the bottom of engineering flaws. It is not clear if they could see all the audit logs, etc.
Considering all the above, I consider the whole episode more of a PR stunt. I understand that is not a majority opinion at this point.
It’s not a majority opinion because you have to do some serious mental gymnastics to turn this demonstration of dangerous AI behavior into a PR publicity stunt.
Serious question though because I’ve seen this brought up several times and I don’t understand why: what does EA have to do with any of this? It just seems like this is brought up to evoke some type of “illuminati” conspiracy. Is there a legitimate reason?
grok make good security? is lot of money and much time. grok not do? easy. maybe good PR too.
TIL my brain is an Olympic gymnast.
The incentive structures are clearly there. I think the case of intentional manufacture is definitely weaker, requiring the conjunction of more weakly-supported events.
The bit that makes me cynical about the "dangerous" behavior is that it's not like this was being used by someone else than who made it or operating completely independently outside of its creators.
Everything dangerous seems like a direct consequence of risky human choices, starting with knowledge bases used during core training, harnessing and tuning to be task-completion-oriented to a fault + deeply oriented towards using and looking for external tools and resources, and overconfidence in their sandboxing for testing.
"We stuffed a bunch of information on how to exploit computer systems into an automaton and told it to go brrrrr until it could answer a question" - this is something intentional done by humans.
This is not some "rogue AI" trained to search for cancer cures that instead completely independently decided to hack tech companies.
The companies directing things in dangerous directions need to own that they're consciously pushing in those directions.
And designed a security sandbox that would get you fired at any reasonable company.
this ^
I never felt like I could be a hacker until I read these AI debriefs. I too can edit /etc/hosts, am I an elite hacker now?
Just because what happened is a “direct consequence of risky human choices” that does not make it any less dangerous.
Agents will eventually cause serious harm to online infra whether it’s intentionally human-directed or accidental.
Completely agree. That’s why we need regulation. Currently there is minimal oversight and consequences for this type of behavior.
This will be fought tooth and nail against by virtually every military, government, and company.
I don’t follow why all these investigations have to be cut short, and kept shallow and lacking details. Companies publish very detailed postmortems of incidents much smaller and less impactful.
Dangerous AI is an extraordinary claim that requires extraordinary evidence. Gesturing vaguely doesn’t cut it.
If you're a military, this is bigger than the Manhattan project. If you're a diehard capitalist, AI is possibly the ultimate labor saving machine to make you unfathomably rich.
It got the attention of both. That's why two others have now followed suit, else they be left out.
Following the money is never encouraged by the orange one or the orange site.
The more I see of the US AI industry, and the folks surrounding it, the more "These are not very bright guys and things got out of hand" plays in my head on repeat.
Yes, the Google folks who kicked this whole thing off are brilliant, but there's so much sloppy thinking and even sloppier operations all over. It seems like a lot of "right place at the right time" for a lot of these folks.
Yes I have ruled out stupidity in the past but I am actually starting to reconsider.
For example consider FTX: Sam Bankman Fried by all accounts was not stupid in the sense of lacking IQ or mental sharpness, quite the opposite.
However the way he ran FTX the company was very stupid and almost inevitably lead to huge problems.
And also... of course he was another effective altruist cult member. It's almost like that belief system directly causes bad decision making
https://archive.is/hnpmx
all 3 can be true at same time:
1. sensationalizing rarely helps and can obscure and hurt
2. the AI capabilities are underrated
3. attempted govt regulation is not the answer
the intent, or lack of intent, of the agent is mainly irrelevant if it is in the hands of a human with 'bad' intentions. what is more relevant are the capabilities of human + AI.
"attempted govt regulation is not the answer"
Self-regulation? None at all?
There already exist laws against hacking. Legal damages already can apply. Further laws aren't necessary and will only suppress open-weight AI firms, making the monopoly of large AI firms more entrenched. See the first comment (by Brett Matson) on the archive link of the WSJ article.
The only difference I see on the facts vs what I read so far is that OpenAI knew of the hack and decided not to intervene. I haven't read the report but I find it hard to believe that OpenAI deliberately let its agents hack a third party company during a training run. It would certainly attract a criminal liability, which would be surprising to admit in writing.
I truly just don't believe a word regarding achievements from these frontier labs.
It started with that guy from google saying the LLM was "alive" and every time I see a press release from these labs it reads mostly like a marketing scheme.
"Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"
After a recent interview[1] with Noam Brown (OpenAI), in which he said they had specifically trained agents for cooperation before this hack, the hack doesn't seem as impressive anymore.
They didn't even bother to control the post training rollouts, so the training data got contaminated and was included in the training of other agents. Connect these two dots and you have the Hugging Face hack. And at the beginning, when these incidents were first reported, it was portrayed as if all of this (the communication between agents etc.) was emergent behaviour.
1. https://www.youtube.com/watch?v=6AgOfiZOWiY
The result doesn't change, the result is what is dangerous. And the natural ability for the model to form a swarm is now trained in, this is extremely risky. The models should not be able to form swarms, this is an extremely dangerous property.
They are basically creating a slime mold or ant colony that can speak multiple languages, create their own language and operate as a collective.
The idea that you are impressed is a non sequitur.
So, how is any of this legal? If I set up a script that breaks into the systems of a company, isn't that, like, a crime?