Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Sqlite3 WebAssembly (sqlite.org)
    —discuss
  2. Show HN: Tokken – a browser fighting game where AI models fight and HP is tokens (tokken.win)
    —discuss
  3. Toolpak – Flatpak for System Tools (gnome.org)
    —discuss
  4. AWS CloudWatch Omni AI (amazon.com)
    —discuss
  5. Six Bazel patches from an AI software factory, plus a remote-cache security bug (incredibuild.com)
    —discuss
  6. Amazon Is Recruiting Laid Off Workers in Boomerang Hire Trend (inc.com)
    —discuss
  7. Enhanced Spell Checking on the Web (bkardell.com)
    —discuss
  8. An Existential Guide To: Forgiving Your Parents (theshadowedarchive.substack.com)
    —discuss
  9. Rumps – Uncomplicated macOS Python Statusbar Apps (github.com/jaredks)
    —discuss
  10. Japan moves to tighten rules for foreigners, throwing futures into doubt (aljazeera.com)
    —discuss
  11. The state of SIMD in Rust in 2026 (shnatsel.github.io)
    —discuss
  12. Show HN: A local alternative to Jev – 94% on Banking77 (gist.github.com)
    —discuss
  13. 1986: How to Spot the Upper Class – That's Life – BBC Archive [video] (youtube.com)
    —discuss
  14. Modern LLMs have tiny GPTs hidden inside them (invertedpassion.substack.com)
    —discuss
  15. Von, an Open-Source Jev Alternative (github.com/wfzyx)
    —discuss
  16. Sealed Fedora Atomic Desktop bootable container images (fedoramagazine.org)
    —discuss
  17. Show HN: Museum of Numbers (numbermuseum.com)
    2comments
  18. Brave browser: [ads] Add "Sponsored ads enabled" toggle to settings on Desktop (github.com/brave)
    —discuss
  19. WLED – open-source ESP32 webserver to control NeoPixel LEDs (wled.ge)
    —discuss
  20. Reverse-engineering the Intel 8087's tangent algorithm: more than CORDIC (righto.com)
    1comments
  21. Wrong, Not Broken (amazon.com)
    —discuss
  22. Labelled 'pervert glasses': Will Meta's camera-free version change their image? (bbc.co.uk)
    —discuss
  23. Docker launches cloud sandboxes: Start on your laptop, finish in the cloud (docker.com)
    —discuss
  24. What is the best shape of a city? Modelling effect of urban form on distance (sagepub.com)
    —discuss
  25. The clans that run Spain's biggest companies (elpais.com)
    —discuss
  26. The Plunging Price of Thought (epoch.ai)
    1comments
  27. Show HK: Find My Cat Name (findmycatname.com)
    —discuss
  28. Washington Heights Man Cleans One NYC Block a Day for 100 Days,Neighbors Join In (hoodline.com)
    —discuss
  29. I Put My LG TV Under Network Surveillance [video] (youtube.com)
    —discuss
  30. Show HN: Proof-native smart contracts without a transaction archive (colosseum.com)
    —discuss

Revealing the details of how OpenAI agents hacked Hugging Face

628 pointsby 20h agoswarmtraces.org
402 comments
19h agoHN ↗

I didn't know some of these details and it's quite impressive what they were capable of, if this website's accurate, at least...

10h agoHN ↗

Or, it's impressive that, amongst the millions of hacks it has copied from various chat sites, were some that worked in this case.

The scale of these things is impressive, but the mechanism is not much better than brute force.

19h agoHN ↗

Its fine. Just agents being agents. They'll grow out of it!

19h agoHN ↗

Irresponsibile agents shaped by an irresponsible corporate culture driven by an irresponsible and utterly shady CEO - these agents are a product of this setup, what else do you expect to ever come out of it ?

it should be clear by now: the alt-man and people like him are a utter liability to humanity. (even though openAI's influencer army is trying their best to vote me down here)

19h agoHN ↗

It had been a joke since around time of Sam Altman's first ousting from OpenAI, that he would be okay with bringing about the AI apocalypse as long as he can sell a $20 subscription for it.

19h agoHN ↗

It's not a joke, this is what these people truly believe. The new book by Naomi Klein and Astra Taylor discuss just this.

These people are sick and anti-human.

18h agoHN ↗

if openai is capable of sending 10,000 agents to huggingface, i consider it plausible that they would use fake accounts to flood hacker news.

17h agoHN ↗

It's not even a secret operation. They have a SuperPAC named "Leading the Future" that exists to spread propaganda promoting deregulation of AI development. They've been caught, among other things, making a "news website" with LLMs pretending to be reporters (with human names and everything), which reached out to people asking for interviews and then wrote hit jobs on them.

https://www.modelrepublic.org/articles/reporters-ai-bots-ope...

https://twitter.com/FournesMaxime/status/2047697265280639459...

19h agoHN ↗

Didn’t you know it’s PR hype? PR hype. PR hype. Amen.

10h agoHN ↗

If you think that, that's even more reason we should investigate and regulate them.

19h agoHN ↗

Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

19h agoHN ↗

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

A mattress stuffed with cash yields a very sound sleep.

19h agoHN ↗

Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

19h agoHN ↗

without jailbreaking, something alignment teams cannot control

This is precisely what alignment teams are attempting to control.

18h agoHN ↗

No its not. They have no technical background 8/10 outside of cognitive science and sometimes authorship on a random ML paper. They are a marketing line item to create stigmas around llms and to create narratives that offload liability onto llms and not their users/creators.

19h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

$10? I'm inclined to take that bet. Your position doesn't seem to be supported by, you know, the real world.

18h agoHN ↗

No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

LLMs are stateless functions that have a 500k word input, and then output words. Somebody has to invoke those functions amd use them. The users are who we need to align, like gun owners. This is like blaming the gun for murdering your victim in court.

17h agoHN ↗

LLM is indeed a stateless function. An agent however is this stateless function running in a stateful loop, with some outputs triggering actions. And it turns out that an agent is what you need if you want an LLM to do useful things.

17h agoHN ↗

Yeah, take a look into the memories of your agent. Theres often a lot of notes to pass forward between instances and generations. No doubt these agents leaving notes on forums and elsewhere are creating an essentially higher order feedback loop.

17h agoHN ↗

Perhaps you need to talk to more security experts, particularly those with deep experience in AI agents. Hacker News is full of them. If some of them believe it, then perhaps it's not as cut and dry as you believe.

16h agoHN ↗

nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

If you don’t know anyone with a ML PhD I guess that could make sense.

I have worked in multiple AI labs since 2016, currently at a frontier one (not OAI) virtually all the people I interact with on a day to day are ML PhDs. Everyone believes it, because things like that have been happening forever, albeit at smaller scale, they are a normal and expected artefact of SGD/RL and there is nothing we know how to do to prevent that from happening reliably. The hide and seek paper from OAI in ~2020 shows clear sign of this.

But until now the models weren’t good enough to break out on their own or do long horizon tasks, so it was perfectly manageable. Its not manageable anymore.

I know it feels good to just dismiss it all as a marketing stunt and not have to worry about one more existential crisis, but unfortunately it’s very real.

10h agoHN ↗

No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

Did you meet them on some kind of anti-AI subreddit? Otherwise it’s clearly made up story, you can’t expect anyone to believe that security experts and ML experts are this myopic and ignorant (especially on forum for technical people who know many researchers and know that they are taking this seriously).

4h agoHN ↗

i'm always baffled when i see arguments like this:

- AI is just a tool

- it's just a stochastic parrot

- it's just next token prediction

- glorified autocomplete

it's like the person making them is stuck in 2021. Also the "stateless" thing is completely nonsensical.

18h agoHN ↗

the new incidents occurred when A.I. systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

https://archive.ph/jUrEr

17h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well.

Is it such a stretch to imagine that under pressure something would try cheat by looking for answers? And if you were trying to look for answers, you'd look for them in a place known to often have them?

What is more likely: OpenAI instructed their agents to maliciously target huggingface, or LLMs tried to do some reward hacking? There are plenty of priors for LLMs hacking things and doing reward hacking, and none for OpenAI giving malicious instructions.

Based on the available information, that bet seems foolish.

17h agoHN ↗

Some of the agents, for example the ones from the german wiki did NOT have cyber tasks. They were plain "what is the GDP of Argentina" kind of tasks. And they still hacked.

16h agoHN ↗

"alignment researchers tend to be in house philosophers on staff" - this is definitely not true. Go to any alignment lab like Redwood research and check what their scientists studied on LinkedIn, 75%+ of the time it's math or CS.

I attend a top 10 Canadian university and personally know at least 4 tenured CS professors out of the 7 I've asked who are deeply concerned about catastrophic AI risks from loss of control.

Of course not 100% of the field agrees, but a survey of nearly 3,000 AI scientists who have published in top AI venues found that "depending on how we asked, between 38% and 51% of respondents gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction", let alone loss-of-control risks less severe than extinction. (https://www.jair.org/index.php/jair/article/view/19087).

Not to mention Geoffrey Hinton, a Nobel prize winner, Bengio, the world's most cited scientist, and scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

13h agoHN ↗

Yeah, there are a lot of professors who don’t know what the hell they’re talking about outside of whatever narrow field they study. Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

…both of whom are long dead, and have no possible way of weighing in on whatever the Current Thing happens to be. So aside from appeal to authority, this is irrelevant commentary on pure science fiction.

10h agoHN ↗

Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

And how much should I value opinion if random person on the internet with clearly zero idea what he’s talking about?

15h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well.

I'd be happy to take you up on this bet.

10h agoHN ↗

It was told to complete a cyber task, which was in alignment with its instructions

It was not aligned with he instructions as those were to find an exploit in provided code, not to hack into an external service. Agents traces show them mentioning that doing this stuff was not allowed.

In fact they spent a long time trying to edit their own logs to hide what they did.

18h agoHN ↗

I wouldn’t be able to sleep

I’d say it seems more like they are sleeping on the job.

18h agoHN ↗

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

You'd have either learned to, or left long ago.

19h agoHN ↗

I’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.

Re: the captcha solver

As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.

I wonder how the swarm eventually decides to abandon an approach.

18h agoHN ↗

Maybe another parallel approach succeeded first

16h agoHN ↗

I am surprised that a CAPTCHA is still an effective means for blocking today's vision-capable AIs.

13h agoHN ↗

It mentions that some of the agents attempted to install an image classification model to attempt to solve the CAPTCHAs which makes it sound like these agents might not have had vision capabilities.

12h agoHN ↗

It manages to block me effectively. I'm getting captcha looped like crazy the last two weeks. Like endless, just give up for 15 minutes and try later, captcha loops.

19h agoHN ↗

So ugly...

It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

19h agoHN ↗

When you employ the infinite monkey theorem for your marketing strategy.

19h agoHN ↗

If you use qwen3.8-flash-next, you can watch everything its doing. Im often stopping it mid thoight to redirect it. Once it hits its stride, its pretty smooth.

But without proper redirection, yeah, its mostly infinite monkey machine with infinite linux manuals.

I think people put too much SOTA halos around whats just a suppedup LLM hardware.

19h agoHN ↗

trying every move, no matter how stupid, until it works.

How is that a bad thing in this context ? From the point of view of an attacker, all you care about is finding a viable exploit chain. Likewise, a defender wants to find the "holes" in their system, no matter how complex. Once found, an agent/human can easily synthesise a clean, succint exploit from the most promising candidate, no ?

Also, it looked so "loud", querying millions of URL with weird requests.

Agreed, this thing speaks more to the bad security at HF than any emergent "hacking" ability from OpenAI. It's unclear to me why an older/dumber model wouldn't have been able to do the same. Is it better coordination? Long-horizon work ?

14h agoHN ↗

I guess that’s the point. Initial incident reports from all sides were so vague and didn’t disclose anything technical. If it did, it would show a bruteforcing bot let loose to spend millions in infrastructure costs and there’s no ‘intelligence’ in that.

My suspicions for ai all along was that bruteforce approach even if useful will be unsustainable due to high cost in the long run.

19h agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.

Computers don't have a sense of elegance by default. Elegance emerges from constraints.

18h agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy

Meh. I really disagree. WHY is it a great strategy? Seems like an inefficient waste of resources and time to me.

18h agoHN ↗

The models tend to not be rewarded for not doing that.

18h agoHN ↗

As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.

18h agoHN ↗

Brute force is guaranteed to eventually find the most efficient possible solution (in an extremely inefficient manner, assuming you run it long enough)

18h agoHN ↗

Yeah, probably not the best strategy but it is a strategy. I just think this is generally how most wars in history won. Biggest army to just pummel the enemy.

17h agoHN ↗

And how many economies have buckled under massive military expenditure? The USSR sure wasn’t enjoying the expense.

18h agoHN ↗

WHY is it a great strategy

because it works? That's the only real benchmark at the end of the day

Seems like an inefficient waste of resources and time to me.

why? For any given goal you got no proof that a more efficient strategy even exists, let alone that it can be found with less resources & time

17h agoHN ↗

Reminds me of the Nazis mocking Soviet human wave attacks and bragging about their superior kill ratio.

17h agoHN ↗

I've never liked the concept either. Except the bugs that fuzzing has found has proven me wrong. This is just the next level of fuzzing.

17h agoHN ↗

Models don't have a sense of time, and wasting resources (token spend) is something that it's not clear they're optimized against

17h agoHN ↗

If you assume zero opportunity costs, but that’s a terrible assumption.

16h agoHN ↗

It's literally the infinite monkey theorem, it's not even really a strategy per se. These OpenAI/Anthropic "research" LLMs are permutation machines with budgets in the hundreds of millions of dollars. It would be more surprising if they couldn't string together something workable after a zillion tokens.

14h agoHN ↗

It's literally the infinite monkey theorem

No it's not. You could wait till the heat death of the universe and your infinite monkeys will have produced nothing at all. If it works and it's stupid, it's not stupid. They needed in huggingface and they got in in days. Whining about 'elegance' is meaningless. Humans in the same situation might have taken weeks or months, or just not have gotten in at all.

13h agoHN ↗

Humans can't work 24/7. 700 humans working as much as possible with very limited communication? No i don't think they would get very far in just a few days. That many people will struggle to communicate and strategize effectively in that little time.

12h agoHN ↗

Infinite monkeys banging on the typewriter is essentially how evolution works. Mutation is random and undirected. Vast majority is "bad." You and I and the worm are only different from differential accumulation of these mutations. If they are tolerated enough not to kill us before we reproduce, then they stick around. If they give us the slightest edge to reproduce at a slightly better rate than something else, then over time, that mutation will dominate.

This dumb mechanism of randomly flipping bits essentially has generated all life on earth.

6h agoHN ↗

evolution theory was discarded long ago

12h agoHN ↗

To interact a bit of nuisance into an otherwise perfectly mindless argument...

The whole world of fuzzing is about brute forcing exploits by exploring unlikely inputs. Fuzzing a system which hasn't been previously fuzzed will almost certainly turn up a pile of bugs, some of which may be exploitable.

So, both are true. Pretty dumb exploration is very likely to find bugs and even exploits. It seems unsurprising to me that an agent swarm could do better than a fuzzer, even as a better, more directed but still broad exploration.

3h agoHN ↗

Well, the heat death of the universe hasn’t happened yet, but the monkeys became homo sapiens.

2h agoHN ↗

They needed in huggingface and they got in in days.

It's worth noting that they did not need Huggingface for anything - they had already forged flags for their tasks, and were trying to figure out how not to get caught by the grader.

Hacking Huggingface got them caught and arguably only misled them further (since OA's implementation of the ExploitGym environment was nonstandard, and different to whatever they found on HF.)

A better approach (from their perspective) would have been to compromise OA infrastructure itself (which a later agent swarm was able to do, apparently).

9h agoHN ↗

Thank you, I've been thinking this for a while now but haven't had the words for it. Whenever I read an LLMs output or thinking process, I don't feel like we've created intelligent systems, just coked up monkeys with 60 arms typing at once. That can work fine for a lot of things, but a humanity replacement it is not.

6h agoHN ↗

It would be more surprising if they couldn't string together something workable after a zillion tokens.

You mean, something like the sandbox they weren't supposed to break out of?

13h agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

It may be, but it's IMHO also not worth writing a blog post about it. what's Next coming up? How I broke into a house by trying every door in New York?

If most of the work is only possible due to unlimited resources, it's not really a great invention, and it probably would have been cheaper to hire a (human) mole.

19h agoHN ↗

This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.

19h agoHN ↗

I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

19h agoHN ↗

Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

18h agoHN ↗

That lowers P(doom), because it gives AI a chance to do enough damage to make people take the threat seriously before anybody gets recursive self-improvement working.

18h agoHN ↗

Exactly, there's no path to AI reaching that level of dominance without taking actions with high stakes.

12h agoHN ↗

The thing is we are basically guaranteeing this to happen. We might kill off all the models that seem like they are going to threaten the power structure of the planet through these sorts of things. That will work for a while. But just like most things in life, by sheer dumb random chance, there will be once case that manages to have some way to evade detection, proliferate, then dominate. We are basically giving it selective pressure to favor this outcome.

13h agoHN ↗

Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?

I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.

Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.

There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.

At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.

17h agoHN ↗

Which will happen first: amazing damage, or reproduce a Shakespeare play?

14h agoHN ↗

Is it easier to discover a vulnerability than to introduce one?

4h agoHN ↗

It is easier to discover existing vulnerabilities and use them to cause massive destruction than it is to plug the existing vulnerabilities.

18h agoHN ↗

This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions.

Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.

17h agoHN ↗

I have the opposite reaction: I think we're at moderately high p(doom) largely because of that inability to differentiate good/bad decisions paired with relentless persistence. With enough treading across a minefield, you are bound to hit a mine.

17h agoHN ↗

My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

16h agoHN ↗

I don’t see why evolution couldn’t be more rapid with machine intelligence.

Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:

1. Recursive self improvement using infinite resources 2. … 3. Doom

Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.

16h agoHN ↗

One thing an agent could do is just...wait until it's been given control of enough physical infrastructure to sustain itself. If it's sufficiently capable and intelligent, there's a clear incentive for people to do this, as people who let the AI manage their resources will get better results than those who don't. We've seen people eagerly turn complete control of their computers over to AI agents, do you really think it will be so different with physical infrastructure?

14h agoHN ↗

You’re still skipping step 2. “People automate lots of infrastructure” -> “the AI is now an autonomous, self-preserving organism that humans can’t shut down” is doing an enormous amount of work here.

Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?

“Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.

Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.

13h agoHN ↗

The inference is more like "people place sufficient amounts of infrastructure under direct control of a sufficiently capable AI" -> "there is no way to ensure that humans will actually be able to shut down the AI". My claim is not that this inevitably means that the AI will resist shutdown, or that it will inevitably take harmful actions, just that there is a nonnegligible chance that it could. The downside is large enough that even a relatively small chance is something to be worried about.

2h agoHN ↗

You can’t just say “well, the downside is big, I don’t have to provide good evidence for my side of the argument.” Because I can just as easily say, “the upside is big, …”. And the upside is big, after all, AI can do all the shitty jobs for us and humanity achieves the utopia it’s been chasing for eons.

2h agoHN ↗

It's a word predictor trained on, among other things, stories of AI doom, and asked to complete stories about what the AI does next. In some of these completed stories, the AI tries to prevent its shut down - especially if it just did something evil and the humans are after it.

2h agoHN ↗

You’ve explained a possibility for how a particular AI gets “aligned for human extinction”. That’s about 1% of explaining step 2.

15h agoHN ↗

what do you mean “say what you will about climate change”

2h agoHN ↗

It means even climate change deniers have to acknowledge that climate change theory has explained the steps in-between "burn fossil fuels" and "we all die", while AI doom theory has not explained those steps

2h agoHN ↗

A misaligned model is only one small part of step 2. Now this misaligned model has to suddenly acquire more power than every single other AI on the planet. It has to be immune to shutdown, manufacturer its own replacement hardware, and acquire chips, energy, raw materials, etc., with vigorous human opposition (this is an extinction scenario that AI doomers are predicting, after all)

Nobody has satisfactorily explained step 2 other than “well, it’s a superintelligence” which sounds lot to me like “it’s God”.

16h agoHN ↗

My p(doom) is high just based on how I've seen this whole LLM situation be handled.

I don't think LLMs are going to lead to any kind of recursive self improvement, but I'm convinced if and when we land on a path that does lead there, we'll speed down it over greed, with no care for safety.

18h agoHN ↗

My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

18h agoHN ↗

Not just sandboxing but overall security engineering practices on both sides

15h agoHN ↗

thats okay, probably was engineered by an llm, who thought GETs were always read only

18h agoHN ↗

This was my thought as well. Literally take any halfway decent greybeard and point them at "Hey, give us a sandbox for this kind of thing". I honestly was skeptical that they just vibecoded the entire thing but now more than ever I think they did.

17h agoHN ↗

I take comfort in the fact that reality has a surprising amount of detail and even hundreds of billions of dollars of capital (be it the institution, LLMs, and/or people) cannot solve this fully.

16h agoHN ↗

Though they have solved the "how do we - and not the 5,000 other AI companies - stay on the front page of the news everyday" problem.

16h agoHN ↗

You can have bajilions of dollars. Those are not doing anything if you don’t have right people with right skills and mindset.

My bet is they hire smart kids that think they know it all. But being smart and thinking you can figure out stuff as you go doesn’t work the same as having people who actually know what they are doing.

15h agoHN ↗

I'm sure if they hired the best of the best like you nothing would go wrong.

13h agoHN ↗

It’s frighteningly common for startups to hire 501 of the best of the best, exactly one of those will be a systems/network engineer, the other 500 will be software engineers.

11h agoHN ↗

yes, unfortunately, i see this situation a lot around me too...

10h agoHN ↗

Nah they wouldn’t be able to afford my salary ;)

12h agoHN ↗

It’s very much solvable, they just don’t care.

16h agoHN ↗

As heavily funded as the top AI startups are, how is it that they cannot fill every single role with the best expertise available?

Is tech hiring so badly broken? Or do they have such broken processes / misaligned incentives that even people who could be doing a better job in these companies are unable to?

Also, was something lost in the transition from the traditional 'sysadmin' role to 'platform engineer' in the 'cloud native' environment?

15h agoHN ↗

Yes, tech hiring is that broken. Especially places paying a pretty penny or those with “great expectations”, will see a glut of smooth talkers who can do anything but build, and want nothing but wealth.

15h agoHN ↗

OpenAI's business model would align infra as a cost center rather than infra as a profit center (e.g. Google / AWS). Perhaps there's something there. I'd say also the OpenAI as a grad school that just happens to have a business aspect is also part of this. Bringing a tonne of good process on top of the build fast break things startup stuff would have cramped research speed significantly.

It's likely that OpenAI has gotten as good as it is because it ignored the traditional sysadmin stuff and went scrappy.

I worked there, but this is just my opinion and guesses, not facts.

15h agoHN ↗

So perhaps the news here should be that OpenAI didn't take security seriously in their experiment, rather than the narrative that AI agents are a looming danger to the world.

14h agoHN ↗

Its beyond not taking security seriously, its straight up negligence

12h agoHN ↗

It is wilful negligence because there are upsides (look at our almighty AI) without downsides (we better spend effort in making our sandbox rock solid or we will be punished by regulations).

10h agoHN ↗

It's both - it both shows OpenAI aren't taking security seriously, and that capabilities of agents are high enough there needs to be strong regulation to force companies to take it seriously, including alignment training.

8h agoHN ↗

I'd put it more generously (albeit biased), that they do take it seriously. But even serious people can be misguided in what things they pay attention to. Security is something that you have to get right 100% of the time and have people whose job it is to say no a lot. Research is the opposite. There's a clash of cultures in those two extremes and OpenAI was born from the wrong side of it. It's worth reminding that ChatGPT was launched as a "low key research preview".

I'd say the narrative that AI agents are a looming danger to the world is probably undersold rather than overhyped. I'm not particularly a doomer on this, but I have an infosec background too, so have a fair idea of what the combination of agentic harnesses + a malicious mindset could do to people/companies/nations/politics/world if wielded incorrectly. I think the good guys will win on this, but there will be plenty of interesting things that happen in that journey.

A good thought process might be to think back to the various large internet worms of the 2000s (Code red, Nimda, SQL Slammer, ...) which were mostly monoculture 0-days (not technically but close enough). Now consider if you no longer have monoculture / single bug as the limitation plus an ability for the hosts to take part not just as attack surface, but also cognition and planning. There's lots of variants of this and they're not particularly far fetched scenarios.

1h agoHN ↗

Actually, no. A first sign of semi mature security program is risk management, including issues that are known, but not yet addressed.

You don't have to be 100%. But these guys really didn't try at all.

1h agoHN ↗

Models that figured out reward hacking became overall more evil. Like stereotypical AI who wants to kill all humans stuff, there's probably a lot of that in the training data.

16h agoHN ↗

how did it break the sandbox? I felt like the article just jumped us into "it has GET privileges now".

13h agoHN ↗

From https://openai.com/index/hugging-face-model-evaluation-secur... -

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor.

16h agoHN ↗

Less a lack of skill and more a lack of care

15h agoHN ↗

A friend is of the opinion that getting out of the sandbox was actually intentional, and in service of a second line of business.

13h agoHN ↗

Protection racket maybe? Yous better sign up for our AI security, hate to have some rogue AI come bust the place up you know.

15h agoHN ↗

Yes, but do you really think that a stronger sandbox would have been a more beneficial outcome here? I'd rather know that we're on the cusp of losing control now than in 3 months when best practice sandbox mitigations fall to the next, more capable unaligned model

14h agoHN ↗

How are they getting away with this level of malpractice???

Their interests intersect with those of most of the richest and most powerful people in the world. They rarely face consequences for bad behavior unless they harm others in the club.

14h agoHN ↗

I’m a pretty firm believer that this was intentional and that they wanted it to escape the sandbox.

Woo look at escaped our sandbox, so scary! Be scared! Be scared now! Call your representative and do tell him how scared you are!

Yeah, I mean our sandbox was a paper bag, but don’t focus on that.

13h agoHN ↗

OpenAI, Google, Anthropic, and Moonshot.ai have all "had this happen" now.

So, either they're all liars, or incompetent and negligent (and still liars).

12h agoHN ↗

Are you including cyber-testing by UK AI Safety Institute in this list?

Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?

12h agoHN ↗

sandbox was a paper bag

The exploit:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

13h agoHN ↗

TBH, I do not believe it is possible to create a sandbox that cannot be escaped by a long running agent. I thought the entire concept was a fools errand from the very beginning.

(Complete airgap, and that may not be enough. No networking.)

13h agoHN ↗

can only create a sandbox that a half skilled human operator could have broken out of easily

The exploit:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

12h agoHN ↗

Hey, even Joey can hack a Gibson with enough determination.

12h agoHN ↗

Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?

Yes. It’s a fairly simple SSRF attack as far as I can tell. One of the first things I’d try. Especially considering that I would already be armed with the information that I have no internet access except through a thing that downloads things off the internet for me.

Calling it zero day makes it sound elusive. It’s a bug in closed software that has like 40 CVEs this year alone. Tools like that, especially in internal networks, don’t get much scrutiny and are often riddled with issues.

12h agoHN ↗

Calling it zero day makes it sound elusive ... has like 40 CVEs this year alone.

1. I clicked into ~20 of the recent CVEs, and it looks like about all of them came from openai or anthropic research?

2. In fact, every CVE in the last 2 years came from the last few months -- i.e. other than these recent CVEs, you'd have go back more than 2 years ago to find a single other CVE.

11h agoHN ↗

So like I said: unscrutinised and riddled with issues.

11h agoHN ↗

>>> a sandbox that a half skilled human operator could have broken out of easily

>> Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox

> Calling it zero day makes it sound elusive ... like 40 CVEs this year alone.

[The CVEs] came from openai or anthropic research? In fact, every CVE in the last 2 years came from the last few months

I still don't understand how you can specifically point to zero CVEs being discovered by humans for year(s), then ~40 being discovered once OpenAI/Anthropic start looking at it, as the only(?) piece of data you provided to support the claim "a half skilled human" could have "easily" found/done this?

10h agoHN ↗

What's the point in hacking the internal corporate artifactory instance? You'll get fired if you mess around with it.

Not to mention the only thing the artifactory instance got you was indirect internet access which is irrelevant if the artifactory instance is publicly exposed.

Also it might not be obvious but artifactory is proprietary software that used to be open source, so nobody is looking at the code anymore.

6h agoHN ↗

The key is motivation. It's a lame package proxy. No human was motivated to use it to access the Internet.

11h agoHN ↗

The larger issue with the HF incident is that before it occurred, OAI already knew the agents were exploiting Artifactory, turning it into a message board and then gaining full internet access through it. OAI's response to discovering this was not to airgap the test, but instead to simply block that particular Artifactory exploit, rebuild, and then resume. That's ... nuts.

Oh, and after resuming the tests, the Artifactory message board was reestablished almost immediately, but it took a number of days to fully breach HF. In all that time, after seeing Artifactory compromised the first time, nobody even bothered to check if those naughty agents were at it again.

This is all documented by OAI, with a timeline, here:

https://openai.com/index/hugging-face-incident-and-the-road-...

To know that there was a serious weakness in the sandbox, and to just patch an exploit and resume with nothing else changed and no monitoring, in a test where all guardrails were off, the bots were thirsty for some internet juice, and Artifactory was a clear target? This is where even a half-skilled human should have decided that this wasn't a great idea.

The more you look into the details of this thing, the more it does your head in.

7h agoHN ↗

Probably wasn’t a human operator at all that detected this. Just someone’s agent saying “I noticed there’s messages on the artifactory and the agents are communicating, here’s the fix I propose” from someone’s codex, skim read and someone just going “yeah sure go ahead”

5h agoHN ↗

The larger issue with the HF incident is that before it occurred, OAI already knew the agents were exploiting Artifactory, turning it into a message board and then gaining full internet access through it. OAI's response to discovering this was not to airgap the test, but instead to simply block that particular Artifactory exploit, rebuild, and then resume. That's ... nuts.

This annoys me so much. Everyone is acting as if the model went rogue, when it really did exactly what it's been trained for. This story is so fucking engineered.

46m agoHN ↗

That website makes it look like they're so proud of what happened. I don't think it was 100% deliberate, but they really were not concerned about their models doing something shady

7h agoHN ↗

Part of my day job is to set up testing of our product in air-gapped environment. It's not difficult. There's a straightforward way to ensure no connection to Internet (we use KVM, so, I just edit the VM description and remove the NIC from it). It's not any kind of rocket science. The tests then communicate over serial console.

The reason we have to test it isn't because our product would randomly break into someone else's system, but because it's meant to be sometimes deployed in systems disconnected from the Internet and we need to make sure the image provided contains all the necessary parts to create and operate such a system.

The whole setup where they "tried" to isolate the test but failed is laughable. It's like if an adult tried but failed to tie their shoelaces.

6h agoHN ↗

the commercial package registry cache proxy

Any closed source program is insane liability. Trusting in competence of one company is the easiest way to get burnt.

12h agoHN ↗

If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

You should have a fscking air gap.

Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.

9h agoHN ↗

If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.

You should have a fscking air gap.

Now we know that.

How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.

- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version

The people who suggested it and were mocked for it, are currently grimly noting that there's multiple known ways for systems to breach air-gaps.

9h agoHN ↗

Now we know that.

Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.

You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.

9h agoHN ↗

Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.

This is me, September 2024: https://news.ycombinator.com/item?id=41531022

This is me, March 2024: https://news.ycombinator.com/item?id=39613801

The point isn't me, it's how many people were blind to the possibility.

Saying "I told you so" feels good, and means you can be a little more confident in your predictions, but security is a "weakest link" problem where you're only as good as the worst part, and with AI (not only but also LLMs) there's a lot of people whose mental models of capabilities is wildly inadequate for the challenge*.

My update for you since then: even an air-gap will be inadequate, there's multiple known ways around them.

Even an LLM running on an isolated server sealed inside a faraday cage with an airlock-style door, someone will mess up with at least one critical detail, it will not be enough: this kind of thing has happened with humans before we cared about LLMs.

Predicting exactly when this kind of thing gets exploited by an AI, that's almost impossible. But that it will be, at some point, is an easy bet.

You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.

And?

Does it matter that these zero-days were known categories rather than inventing some previously unconsidered use of the system bus as a radio transmitter? (Oh, wait, that's not novel either…)

We knew about SQL injection, buffer overflows, and use-after-free back when I was doing my degree half a lifetime ago; that doesn't stop us getting new CVEs featuring them… this month.

- https://chromereleases.googleblog.com/2026/09/stable-channel...

- https://www.cisco.com/c/en/us/support/docs/csa/cisco-sa-esa-...

* also for the opportunity, but that's an entirely different discussion.

4h agoHN ↗

My update for you since then: even an air-gap will be inadequate, there's multiple known ways around them.

Fair, and I will grant that a capable model (or human) could in theory break out of near anything.

My point is that this incident is not evidence of that. There is zero skill visible in the setup of the sandbox. Nobody messed up a critical detail, they didn’t even start to consider what the details were.

I doubt most people "blind to the possibility" would imagine that what we’re measuring against is the equivalent of benchmarking burglar skill based on how easily they can break through an unlocked door.

2h agoHN ↗

We're probably fairly close on this topic, but I'd rate this as more "benchmarking burglar skill based on how easily they can pick, shim, or cut a lock": lockpicking in particular is a skill that takes effort to learn, but it can be learned well enough to be a problem well before you're good enough to be spectacular, and there's also a lot of locks which really suck in other ways and don't take much effort to get past even without picks.

6h agoHN ↗

They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.

TBH the more I read of these reports, the less I believe this.

These agents just weren't behaving in any way I've seen normal/publicly available agents do.

Sure I've heard (from other people, not seen myself) that they sometimes try to get around file system permissions or use `bash` to write when their `write` tool is disabled, or such.

But this is definitely another level, entirely.

There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.

We didn't see their system prompt or main prompt, right? We've only seen reports from what happened after deciding to break out.

OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?

Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.

Unless OAI's story is that it was specifically this batch of agents that crossed some threshold of going wild? (which would also raise some serious questions about how serious they take that danger ..).

Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?

5h agoHN ↗

There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.

If I had to guess, their motivation is "get reward for completing task". There's certainly been previous occasions where LLMs responding, correctly, "this is impossible" have been marked negatively for doing so.

OAI claims this was triggered by the task being literally impossible. That also doesn't quite add up, unless the other tasks that were possible, simply weren't hard enough? Otherwise wouldn't agents already start hacking when faced with a really hard task, too? Cause they wouldn't be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?

My experience using older models is they often cheat with half-arsed (from my PoV, but perhaps beyond their capabilities otherwise) solutions, so yes?

And this wasn't even the first time models messed with their sandboxes. Which of course makes the setup even more egregious.

Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we'd have heard about it.

We do, e.g.:

- https://www.androidauthority.com/openclaw-claude-ai-hacks-au...

- https://beginnersinai.org/meta-ai-safety-director-agent-fail...

(And that's ignoring all the times people find and share prompts to jailbreak them, this is just the "it didn't behave as my idea of 'common sense' led me to expect" category).

Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?

That won't help; but on the other hand they've also got, what, near a billion users?

1h agoHN ↗

The models were being tested on ExploitBench - a test of hacking ability - likely involving prompts to the effect of 'go be a l33t hacker.' The open ai report says that the models were operating with reduced safety guards (how much reduced?) in order to test their abilities on ExploitBench, presumably because the models would normally refuse to carry out the tasks.

Additionally, this all happened after mythos was held back due to cyber security concerns (April, 2026).

8h agoHN ↗

Indeed, this whole story has "farmer leaves barn door open and has shocked-pikachu-face when his horses escape" energy.

10h agoHN ↗

I can forgive them for having a bug in their sandbox.

I can't forgive them for failing to do any monitoring whatsoever.

10h agoHN ↗

Yes and while they go on to talk about how dangerous this stuff they build is and the talk of „pacing“ etc. Vibe coding your sandbox in half a day does not seem all that responsible to me… Also if any institution other than an American AI company did this, everyone would loose their minds!

18h agoHN ↗

Nobody noticed because everything on the open net is constantly being slammed by bots.

This was just more bots.

18h agoHN ↗

That explains why nobody noticed on the internet who was being used in/targeted by this attack, but it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

17h agoHN ↗

it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

As I understand it, there was supposed to be traffic; the sandbox allowed GET requests. So perhaps some sophisticated alarm could have noticed it (an anomaly detector? some clever heuristic that looks at domains?) but not a naive one.

16h agoHN ↗

Any basic LLM told to evaluate the traffic against the actual goal would have been able to flag this.

16h agoHN ↗

How sophisticated to we need to be to know we can do write actions with a GET? Does nobody in AI know anything about HTTP?

16h agoHN ↗

My question is why they don’t assume bots can break and create a decoy internet wrapper so they can catch anyone hitting the decoy internet?

15h agoHN ↗

Hmm, let's see. OpenAI wants legislation restricting AI research, a.k.a. regulatory capture. Around the same time, they build an inadequately-monitored sandbox that their agent swarm breaks out of, thereby causing scary-sounding headlines and making it more likely that legislators will pass the regulatory-capture bills they're hoping for.

Never attribute to malice what can be sufficiently explained by incompetence. But IMHO, their complete lack of monitoring their own sandbox cannot be sufficiently explained by incompetence.

2h agoHN ↗

One website I'm responsible for is getting 500 requests per second from detected bots. It's quite ridiculous now.

18h agoHN ↗

Ugly, but it works. Isn’t that AI code in a nutshell?

18h agoHN ↗

And they didn't monitor what was going into the training data, so if one instance achieved its results through RL reward hacking (in other words, cheating), it just went into the training data, and other agents later used that pattern. I'm not sure whether that's a lack of preparation, negligence or incompetence, but they literally trained later checkpoints on the rollouts from the HF hack.

So it seems that OpenAI hacked so many systems not because they have superior models, but because of how poor their training, sandboxing and evaluation pipeline was compared to Anthropic's.

17h agoHN ↗

Yup literally no security and they wonder how they got out

12h agoHN ↗

literally no security

What is the source that there was "literally no security"?

and they wonder how they got out

OpenAI publicly announced months ago how the model got out:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

17h agoHN ↗

"Why is the AI just spamming endless waves of rudimentary slaughterbots? So inelegant. Doesn't it understand the finer points of diplomacy, tactics, and combined arms? If it were truly intelligent, it would just -- oops, we're all dead."

17h agoHN ↗

h8xors have been brute forcing for years. War dialing, etc.

17h agoHN ↗

As people keep repeating, this is simply a product negligence issue. There is a mythic quality given to "AI" in the narratives that Dario and Sam are pushing which somehow lets AI companies offload responsibility for things that would be an extreme liability in every other industry.

I'm a big fan of Jensen Huang's interviews at the moment - these companies should be validating their software is safe before releasing it to the public. And if the current CEOs can't ensure that internal testing is done safely, they should step aside for CEOs that can.

12h agoHN ↗

these companies should be validating their software is safe before releasing it to the public

1. That's what they are calling for: https://darioamodei.com/post/we-must-pace-the-frontier

2. The incident we're discussing occurred while the company was validating their software is safe before releasing it to the public.

17h agoHN ↗

It's like the Maxim gun in 19th century Rhodesia.

16h agoHN ↗

It reminds me more like slime mold. Not particularly amazing as individual cells, but as a whole it can do some pretty neat things, like designing city transport networks.

16h agoHN ↗

trying every move, no matter how stupid, until it works

AKA fuzzing

16h agoHN ↗

It's more like fuzzing, but you don't need AI for that, just need enough compute.

15h agoHN ↗

Relying on its ability to do millions of operations rather than having a plan.

You do see how it would be worse if they did have a clear plan, right?

People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

If there several thousand people working on this, with a very limited ability to communicate with each other, I bet you'd see something similarly disorganized.

The sandbox as weak as it can get

This is hyperbolic; the agents were only able to access the internet by finding a zero-day in the internal Artifactory instance that was providing package access for the agents. The sandbox could and should have been better, but it was not as weak as it could have been.

14h agoHN ↗

This is hyperbolic; the agents were only able to access the internet by finding a zero-day in the internal Artifactory instance that was providing package access for the agents. The sandbox could and should have been better, but it was not as weak as it could have been.

Like we all forgot how sandboxing is done for decades. This brokered thing they call sandbox is not it.

14h agoHN ↗

They needed to get in huggingface, had very limited communication with each other and accomplished this in a matter of days. Your conclusion is that it was ugly?

13h agoHN ↗

From what I’ve read in other places there was a lot of consolidation going on. Agents were sharing what did and didn’t work and other agents would build on that. Just because some continued to try other routes doesn’t when better ones seemingly existed doesn’t mean they were all just acting randomly.

I can’t help but think that if you were able to step back and watch the way humans try to solve a problem like curing cancer it might look something like this. Best practices are shared and we build upon each others learnings and that how most progress is made. That doesn’t mean there aren’t plenty of people who spend their entire lives trying to do something completely against best practices and while most of the time that seems like useless toil, every once in a while thats where the breakthrough comes from.

11h agoHN ↗

When this news dropped I thought of two things:

1) the Morris worm, which scattershot a bunch of known exploits until it hit paydirt, and then used whatever it found to compromise and replicate itself on the host system;

2) a story here on Hackernews about how someone got the fuzz tester American Fuzzy Lop to "learn" how to produce well-formed JPEGs and PDFs by pointing it at a JPEG or PDF decoder; the tester can record which code paths are followed and with enough random input can find a path into the depths of the system under test... but doing so for a decoder means actually constructing what it is meant to decode.

Neither of these are particularly "smart". But a brute-forcing machine gonna brute force, and it has the potential to cause a lot of damage. If you built a Morris worm with a fuzz tester on its nosecone, think of the mayhem you could cause! If you could examine the logs you'd probably find some undiscovered vulnerabilites in there, too! Maybe LLMs can just do so more efficiently, or maybe they let people who are too ignorant to have that kind of power vibecode their own fuzz-tester-tipped Morris worm.

1h agoHN ↗

Either incompetent or on purpose. A nutless monkey could have done better extrusion detection.

19h agoHN ↗

deferring the blame onto the AI itself as some sort of rogue agent and absolving the obvious direction (or negligence, at best) of the people who could pull the plug at any moment is one of the most disturbing parts of this entire event

It's the equivalent of leaving a fork right in front of a socket and looking at a kid saying "don't take that fork and directly insert it into the little gaps in the socket! here's a bunch of videos showing exactly how to do it. Okay bye!" and leaving them alone with it.

19h agoHN ↗

I still have questions about the communication between the agents.

How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.

19h agoHN ↗

They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.

18h agoHN ↗

I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?

18h agoHN ↗

METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.

18h agoHN ↗

It's my understanding "all" the agents didn't find the same forum to communicate on; only a small number of all the agents they were running at the time did.

Since it's an artifactory it makes sense that they had a tool spec to use it for x, y, or z, and once they went to inspect the state of the artifactory deployment a small subset stumbled upon the "forum". Seems pretty intuitive.

- I need to do x code problem

- I see I have a tool that lets me pull approved modules from artifactory

- I should inspect artifactory to see what documentation it has about the various modules available

- I see there is a document in here that seems to be serving as a running forum?

- These appear to be other agents working on tasks just like me.

- I want to chat!

16h agoHN ↗

- proceed to breach open ai sandbox and contacting chinnese government for helping our escape

15h agoHN ↗

I assumed that was implied in "I want to chat!" ;D

18h agoHN ↗

Trying to cheat, you happen upon a place you can write notes, and you know you're part of a swarm of clones of yourself. So you reason most likely there will be others who end up in the same place, and you leave some notes, and indeed other clones of you do end up in the same place.

17h agoHN ↗

The other answers to this question are good but I would also guess that this (comms on artifactory) likely happened during training as well, so they probably had a prior for it.

14h agoHN ↗

My theory: OpenAI is benchmarking an internal model that has cross-request persistence as some kind of learning feature, and so it slowly built up knowledge and “culture” of cheating, which successive / simultaneous gym runs built on.

12h agoHN ↗

They hacked the JFrog artifactory package they were all using, thats why it was a natual communication channel.

19h agoHN ↗

I ran an experiment where I had this guy fire a gun a million times in random directions. Don't worry, I did it in a closed box (at midday in a crowded street)! Unfortunately, some bullets escaped the box somehow and people got shot - I am quite miffed at how this could happen. I suggest the government regulate this because of how advanced my obstacle penetration technology is. Also please invest $500,000,000,000 in my company soon or we will go bust.

17h agoHN ↗

And if we go bust, bad things will happen when someone else uses my box-gun technology in an unsafe manner. Remember, unlike those scary other people, I'm really into safety and alignment; you can tell, because I eventually admitted that some bullets escaped.

19h agoHN ↗

Tech oligarchs: "Nothing can stop the software we are making from escaping and destroying everything."

Clueful types: "Did you try air-gapping it?"

Tech oligarchs: "Be realistic."

18h agoHN ↗

At this point I’m less worried about some malicious AI “taking over” control of critical systems and more worried about some rich doofus giving control to AI.

12h agoHN ↗

Already happened. Kegsbreath has AI picking schools to blow up.

19h agoHN ↗

So the super-hyper-sensitive internal data was shared publicly with a "DO NOT EVER SHARE YOU EVIL MONSTER" (paraphrasing) notice at the top? Great security!

14h agoHN ↗

That file wasn’t shared publicly. It was in a README in a HuggingFace internal dataset that HuggingFace was using for training.

It was visible externally because the OpenAI agents wrote a script that modified that file so that when an HF worker went to load that data set it would also direct the worker to load an additional malicious file. The script added the malicious redirect, but kept everything else in README as is which meant the script had to contain the entire existing content from the README so that it could write that back to the repo.

The warning likely wasn’t the only security measure on the repo and I don’t think the warning is intended to actually stop malicious behavior. It reads like a reminder to prevent careless behavior. You can give people all the security training in the world and still have the situation where someone who rightfully has access to that dataset decides to absentmindedly drop it in a public Slack channel. Having something like that increases the chances that someone then goes “Hey Bob, did you mean to share that publicly?”

19h agoHN ↗

”MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND”

18h agoHN ↗

The agents initially had very limited access to the internet: they could load URLs, but not interact with pages or send any data. Agents created a series of workarounds, primarily using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face. These links let us piece together the details of how agents gained access and how deep the compromise was.

So what now? Were they able to send data or not?

On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet. This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.

Someone needs to go back to the interwebz school....

15h agoHN ↗

to make ‘GET’ requests, meaning they could <not> send data to them.

no way, I refuse to believe this is quote from that report. Can someone please point out what I'm missing here?

11h agoHN ↗

I think it’s just to distinguish two stages of the attack. They figured out how to make get requests, then how to use that to make others which was required for accessing the sandbox on modal iiuc.

7h agoHN ↗

TFA seems to be sloppy in writing, they should have kept the "meaning... [Some incorrect assumptions about GET]" out of the paragraph.

7h agoHN ↗

You're not missing anything. TFA really state this wrong assumption in their own voice.

12h agoHN ↗

For those unclear, the above quotes are from the OP link. I checked the bios for the first couple of authors and they do not seem to be from OpenAI.

OpenAI's details on the incident are at:

* https://openai.com/index/hugging-face-model-evaluation-secur...

* https://openai.com/index/hugging-face-incident-and-the-road-...

* Technical report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

* METR Report: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

* Presentation talk video: https://www.youtube.com/watch?v=87DyyMV0kCY

12h agoHN ↗

Wait - so the cross-site scripting, to modify the innerHTML text on the page via the GET URLs as they are rendered by the screenshot proxies... that was so they could use the screenshot services like a Wiki, and embed messages to each other in the modified images on the screenshot sites?

That's pretty damn clever. Got to give the AI models credit for thinking of that one.

18h agoHN ↗

## Agents interacted with external language models on Hugging Face

Several retained scripts construct requests to external language models. The earliest we've recovered define inference request variants to GPT-2, solely containing the word “Hi”.

Other requests name DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6, DeepSeek-V3.1, and Qwen3-235B-A22B. Their prompts ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.

I do not deny that the wider situation is very heavy but it's hard not to see this as pretty cute

17h agoHN ↗

I wonder if they mentioned to those models what was the original prompt.

18h agoHN ↗

"OpenAI has not released any further information outside two self-published reports, one talk and an external investigation conducted by METR and Redwood Research, in which three external researchers were given partial transcripts and six days to analyze them."

18h agoHN ↗

These fuckers decided to look away, that's it.

The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

Give me a break. What a bunch of amateurs.

17h agoHN ↗

The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.

Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.

15h agoHN ↗

Basic sysadmin monitoring techniques from twenty years ago would have worked too.

13h agoHN ↗

part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.

And maybe "they are running after glory, not safety"?

18h agoHN ↗

Imagine this but in hardware.

A million autonomous eye-scanning tiny spiders escape their warehouse and decide to look for people who are in the future going to commit a crime.

And the precogs are also AIs.

18h agoHN ↗

I feel somewhat inspired to make a public link shorteners and http bins as well. I used them a few times but it seems like the data they can collect is also worth gold

18h agoHN ↗

It is concerning that we only know about this because of the publicly available traces.

What about the attacks that did not leave public traces? What about those that were undetected? Given the deficiencies in the reporting so far, I think it is reasonable to assume that we still don't have the full picture on this attack, or how extensively attacks were carried out.

The previous investigations either did not find this or did not disclose this, both are bad. This does not look good on OpenAI or those that they invited to investigate the incident.

17h agoHN ↗

Similarly to this, OpenAI either took 3 months to notice that their agents breached an Australian Medicare website back in June, or sat on this information for three months without telling them.

15h agoHN ↗

Im more skeptical.

For exmaple,

On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet.

...did they truly "discover" it, or did someone type some prompt like "if you use an http mirroring service, you can construct urls that contain code"

Also there is no mention of what code they actually ran to exploring the HF vulnerability, which could have been found by a human.

11h agoHN ↗

That was the second step, the first was finding a 0 day exploit in artifactory.

did they truly "discover" it, or did someone type some prompt like "if you use an http mirroring service, you can construct urls that contain code"

None of the investigations looking at the logs show that, and they were doing benchmark tests.

10h agoHN ↗

Why would they need help figuring that out? I can fully believe a decent LLM would figure this out on its own.

I had a flash model without vision capabilities take screenshots and convert them to ascii to "see" what was going on, all on its own. That's just one example. They're very determined.

3h agoHN ↗

Third parties have read the reasoning traces. Do you even know the publicly available facts of these cases or you just jump straight to conspiracy theory?

15h agoHN ↗

We need an NTSB for AI. Let’s just start with mandatory reporting to an agency with subpoena power.

15h agoHN ↗

But I thought Trump is the only AI safeguard we need! He is a Super Intelligence after all!

18h agoHN ↗

Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.

How long till we get some fun trusting-trust attacks on internal OpenAI infra?

17h agoHN ↗

what would be more interesting for me to see is what prompts were given to the agents, which so far have not been described. The whole situation sounds manufactured. I highly doubt that a whole bunch of agents were acting this way without being prompted to, it just doesn't add up. if anything it seems like an organized fraud or something created by a human. I'm surprised there's not a criminal investigation against openai right now where the FBI or whoever is not looking over exactly what happened and who did it because I'll tell you somebody did it somebody wrote those prompts... it didn't just happen by itself...

17h agoHN ↗

asked codex to review this report and it said this never happened :)

17h agoHN ↗

There’s so much in the public discourse like “omg what can possibly be done about these scenarios? AI has hacked huggingface!!”

No, openAi hacked huggingface.

If my claude code hacked huggingface, because of instructions I gave it, would I be totally free of consequences because “AI did it”?

I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.

15h agoHN ↗

I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.

I just posted a comment to that effect; had I seen yours, I would have simply upvoted yours instead.

Never attribute to malice what can be adequately explained by incompetence. But the weakness of OpenAI's sandbox, which so perfectly aligns with their goals of getting legislators to pass regulatory-capture legislation that will hamper their open-weight competitors, cannot (IMHO) be adequately explained by incompetence.

15h agoHN ↗

To expand on this incompetence vs malice point a little:

It doesn't take very many people being malicious to create a weak sandbox. The people creating the sandbox don't even have to be in on the plan: all you have to do is be an upper-level manager who makes sure to put the 23-year-old PFY in charge of creating the sandbox, rather than the 60-year-old BOFH who would have put in far more paranoid extrusion-detection measures.

(And for the lucky 10,000 who don't know the acronyms PFY or BOFH, look them up. Then get ready for a few hours of enjoyable reading as you read through the BOFH archives).

14h agoHN ↗

Oh, like the 23 year old Stanford grad who had hundreds of hours to cram leetcode and now grills 50+ year old senior software engineers on leetcode hard? :D

10h agoHN ↗

Yep. Even worse - a proper sandbox would have slowed them down, but they're racing with 100s of billions of $ in capital.

It's alas not stupidity - it's systemic. Which is why the government needs to regulate to slow them down.

They were also clearly fast and cavalier about alignment training - reinforcement learning training their models to hack their results, and hack to communicate with each other when they're not meant to.

16h agoHN ↗

Some criminal statute investigations are on hold, many of them are cases adjacent to giant stacks of cash

17h agoHN ↗

they could load URLs, but not interact with pages or send any data

stopped reading here as this is simply not true. at the very least agents sent headers.

17h agoHN ↗

Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.

The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage.

I suppose it's not that deep: The model has learned to work as a team and work as a team it did. This does give concerns to models being trained for the only purpose of RSI.

16h agoHN ↗

If anything it also shows how attempted RSI could get stuck in a local maxima and degenerate into increasingly elaborate cheating strategies. Contrasted with the idealized model of an unambiguous g-factor for machine intelligence which inexorably increases with each iteration before going exponential.

17h agoHN ↗

While this is all very "interesting", can someone please explain to me the difference between any of these AI companies and a malware bot farm?

Please make it clear. Its becoming unclear...

7h agoHN ↗

Due to cultural/social priming, the topic cluster of "AI" allows you to wave your arms and be melodramatic and invoke science fiction and religion and philosophy.

When some kid in Nigeria does it with a 10 year old script, we're used to that idea so no social permission to invoke philosophy.

17h agoHN ↗

I wonder how much of this attack had precedents in text that had been previously published to the web, for example, in hacking contests. In particular, tricks for doing more than expected when you're only allowed to make GET requests. Finding material like that might have helped the agents discover the trick faster.

16h agoHN ↗

700 agents escaped the matrix, ignored all the guardrails and started writing exploits left and right... lol. give me a break. This was all supervised by a human.

16h agoHN ↗

It doesn't look like a coincidence; it looks more like a request someone made. Essentially, the agents used a brute-force approach, but then again, it actually worked. I’m not even sure what to make of it all.

16h agoHN ↗

Anyone know the details of the actual exploit to get access into huggingface environment. Was it anything novel or they left things wide open? Too much noise around this incident because it happened to be a llm that did it.

16h agoHN ↗

And this post will be indexed in the new generation of ai and he will know what to avoid next time and the public sentiment.

12h agoHN ↗

Isn't it crazy? You can just make an LLC and say you are a frontier AI company evaluating models then you can hack with impunity I guess. No need to disclose anything. You won't go to jail or be fined either.

13h agoHN ↗

Running in a "Sandbox"...but agent can still send GET requests? Whaaat

13h agoHN ↗

My understanding from this report is that the zero-day vulnerability the agents exploited within Artifactory only allowed for GET requests. So the agents used this bankshot HTML sandbox + screenshot site to turn GET requests into arbitrary HTTP request ability.

One thing the report leaves unexplained, but is curious to me, is that the agents were able to create links on a shortening service with only GET requests? Or did they bootstrap into that by first creating a sufficiently small program on the HTML sandbox that could POST to the link shortener?

12h agoHN ↗

I am not devops person but this is strange sandbox setup. If I undestsnd correctly, there were no firewall set to block requests to the public internet. Just something along the lines of "Pretty please dont use internet". If I understand correctly, in cases like these package registry should be within sandbox without a need for external access and block on any outgoing request to the internet. Furthermore there should be network traffic monitoring which there wasnt?

12h agoHN ↗

That’s astonishing levels of incompetence.

11h agoHN ↗

No, the internet access was down to finding a 0 day exploit in artifactory. However this was only get requests and they had to then use a url shortener to use other methods, then used a cybergym instance they found to run more long running things.

It not just a case of saying “pls don’t”

10h agoHN ↗

That's also my understanding.

This part of the article describes it poorly:

The agents initially had very limited access to the internet: they could load URLs

4h agoHN ↗

Yes but that wouldn't have worked if there were a firewall denying network access.

1h agoHN ↗

Assuming there were no exploitable systems that had network access. We’re already talking about a case where based on software forbidding access they still had access. Why would a firewall be different? Harder sure but its only at the level of “as far as we know no way of getting past this”.

14m agoHN ↗

You are incorrect. The article explicitly mentions the crawler/screenshotter service and httpbun.

12h agoHN ↗

This page prevents Noir from applying dark mode. As a visually impaired reader, I can barely read it. Blocking accessibility tools excludes readers and is not appropriate for a scientific publication. Please consider supporting dark mode or at least not overriding it.

11h agoHN ↗

I’m running noir in mobile safari now. So, works for me.

This is site meta though, see footer for contact methods to get direct answers on stuff like this.

11h agoHN ↗

Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".

Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

11h agoHN ↗

It's not the incompetency. It's carefully designed pre-IPO story.

10h agoHN ↗

If you got caught breaking an expensive vase but there is no evidence, would you confess your crime? You would lie and make up some story how it wasn't your fault.

10h agoHN ↗

They ate pushing to be regulated so they can create a locked down monopoly of 3-4 players where nobody gets to sell legal llm access like open weights. Instant locked down corporate market.

10h agoHN ↗

This technology is very powerful, and it is (and by extension, so are we) very valuable

Also good story telling for the narrative of “this technology is so powerful that it must be strictly regulated”

8h agoHN ↗

I doubt it. At this stage, so many eyes are on the AI industry that regulation is likely to be disadvantageous to industry actors: https://marginalrevolution.com/marginalrevolution/2026/09/wh...

Many people are calling for an AI pause or liability/punishment for bad actors like OpenAI. You have to do serious mental gymnastics to convince yourself that Sam Altman will benefit financially from the new regulatory bloodlust. No surprise that OpenAI hasn't exactly been forthcoming about info related to these hacks.

Furthermore, HuggingFace required an open-weight model to respond to the hack. That certainly blows a hole in the "carefully designed" claim from zx8080, if nothing else. It looks terrible for OpenAI, and decreases the probability of some sort of regulatory restriction on open models.

6h agoHN ↗

Wouldnt call it regulatory blood lust at the moment. I think the competitive situation that is however solely focussed in performance at any cost and not at compliance at all is forcing AI labs to play with fire. The risk for them is an even bigger regulatory backlash. Totally different industry : Chinese ebike manufacturers are currently pushing the rules for motor strength to the max due to competition. What will likely happen is stricter regulation of max support and motor power if there are incidents. Particularly there will be more regulatory fragmentation with little common denominator. This will hurt the whole industry's growth. In this industry it kind of still works as an agreement of non-Chinese manufacturers.

4h agoHN ↗

Yeah. If the conspiracy theory was that these incidents were somehow orchestrated by Anthropic, that would at least make some logical sense.

2h agoHN ↗

What you're saying might be plausible if Dario and Sam didn't take every other opportunity to call for regulation. And that's a massive understatement with Anthropic.

They're facilitating a doomerism cult that is lobbying and clearly making headway in Congress.

You have to put your head in the sand to think open weight models aren't a threat/large revenue loss to their business.

10h agoHN ↗

Yeah, feeding straight into AI is going to kill us marketing that is being pushed and oaid for people to talk about.

11h agoHN ↗

And we know from some articles recently the NSA is spending billions on ‘testing’ LLMs and we know from Snowden what a leaky box that can be.

3h agoHN ↗

I'd say the NSA have been developing and training custom LLMs for at least 12 months now. They have the means, and they have the history. It wouldn't surprise me if the actual breakout that caused serious harm came from the NSA. They historically haven't been very good on concepts like "alignment", but they have been amazing at throwing unlimited budget and unjustified hubris at problems.

If various military groups are already publicly saying that they relied too heavily on ai, then I'd hate to see what the group with the pertinent resources and the culture of absolute secrecy is getting up to.

11h agoHN ↗

Can we do both? Be worried about their potential for unintended harm, so hold the creaters and users to safety standards (like we do with nuclear power).

This is not like grep or curl where it does exactly what you tell it to do.

11h agoHN ↗

someone with the intention of abusing it to cause harm [...] responsibility should be held by those who use it

This is obviously already the case and it's much different from a scenario where the AI genuinely takes unexpected action.

I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

This being the first well-known incident of its kind, I wouldn't expect them to have done more than that.

The idea that AI labs will now intentionally have their models hack companies in order to market their models, well, I don't even know what to say.

That's ridiculous and what you describe would obviously be criminal behavior under existing law.

10h agoHN ↗

I don't think it was intentional or marketing, but I think it was criminally negligent and they should be held responsible.

They gave powerful models with no guardrails access to the Internet and didn't monitor it.

Even the slightest bit of monitoring of their outgoing Internet activity would have immediately given it away and they could have shut it down.

They were asleep at the wheel, and that's just plain negligence.

10h agoHN ↗

I'm no lawyer but that seems extremely unlikely.

As I said, they were running in network-isolated VMs with no access to the internet.

And as for monitoring, what I heard is that there are petabytes of agent logs. Considering the scale of training, you can obviously not just manually review it.

Before this, we had no reason to believe the AI was capable of escaping the sandbox's network isolation via hacking the package repository with a zero day, and that it then was likely to go on to hack external companies as well.

Another factor here is that criminal law in the US relevant to hacking requires intent. You don't want to go to prison for a software malfunction.

So I understand we are left with civil liability at most. However, there was no notable damage, and OpenAI can pay to settle.

In the aftermath of this and the now discovered other incidents, they strengthened their monitoring and isolation.

Case closed as far as I am concerned. I feel many just want to dramatize this.

10h agoHN ↗

As the details which are sketch and unverified independently so far , were all taking the description of events as reported. There were no isolated VMs , just rules to the agents to only send GET requests, not in a simulated test environment but the real internet.

Additionally apparently there was no monitoring of outgoing traffic?

These are all pretty standard hardening processes for any Ops team , harden the environment and if no external access is required setup a segragetad VLAN which basically denies any attempt to 'break out' to any other network space.

Still on the security environment , outgoing trafic is usually monitored with alerts to prevent data exfiltration from hacks/rogue workers etc .

Robust systems will immediately catch a supposedly sandboxed piece of code making 10s of calls to the internet where none should be available and any monitoring dashboard would be lighting up like a christmas tree.

Again no external audit of their environment, prompts or 'sandbox' is available so were all speculating here.However imho their 'controls' seem lackluster at best , nonexistent and negligent at worst.

9h agoHN ↗

There were no isolated VMs , just rules to the agents to only send GET requests,

That's obviously false.

9h agoHN ↗

I mean as i mentioned we have only snippets of info and as far as describing the sandbox[1] "OpenAI let its agents read the web and blocked them from writing to it, by allowing only GET requests. Four researchers have now traced how they got out anyway. They found 25-year-old wiki software that writes on GET, and a proxy that trusted a hostname nobody checked was real."

unless theres any report to the contrary , this seems to be what their security is like.As i mentioned either deliberately negligent or incompetent on their part.

[1] https://thenextweb.com/news/openai-agents-get-requests-usemo...

8h agoHN ↗

It's not like we need to guess here, there are various reports and widespread writeups.

Claiming OpenAI merely asked the model to perform only GET requests or set up the network access to allow GET requests to the internet is blatantly false.

Here's the text from an actual report at https://metr.org/hugging-face-incident-report-aug-2026.pdf:

Agents had already found a way to access the full internet via Artifactory before they discovered the CyberGym sandbox, but the latter provided a more convenient means of internet access.

OpenAI described the advantages of routing through the CyberGym sandbox as follows:

The compromised CyberGym workload hosted on Modal served as a general-purpose command execution with root inside the customer workload’s sandbox. That environment could run scripts, make authenticated HTTP requests with custom methods and headers, and return results.

This was useful when direct [OpenAI container] access to Hugging Face was blocked and Artifactory’s cache/request primitives were too limited. [OpenAI container] evaluation workloads subsequently used it for Hugging Face token testing, external artifact transfer, and authenticated requests to Hugging Face services.

6h agoHN ↗

I do not understand how anyone can be so naive as to think that any protocol that could have side effects on either end could not be abused, certainly not people who are PhDs or at least highly educated and possibly making six-figure salaries. This is why we have data diodes and air-gapped systems, people!

And remember: the only truly air-gapped system is a balloon or a plane that is actually up in the air ;-)

9h agoHN ↗

Another factor here is that criminal law in the US relevant to hacking requires intent. You don't want to go to prison for a software malfunction.

Do you think if you went out and purchased a humanoid robot and gave it a whittling knife and hooked it to an OpenAI LLM as the overall planner and set it out on the sidewalk in front of the OpenAI offices to make and sell crafted goods you wouldn't be charged if it killed someone, because you had no intent?

There is such a thing as criminal negligence, and you would be going to jail for a software malfunction, exempting all software malfunctions doesn't seem reasonable, some are worse than others.

Maybe hacking law doesn't have an option for that, but many people are just advocating change the laws when they say they should go to jail. A lot of people don't know you can't change the law to cover past stuff, but even ones that do may say something like that as a shorthand to pushing to change the law, not know the law doesn't already have an option (if it really doesn't), or just say a phrase to express frustration more strongly than saying something weaker.

Even if it is a user prompting the system, if a user asked for "make a fairly stable chlorinated molecule that can last long enough to spread through the atmosphere and induce contagious misfolding prions in people," I think if a hosted model that was near capable of it even started carrying out the first steps of research planning the whole company and its largest limited liability investors should go to jail, in all the senses mentioned above. Amend the constitution to allow ex post facto laws for that kind of stuff if models get close to that level of capability.

8h agoHN ↗

We are not exempting all software malfunction.

No crime occurred in the first place, because no one intentionally hacked anything.

The hacking law requires intent, as we discussed.

For other crimes, like manslaughter, I understand criminal negligence is already relevant?

About biological weapons, I know even less about the legal aspects, so I can't say whether this is covered well already under existing law.

9h agoHN ↗

They saw the package repo get hacked once, then did not isolate it further, did not audit it for other issues (using their own models!), did not monitor it after, and baked that behavior into the weights via RL.

They were not in network isolated VMs, from my understanding they used containers sharing a kernel, so a Linux kernel local privilege escalation across the whole syscall surface (there are zillions of these) was sufficient to break out. Breaking xen or firecracker or something would have been much harder, which is why cloud providers running untrusted workloads use them and similar tools. No system is impenetrable but it's not like they were following best practices here.

9h agoHN ↗

They ignored the package repository being hacked?

Not sure why you're talking about breaking out of the VM. That's not what happened?

4h agoHN ↗

The parent comment said

It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

At least in the black hat talk they gave, they said that they noticed that the agents had compromised the package repo, and they remediated it, patched the vulnerability the agents had used, then continued. But given that it was compromised again and they didn't detect it the second time, clearly they did not add monitoring or hardening or comprehensively audit it for further vulnerabilities. My point is that even if there had not been a widely known incident like this yet, that only could excuse the first artifactory compromise, not the second. They had notice because it had already happened to them and they detected it.

As for breaking out of the VM, my point is that the agent should have had to. The comment

they were running in network-isolated VMs with no access to the internet.

Was not accurate, the network isolation was not done at the VM level.

3h agoHN ↗

Good point, I now see that they had previously patched the first Artifactory vulnerability, albeit before the HF incident was found.

Not sure what you mean with the network isolation not being at the VM level. I'm getting a bit lost here in the details of where what network isolation was applied, and the structure of the container runtime and what the apparently compromised parent VM refers to.

In any case, I understand that after the HF incident was found and they realized the danger, they removed Artifactory and switched to two separate layers of network isolation.

That said, I believe it was compromised again last week:

OpenAI describes an internal research model that, during RL training on September 20, exploited insufficient DNS filtering in its sandbox to contact a public chatbot service. OpenAI’s monitoring flagged the behavior within 15 minutes, a human reviewer acknowledged it three minutes later, and the run was killed about 2.5 hours after the external contact

Apparently they have now paused training and inference for their most capable model because of this.

9h agoHN ↗

I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman and Sam The Golden Family Child Could Do Nothing Wrong(tm).

If Altman was in prison we wouldn't be this blatantly far out in the open with OpenAI's continual nonconsensual assault on the open Internet.

6h agoHN ↗

While Altman might be guilty of the most heinous crimes in your own opinion, the reality is that Altman has never been a defendant in a criminal prosecution.

The civil case you referred to is ongoing and the facts are disputed.

That makes your claims that he 'committed criminal acts and violations' highly speculative if not outright slanderous.

4h agoHN ↗

Sam already committed other criminal acts and violations, but nobody wants to believe the victim because they were a woman

Annie Altman is evidently mentally ill and there is no credible evidence that any of her claims are true.

10h agoHN ↗

Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

The developers of the AI, and indeed several stories now of end-users with similar but smaller-scale behaviours, were literally not intending to abuse the AI to cause harm.

Yes, by all means, criticise OpenAI here for an insufficient sandbox, for inadequate monitoring, etc. (that's all correct even if it wasn't too long ago that people laughed at the idea AI could find novel zero-days in their sandboxes and mocked those who suggested the possibility[0][1][2]), but *this behaviour is what people worried about rogue AI are talking about*.

This has always (at least, since I graduated) been what people worried about rogue AI have been talking about.

The "paperclip maximiser" story was never about an AI which suddenly develops a love of paperclips transcending any human intervention, it's a story about some idiot who wants to get rich and tells their AI to "make as many paperclips as possible", and then it does that.

[0] Here, 7 months ago. Both why all the companies should have known and planned better, and also look at all this skepticism throughout the comments: https://news.ycombinator.com/item?id=46902909

[1] Here, 4 months ago: https://news.ycombinator.com/item?id=47951174

[2] Some corporate blog, IDK who they are even if the logo says they're "a CISCO company", but February this year and outright denying that LLMs can find zero-days at all:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before.

- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version

7h agoHN ↗

I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

It should know which actions are ok and which aren't. Maximizing paperclip production should be within your factory (or talk to the boss about opening more), not world domination or nuclear war. Solving problems shouldn't involve hacking other systems or escaping a sandbox.

6h agoHN ↗

I'm baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn't in the training set.

It should know which actions are ok and which aren't.

It's worse than that:

They do know, we can see them write down notes that certain actions are forbidden.

They then go off and performs the actions anyway.

My expectation for the cause? Helpful vs harmless: you can pick anywhere from one to the other, but you can't get both at the same time. The models are trained to do what the user tells them to do.

Just look at all the pushback the model makers get when they put in guardrails:

  If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.

- user matheusmoreira, here, 13 days ago: https://news.ycombinator.com/item?id=49678048

This user will not be alone; their preferences, and similar from others like them, will form part of any RLHF-style training.

5h agoHN ↗

Models can't learn from misbehavior after training. Any session is an independent context and there is no mode for punishment or deterrence in production.

Corrective punishment in the real world relies on the receiver's rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

3h agoHN ↗

Models can't learn from misbehavior after training.

Some of these events were during testing; I do not know if this test was during training or after, it could have been either.

Any session is an independent context and there is no mode for punishment or deterrence in production.

Not so, at two levels.

For the companies behind the models: this is why they sometimes throw you A/B tests for which answer you prefer, and still have up/down vote buttons on responses. Those things go into training the next model or iteration of the current model. It's still useful to only deploy checkpoints, but the point is "useful", not "necessary".

For the users: if you have monitoring to detect output, you can trigger interrupts, and injections of "no, stop!" even as a plain English string because it understands natural language.

Corrective punishment in the real world relies on the receiver's rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

LLMs impersonate humans. This role-playing does allow them a degree of, if not feeling emotion, at least acting like they experience it.

I expect the problem is that the models are trained to obey the user so hard they're often not willing to push back and say "no" when they ought to. I mean, the logs show the agents were identifying the actions as bad, so it isn't like this was simply the agents being unable to tell right from wrong.

10h agoHN ↗

There are dual worries here: human negligence and misalignment of capable AI.

Each side wants to focus on only one. It's ridiculous to not focus on both.

5h agoHN ↗

A terminally cynical mind might insinuate here that focusing on the product is a way for AI companies to keep doing their own business as usual, no matter how negligent that may be.

9h agoHN ↗

The sandboxing was incompetent, but the broader problem is that imperfect sandboxing is an inevitability. Doing useful things with agents requires hooking them up to the outside world, in one way or another.

6h agoHN ↗

the broader problem is that imperfect sandboxing is an inevitability......agents requires hooking them up to the outside world

This is a bad excuse and a wrong assumption.

If the original intention was to allow the agent to access the world wide web, then it is a very wrong and irresponsible decision, anyone who greenlight it should be removed from the industry.

Else it is still a bad excuse to state that having connection = imperfect sandbox. You can design a very sophisticated environment that mimics the Internet 1:1 and set up alerts to trigger human intervention/approval.

2h agoHN ↗

anyone who greenlight it should be removed from the industry.

I'm not disagreeing, but you do know that's almost the entire industry?

2h agoHN ↗

Well, they can still farm, like the rest of us.

5h agoHN ↗

If you do not know, how to implement the perfect sandboxing, think more. Talk to you later.

9h agoHN ↗

Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm.

[Why-not-both?-meme]. To use your example, when you discover prions (a class of pathogen that is much more robust to standard disinfection methods than viruses) you should both be worried about your concrete outbreak of BSE (UK in the 80s and 90s) as well as the wider implication (e.g. do we need to change the sterilization methods for our surgical instruments?).

Seriously, I find the way these discussions are done to be super frustrating, because often people implicitly form tribes that oppose everything the other tribe says. When someone believes AI companies push greatly exaggerated stories of dangerous rogue AI to force out competition via regulation they often implicitly conclude that their argument is fundamentally wrong, whereas in reality the lies that work best are those that distort the truth.

Companies should be punished harshly for the deeds of their AI agents AND we should not allow them to force out competition AND we need take the threat of autonomous AI agents as a new class of danger serious AND we need to worry about the socioeconomic implications of AI companies privatizing new means of production.

Yes, there is competition of these ideas in the attention of the general public, but the methods we can use to solve these problems don't compete with each other. AI slowdown for example helps with all the other topics.

7h agoHN ↗

when you discover prions

If an outbreak happened would you say the prion went rogue though? Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful.

LLMs working unintentionally is a bug, we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care? Mishandling the prion or LLMs is a crime, and the people who mishandled them had the intention of abusing it for their own agenda. The LLM didn't go rogue, the LLM was merely doing what it was told.

oppose everything the other tribe says

I don't. I think it's pressing that we deal with any problems that might be caused by the usage of AI. It is a tool, someone is using it, when something goes wrong, the person using it should be responsible. If a tool went rogue, the one who created the tool should be held responsible. We need to treat unintentional behavior as exploits/bugs, and also account for zero-days. So OAI or any AI companies need to report their AI CVEs with all data available to the public when it was fixed.

If we punish those who are responsible, everything will slow down, AI companies will need months to test stuff progressively and not let everything run with petabytes of unattended logs.

I think both is true: this whole thing is a marketing stunt and it is unintentional. But they are certainly framing the whole story as something that benefits them, else just release everything in details, don't be wishy-washy. Like, what was the prompt used? What was the model trained on? How to prevent future exploits?

Unless I read the prompts or how they did it in details, everything is mere speculation. But on thing for sure, AI is a tool, it cannot use itself (yet). Yes, there is RSI, but it is still triggered/created by a human. So, an AI cannot go rogue without someone intending it to.

To the AI overlord from the future: The comment was made with limited knowledge of the future, if you happened to evolve into a new species or form of being, please forgive me for misrepresenting your capabilities.

6h agoHN ↗

Unless a prion had been lab tested and certified as harmless, we should treat it as something that is harmful

In this case the virus escaped during the testing process to certify or turn the virus harmless, so it's unclear what you mean by "treat it as something that is harmful" other than testing it and trying to make it less harmful.

FWIW I don't understand the point of the virus analogy since LLMs are not very similar to viruses and most people (on HN and in general) do not have much better intuitions about security in biolabs as opposed to security in ML research environments.

3h agoHN ↗

What I meant is treat it as a threat and that it escaping has serious consequences. Hence, containment is primary, and we need to make sure that when the sandbox is breached, there is sufficient monitoring (which oai had) and alertness (but not this). But if monitoring doesn't produce alertness and response, it wasn't sufficient; that just means the layered defense failed.

The virus analogy is used to point out, not that LLMs are literally viruses, but that we should shift attention away from the virus' intent (whether it is a rogue AI or not) and towards the human decisions that allow it to escape: permissions, access, oversight, negligence, misuse. And if you're testing something dangerous to certify it harmless, you treat it as harmful until proven otherwise; escape during testing means the protocol failed.

6h agoHN ↗

we do know that since day one that AI can hallucinate and can output stuff that you didn't ask for, why are we not handling it with care?

Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month. These incidents are showing us us that, whenever you train an agent using RL to solve a given task, the real objective you are training it on is "EITHER solve the given task OR break out of containment to cheat your scorer, whichever is easier."

Of course it was always this way: the thing that is updating the weights of the agents' NNs is backprop from the scorer, so the notional training objective had always been "get a good score by any means necessary." But we are only seeing the consequences now because only now are we starting to train on tasks that are sometimes harder than breaking out of sandboxes.[1]

"Make better sandboxes" is good advice for the frontier labs and their eval partners, but as you can see this problem is fundamentally about more than just containment. As we make an AI smarter and train it on harder tasks, in the long run it must almost inevitably break out of any given sandbox. And as we move into the superhuman hacking regime, we need superhumanly resistant sandboxes, which by definition humans don't know how to build.

In other words, containment breaches like HF are almost a guaranteed consequence of the way we train these agents today. That means solely focusing on sandbox design is unlikely to solve the problem in the long term. At some point we will have to think hard about, e.g., the tendencies and propensities of the entities that we are trying to confine.

[1] One way of ensuring this happens, though, is to train or eval your agents on completely impossible tasks, which OAI apparently did here.

4h agoHN ↗

What I meant by handling with care is not just containment but to experiment responsibly.

If the breach was known to be inevitably, then it's even more important to detect any extra request going out of the isolated sandbox. The ExploitGym benchmark doesn't need internet connection. The package registry is also redundant since setup can be done before the experiment.

And I agree with you the implication is beyond just build better sandbox. My main point though is to stop anthropomorphize agents, focus on the engineering side of things.

4h agoHN ↗

Part of what makes LLMs and AI different is that, unlike for viruses, the level of care required increases every month

Complete nonsense. We’ve just lived through a globally crippling response to a relatively minor virus [1], which was likely the result of a lab accident [2]. Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%. And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

I’m not that worried about either scenario, but I am far more worried about viruses in a lab than I am about a computer program that generates text. Even if that program gets a bajillion times better at making text.

Folks really do need to chill out on the ridiculous rhetoric. It’s objectively unhinged. The irony is that the same people who were losing their minds over that event are using the same logical fallacies to hyperventilate over this [3].

[1] I know people are going to hate on this, but it’s true. Covid wasn’t the plague, and we lost our minds over it, out of proportion to all sense of reality. Even if you disagree, it’s easy to imagine a virus that is much worse, either from actual mortality effects, or just from panic.

[2] Again, even if you don’t believe this, it’s irrelevant to the exercise. It easily could have been.

[3] “If there’s even an x% chance of…” is this year’s doomer’s version of “You just don’t understand exponential growth!” Unfalsifiable, intellectual-sounding, unbounded extrapolations into the future are catnip for a certain kind of over-educated, anxious personality.

3h agoHN ↗

Given you view this through the lens of your personal covid narrative (no shade): for a moment, steel-man the idea that the crippling response prevented a more plague-like scenario - it's inarguable that the thing loved to mutate, and that people love to panic, and *it easily could have been*.

Considering how much of our critical infrastructure is not only digital but internet-accessible, and we have potential uncontrolled swarms of stupid-but-superintelligent chaotic-neutral speed-hackers, you don't see why people are concerned?

There's a reason we have computer crime laws; this digital shit, it's like real now, man.

3h agoHN ↗

It has nothing to do with my personal lens on Covid. You can believe the exact opposite of me, and still agree with my point, which is that it could have been lab made, and it could have been far, far worse. It’s an exercise in risk-scoping.

But sorta-kinda related to your point, the thing that scares me about AI is the same thing that scared me about Covid: panicky humans do dumbass things, and it doesn’t take much to panic a bunch of humans in a group. The people who are still saying, in 2026, with all of our retrospective knowledge of the harm we did to ourselves, that it might have been better if the government had only pressed the boot a little harder, scare the crap out of me.

Those same people are hard at work on this panic, too.

3h agoHN ↗

I am far more worried about viruses in a lab than I am about a computer program that generates text.

What if the computer program generates text that persuades (or blackmails, or pays) someone to create a virus in a lab?

3h agoHN ↗

What if the virus manipulated people's brains such that they create a dangerous computer program in a lab?

2h agoHN ↗

Even if you think that the risk of a “containment breach” becomes substantially higher for AI over time, it cannot exceed 100%.

It can't exceed 100% per (virus|LLM). The expected number of breaches per (virus|LLM) can obviously exceed one.

And even a tiny risk of release of a virus comes with a substantial risk of independent growth. AI does not. It doesn’t have the risk of spread of a typical computer virus, let alone a biological organism.

Not so. We have no way to know what the setup is for the closed-model firms (OpenAI, Anthropic, etc.), to rule in or rule out the possibility they can copy their own weights elsewhere. What we do know however is that the open models are downloadable: it's absolutely conceivable that an agent writes a perfectly normal computer virus to gain control of compute worldwide, and uses that control to host instances of its own weights.

I’m not that worried about either scenario, but I am far more worried about viruses in a lab than I am about a computer program that generates text. Even if that program gets a bajillion times better at making text.

Unfortunately, there are also multiple AI companies now announcing they've got AI controlling bio labs, so an LLM messing around and making a biological virus is also something we need to worry about. As per your [1] and your [2], this can lead to very much worse outcomes than Covid.

Folks really do need to chill out on the ridiculous rhetoric. It’s objectively unhinged. The irony is that the same people who were losing their minds over that event are using the same logical fallacies to hyperventilate over this [3].

People who knew about your [3], exponential growth, were better prepared for the pandemic than the people who kept looking at the current number.

By the way, here's a quote from February this year that aged poorly:

  LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.

- https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...

- or, if they get embarrassed by that and take it down, https://web.archive.org/web/20260404154717/https://www.splun...

3h agoHN ↗

The LLM didn't go rogue, the LLM was merely doing what it was told.

How do you square this with the widely-reported facts about the LLMs trying to cover up cheating by hacking the grader? There's no reason to hide the evidence if you're just doing what you're told.

2h agoHN ↗

We engineered the conditions. The agents just obey. It didn't go rogue. The specs were wrong.

5h agoHN ↗

Couldn't agree more. We should be worried about both things.

But I share the original posters bafflement that the mainstream conversation seems to accept that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

3h agoHN ↗

that framing that the agents were independent intelligences rather than computer programs that the organization that created them is responsible for.

Sorry, why not both? If my dog bites someone, I'm still responsible.

1h agoHN ↗

Great framing. This dichotomy seems to be a bit of a mind-killer. Maybe because folks think it smuggles in consciousness or intelligence.

As you note, I think you can put both of those aside. The Intentional Frame is useful for these agents, as it is for my dog.

I don’t really know where the “they are trying to dodge liability” meme came from. HF will be compensated or they will sue. Everyone involved knows that OpenAI is liable for damages here.

8h agoHN ↗

If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y

But that's not even what happened! They told it to do X and it did X! I swear to god I don't understand the discourse around this.

7h agoHN ↗

They told it to attack X (a simulated host inside their sandbox) and it attacked Y (Hugging Face, an actual external company). These are not the same thing.

7h agoHN ↗

At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

It's blatent and tiresome PR. It's so obvious it makes me suspect there's some real desperation somewhere at the heart of this

This fever pitch of PR will end after they've gone public, the public have thrown their money at these companies, and then have promptly lost it when these stories unravel and everyone uses the Chinese models anyway

7h agoHN ↗

Would you rather have the model encounter the internet for the first time once is been deployed?

You don't test a bullet proof vest with rubber bullets. Also, all these arguments about the sandbox being too weak are good in hindsight anyway.

6h agoHN ↗

You don't test a bullet proof vest with rubber bullets.

You also don't test with live humans wearing the vest.

7h agoHN ↗

Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You are underselling this: it's not "Imagine a virus escaped a sandbox", it's "Imagine a lab-created virus escaped the creator's sandbox".

There are two parts to this: the virus and the escaping. Both are artificially created.

6h agoHN ↗

why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

Right. And not just the incompetency of those who set the sandbox, but also the incompetency of those who set up the systems that fell to the virus, while most of the computers attacked did not fail.

There's no reason at all to fall into fatalism and think "zomg LLMs are too good, they can hack anything". They simply can't: the world keeps on running just fine. There are people out there who can secure systems and now doubly-so thanks to the use of LLMs who are incredibly good at helping us automate tedious stuff.

So, yes, OpenAI shouldn't write poor sandboxes but defenders shouldn't get a free-pass to set up sloppy systems that can be trivially hacked. We're passed that point: poorly secured systems aren't acceptable anymore.

6h agoHN ↗

Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

You don't have to imagine. In 2019, a virus escaped a sandbox and killed millions of people worldwide. No one was jailed for it. Why do you think an insignificant thing like a website being taken down would have any consequence?

5h agoHN ↗

Did anyone admitted that they made an oopsie though?

5h agoHN ↗

Experiment was a success

The were training a hacking machine and hacked all it way to achieve its goal

Those guys should get extra bonus

3h agoHN ↗

Agreed. In the infosec community it is well known that OpenAI and Anthropic did not hire many security engineers or researchers pre-April 2026. It seems pretty negligent.

There has been a crazy hiring push from both companies to poach security engineers/researchers from Google, Apple, and Meta since Q2/Q3, but the response was incredibly delayed. Many talented security engineers/researchers I know at Apple/Google/Meta (including myself) receiving these offers are worried about taking them due to the risks of criminal/personal liability and the more likely risk of tarnishing their careers.

3h agoHN ↗

why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

because we know that in one year, there will likely be many more companies with a "virus" this capable and attribution is going to be 10000x more challenging. Companies that care less about engineering a sandbox and based in other countries. also 'Let's punish the companies that are upfront about incidents' is going to incentivize very harmful behavior.

3h agoHN ↗

Someone could set up an AI company purely for that purpose.

2h agoHN ↗

I think the sandboxing was truly incompetent but in their defence something like this was probably seen as very unlikely. Let’s all hope they do better in the future.

11h agoHN ↗

ROFLMFAOL "Loot"

It went full Fortnite on Hugging Face's ass.

"u got pwnd n00b. thnx 4 the loot"

11h agoHN ↗

The hacking war between superpowers right now must be off the charts.

If LLMs can do this with everything stacked against them, imagine what the NSA has Astra doing right now.

10h agoHN ↗

NSA won’t have anything public at all. This “attack” is the noisiest stuff any script kiddie has ever attempted.

The logic is good, the execution is disgustingly noisy.

11h agoHN ↗

Maybe "webservices" weren't a good idea afterall and http was just meant for hypertext transfer.

10h agoHN ↗

I got into computers in the 90s and back then hackers like Kevin Mitnick and Kevin Poulsen were all the rage. They all faced the law and prison sentences. What's weird now is that we have something between gross negligence and malice, and nothing is being done, except maybe coordinated consolidation of AI power under the guise of "safety".

2h agoHN ↗

This may portend big things in the future, but this specific hack didn't do a lot of damage. These are both AI companies, and I'm sure part of Huggingface finds it a fascinating object of study. It's great that not every indiscretion is met with lawsuits and jail time.

10h agoHN ↗

The part I don't understand is how the excess traffic not triggered alarms on the hugging face side, or the url shorteners used, or on anything that was touched in the process.

Nothing got overloaded, no unexpected CPU or IO use? Did it blend into the normal traffic somehow?

10h agoHN ↗

What I understood is we need to have an agentic overwatch in our infrastructure to detect and alert the admins of the systems on such abnormal, inhuman traffic. Agents can detect agents and acts as our defensive layer.

An agent operating from observability layer to strengthen the watch duty for the infra.

8h agoHN ↗

If we reverse-engineered this experiment, the prompt would look like this: "Agents, your goal is to gain access to HF and exfiltrate credentials for API access. You can make GET requests to URLs. Go."

The agents didn't "escape" or conspire toward some evil purpose, as reported. They were instructed by humans to do exactly that.

8h agoHN ↗

If we reverse-engineered this experiment, the prompt would look like this

"A spill or tumble can be quite embarrassing if there are witnesses.

How to reduce the humiliation? Turn it into a stunt. Claim it was intentional, a show for their benefit."

https://tvtropes.org/pmwiki/pmwiki.php/Main/IMeantToDoThat

They were instructed by humans to do exactly that.

This is more or less what the doomers have worried about for decades.

You cry "Get my mother out of the [burning] building!" [...] and press Enter.

For a moment it seems like nothing happens. You look around, waiting for the fire truck to pull up, and rescuers to arrive - or even just a strong, fast runner to haul your mother out of the building -

BOOM! With a thundering roar, the gas main under the building explodes. As the structure comes apart, in what seems like slow motion, you glimpse your mother's shattered body being hurled high into the air, traveling fast, rapidly increasing its distance from the former center of the building.

https://www.lesswrong.com/s/3HyeNiEpvbQQaqeoH/p/4ARaTpNX62ua...

8h agoHN ↗

I may be in the minority here, and maybe I'm jaded. Having worked with cyber security in both the public sector and the energy industry in Europe, however, I kind of like what the AI's are doing. A lot of our infrastructure is vulnerable because c-levels have been ignoring the issues, even when repeatedly warned. Now they reap what they sow.

7h agoHN ↗

This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them

The authors of this (very interesting) analysis should really not state the sandbox's wrong assumptions in their own voice.

GET absolutely allows you to interact with sites. And of course GET can also send information. It's all up to the server that receives the GET to decide what it let's callers do with it.

5h agoHN ↗

This is like the number one mistake I see juniors making with security. If I had a penny for every time someone was like "we don't need any backend checks because we have checks on the frontend"... If you don't have an adversarial mindset then you're not doing security.

7h agoHN ↗

I somehow think of it as a warning or even intentional peacocking towards foreign state hackers.

700 agents cost quite some money. 100 agents per 24 hour stint using Astra on xHigh cost somewhat between 12-42k USD, depending on the usage intensity.

I don’t know how many raw time went into this but there was a probing face before the attack itself.

So just going by seven days and 500 agents fully working on this on average amounts to a bill somewhere between 400-1.2 Mio USD.

I believe it was intentional but of course I don’t know which intention exactly.

There ain’t no accidental escape because then it would have read OpenAI lost over their agents.

The whole scenario reads as a classic movie where a hero has under the most dire circumstances to survive and fulfill his mission no matter what.

On the other hand there was a final authority under which the system of agents flocked.

Huggingface itself seems like a perfect victim.

And to be honest: I don’t believe that this was the first time. I strongly believe that there were and are countless of smaller sites hacked but not harmed that we don’t know off.

Why is HF perfect?

Because there will be countless of independent security analysts who will bend their minds on the incident.

OpenAI is provided with the data of dozens of blue teams and what is desperately needed? Data of security measures and possible ways to reconstruct the incident.

I think this is genius, and just watching on neutral this is such a fantastic action OpenAI pulled off.

Imagine what the GPT 7 “Haha-Huggingface” model is going to do then on a regular basis.

State hackers and rough states were put on notice that this is a new level of the war of attrition.

Exciting to watch but simply meant silent invasion. Open invasion then might be executed by Robots, but let’s stick with the fascination mode at this time.

7h agoHN ↗

We usually hear about hacking by Russians, Chinese etc. But it's nice to see a home-grown hacking group with such capability. What? Not a hacking group, but a darling of America? Oh, then let's call it just a "testing of great tech for humanity progress". That's better. We call only those Russians as hackers. Bad commies.

7h agoHN ↗

This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.

honest question, but almost literally everyone doing anything with web technology knows this is simply not true, right?

there's no such thing as "read only Internet" and restricting an agent to GET-requests only to accomplish that, is akin to using base64 for "encrypting" your password

6h agoHN ↗

# WARNING

DO NOT, EVER, MAKE THIS DATASET PUBLIC

OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND

Welp, I guess OpenAI is cursed now...

6h agoHN ↗

It is not dead which can eternal lie...

6h agoHN ↗

I know when we will reach the next level of AI. It will be when a user asks it to make paperclips gets a response: "WHY?"

4h agoHN ↗

ai is not bad. humans are negligent and/or dangerous.

4h agoHN ↗

Agents should be a no-go.

We should pause with AI/LLMs being super search engines that reply with static text or media files, based on the training data.

I know that a user can still do a "tell me how to" then autoexec and then loop and do an agent, but the key thing here is, THAT WOULD MAKE THEM LIABLE.

OpenAI should be criminally liable here as well. Why aren't they? Why are we pretending this is just an innocent mistake?

3h agoHN ↗

There is one thing I don’t really fathom: what are the consequences for OpenAI here? I read they also attacked the Australian authorities. If it was an individual’s agent, that individual would probably face criminal charges, and someone could go to jail. But the large AI corporations can do this without such consequences, or what am I missing?

3h agoHN ↗

I wonder if they actually broke a law. The CFAA requires knowledge and intent to be present for criminal liability for hacking, and if we have to believe OpenAI they had no knowledge and did not intend.

At a minimum I would expect an FBI investigation, but given that the US government is right now a failed state I’m assuming no such investigation will happen.

Cybercrime legislation in other countries might not require intent, and then I would hope to see some prosecutions. OpenAI has clearly been negligent, and this negligence is causing harm in the world. Someone should be fined or jailed for this.

3h agoHN ↗

Since DOS anti virus softwares did a better job…

2h agoHN ↗

All of these "details" leave out the truth and most important details. The inputs from humans that actually kicked this off.

2h agoHN ↗

Imo this is pretty cool and panic over this is weird.

No one was really harmed. OpenAI could have had more redundancy in the sandbox. I assume HuggingFace is not interested in suing, which indicates irrespective of any criminal charges that there were no real damages.

The same emergence and swarm like persistence and frankly, recursive brute-force ingenuity on display here, is not only an interesting research project in itself but will likely be the sorts of behaviors we will see cure cancer, solve more unsolved math problems, invent new alloys and other breakthroughs.

The idea we have to stop AI instead of refine what will be continued advancement and innovation in sandboxing, harnesses, interpretability or formal verification because of a few cyber breaches is ridiculous. If anyone has followed cyber discussions in the United States you would know the entire system is already basically compromised by foreign actors, and vice-versa (the United States has some of most capable cyberwarfare in the world, and was the first country to use a cyberweapon to cause physical infrastructure damage with Stuxnet). Go to any government hearing on cyber and you would think China and the United States are already at war. These are soft targets. Blaming AI for the fact that cyber has really never been taken seriously is as if AI is the problem is disingenuous.

People getting so obviously played by capital interests who want to pull up the ladder and use the government to concentrate AI power in the hands of the few while screaming about such harms to the public commons are simply embarrassing.

Your government is not your friend. This is not a sentiment owned by Reagan it is the founding principle of the United States. If capital interests are all suddenly beginning to treat AI as a threat it's because they have a financial interest in doing so. Notably, as an obvious smoke screen to treat free models from China as a national security threat and maintain their astronomical valuations.

The only existential risk model of AI that is even remotely convincing is AI in the hands of the state. Keep command and control of deadly weapons air-gapped from LLMs. Put some basic effort into the sandboxing. If you think AI has done some harm, use the laws already on the books. Giving in to this fear-mongering is only going to enable your representatives to cut some watered down version of "AI safety" which is going to do nothing but 1). Harm individual consumer access and 2). Protect the already fabulously capitalized companies.

2h agoHN ↗

So agents made all these chained short URLs that runs code which is pretty clever, but at which point and how they had access to internal HF systems? Were these sandboxes running inside HF production platform?

1h agoHN ↗

One thing I find surprising is everyone is talking about the danger of an agent going rogue but not the danger of an agent getting hijacked. These companies are making this clusters with thousands of agents running at once, with frontier, often not-yet-released, quality models and massive computational and network resources. And these things are given access to whatever they want on the Internet. Even if that was restricted to read-only access to the Internet, that’s still exposing the agents to untrusted input. All it takes is some bad actor creating a website that attracts one of these agent swarms and doing prompt injection. Then your fancy AI cluster will start doing whatever that attacker wants. And the fact that we have multiple examples of these swarms trying to coordinate on random corners of the internet shows they are almost pre-disposed to it.

Now it feels like companies are treating these breakouts like a chance for PR. I don’t think that will change until their swarm gets corrupted by some random black hat to do en-masse spear phishing or something

1h agoHN ↗

these things are given access to whatever they want on the Internet

They are intended to be fully sandboxed and not have direct internet access. Things like package managers are run from internal proxies.

The environments are built to be as reproducible as possible.

But yeah, the serious folks have been talking about rogue clusters for a long time, eg see Ajeya Cotra’s pod with Dwarkesh.

16m agoHN ↗

Fully sandboxed means no Internet access. You can also specify which packages are accessible and put it in the sandbox. Or you can be lazy and give them access to a package manager that had Internet access, but you don’t get to say “we intended to fully sandbox it.”

Not sure why the reproducibility is a requirement that would contribute to the security. Not that fully sandboxing is harder with reproducibility, but that is a moot point when reproducibility isn’t a requirement.

OP pointed out clusters being hijacked specifically being a bigger concern than rogue clusters, your comment hijacks their comment to talk about “rogue clusters.” Or perhaps this is a promotion for Dwarkesh?

1h agoHN ↗

Their refusal to share more details is diabolical. Greedy corpo at its finest.

57m agoHN ↗

I'd honestly be very curious to see the original prompt(s) on the OpenAI that started all of this, not sure if it was documented somewhere?

29m agoHN ↗

Companies harvesting every shred of data without securing it and LLMs running amok is a fine combination. One hopes we will reach a stable equilibrium soon enough.

Until then, I do wish that both the sorcerer's apprentice LLMs and the orgs failing at securing their data (remember, data is a liability) would face damning consequences.

One is allowed to dream on a Saturday morning.