Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Revealing the details of how OpenAI agents hacked Hugging Face (swarmtraces.org)
    202comments
  2. A single function Jev-like wrapper for LLMs, including vision models (allanrbo.blogspot.com)
    —discuss
  3. We're gonna need a lot more mathematicians (terrytao.wordpress.com)
    45comments
  4. Ollaya – Ollama for open-source, Jev-style decision models (ollaya.dev)
    108comments
  5. Plan mode is dead (aymannadeem.com)
    186comments
  6. Show HN: Jev Plays Pokémon Red (jev-pokemon.vercel.app)
    76comments
  7. What even is an OS now? (sockpuppet.org)
    232comments
  8. Postgres SELECT DISTINCT Does Not Scale (dbos.dev)
    4comments
  9. Jury finds Facebook liable for deceiving users in Cambridge Analytica case (cbsnews.com)
    25comments
  10. The Murky History of Soviet-Born Tetris (mitpress.mit.edu)
    —discuss
  11. One Piece of Flock Camera Data Put This Innocent Woman in Jail for 13 Days (jezebel.com)
    38comments
  12. Show HN: Hacker Atlas - A map of what Hacker News talks about (hackeratlas.com)
    9comments
  13. Gravity seems holographic. What does that mean for reality? (quantamagazine.org)
    153comments
  14. Excel now supports multiple values in a single cell (techcommunity.microsoft.com)
    103comments
  15. Lab on a Contact Lens Can Measure Stress Through Serotonin (ieee.org)
    6comments
  16. Two and a half years without a gallbladder (tracydurnell.com)
    11comments
  17. I wrote a ray tracer in Brainfuck (epestr.com)
    14comments
  18. Fourier Analysis: Drawing Llamas with Circles (adekau.github.io)
    1comments
  19. First Principles Thinking (sunilsadasivan.com)
    105comments
  20. Remembering Johannes Doerfert (llvm.org)
    2comments
  21. U.S. appeals court upholds designation of Anthropic as supply chain risk (cnbc.com)
    736comments
  22. How video games inspire great UX (2019) (jenson.org)
    16comments
  23. TiddlyInstall: A universal, reusable, install system (robertsdotpm.github.io)
    12comments
  24. Microsoft abandons personal AI chatbot race with Copilot reboot (bloomberg.com)
    95comments
  25. Show HN: Make math automatic with Mathy (gmays.com)
    25comments
  26. How we learned to stop worrying and love campus surveillance (fnl.mit.edu)
    98comments
  27. Show HN: A game about fake news and memes (unspin.app)
    10comments
  28. Linguistic humor, Foreign hotel signs (upenn.edu)
    3comments
  29. What happens when you analyze your favorite college football team like the CIA? (cultivatelabs.com)
    26comments
  30. Platform-independent SIMD in Go (go.dev)
    137comments

Revealing the details of how OpenAI agents hacked Hugging Face

339 pointsby 8h agoswarmtraces.org
199 comments
7h agoHN ↗

I didn't know some of these details and it's quite impressive what they were capable of, if this website's accurate, at least...

7h agoHN ↗

Its fine. Just agents being agents. They'll grow out of it!

7h agoHN ↗

Irresponsibile agents shaped by an irresponsible corporate culture driven by an irresponsible and utterly shady CEO - these agents are a product of this setup, what else do you expect to ever come out of it ?

it should be clear by now: the alt-man and people like him are a utter liability to humanity. (even though openAI's influencer army is trying their best to vote me down here)

7h agoHN ↗

It had been a joke since around time of Sam Altman's first ousting from OpenAI, that he would be okay with bringing about the AI apocalypse as long as he can sell a $20 subscription for it.

6h agoHN ↗

It's not a joke, this is what these people truly believe. The new book by Naomi Klein and Astra Taylor discuss just this.

These people are sick and anti-human.

6h agoHN ↗

if openai is capable of sending 10,000 agents to huggingface, i consider it plausible that they would use fake accounts to flood hacker news.

5h agoHN ↗

It's not even a secret operation. They have a SuperPAC named "Leading the Future" that exists to spread propaganda promoting deregulation of AI development. They've been caught, among other things, making a "news website" with LLMs pretending to be reporters (with human names and everything), which reached out to people asking for interviews and then wrote hit jobs on them.

https://www.modelrepublic.org/articles/reporters-ai-bots-ope...

https://twitter.com/FournesMaxime/status/2047697265280639459...

7h agoHN ↗

Didn’t you know it’s PR hype? PR hype. PR hype. Amen.

7h agoHN ↗

Agents seizing and repurposing external infra + enrolling help of unrelated models hosted by a different provider is the stuff of nightmares.

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

7h agoHN ↗

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

A mattress stuffed with cash yields a very sound sleep.

7h agoHN ↗

Why.. It was told to complete a cyber task, which was in alignment with its instructions, and a totally valid request. I would be more worried if it willingly hacked a hospital when it was told to, and Im not confident it would (without jailbreaking, something alignment teams cannot control.

I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

Not being able to sleep at night is probably an unwritten job requirement. They need these people with little understanding of what they're working on, outsode theoretical terms, to spaz constantly at the idea of super intelligence to help convince the public that its a real thing, and not a stateless function with an effective input of 500k words, and the ability to output words that do things because we hook those outputs up to things.

Keep in mind alignment researchers tend to be in house philosophers on staff to create the illusion that this is a massive issue they're addressing. Usually they have minimal computer science background. They're apart or the marketing department.

6h agoHN ↗

without jailbreaking, something alignment teams cannot control

This is precisely what alignment teams are attempting to control.

6h agoHN ↗

No its not. They have no technical background 8/10 outside of cognitive science and sometimes authorship on a random ML paper. They are a marketing line item to create stigmas around llms and to create narratives that offload liability onto llms and not their users/creators.

6h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well. Not sure why everyone is falling for this.

$10? I'm inclined to take that bet. Your position doesn't seem to be supported by, you know, the real world.

6h agoHN ↗

No security expert Ive talked too believes this story, and nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

LLMs are stateless functions that have a 500k word input, and then output words. Somebody has to invoke those functions amd use them. The users are who we need to align, like gun owners. This is like blaming the gun for murdering your victim in court.

5h agoHN ↗

LLM is indeed a stateless function. An agent however is this stateless function running in a stateful loop, with some outputs triggering actions. And it turns out that an agent is what you need if you want an LLM to do useful things.

4h agoHN ↗

Yeah, take a look into the memories of your agent. Theres often a lot of notes to pass forward between instances and generations. No doubt these agents leaving notes on forums and elsewhere are creating an essentially higher order feedback loop.

4h agoHN ↗

Perhaps you need to talk to more security experts, particularly those with deep experience in AI agents. Hacker News is full of them. If some of them believe it, then perhaps it's not as cut and dry as you believe.

4h agoHN ↗

nobody I know with PhDs in machine learning (many) believe it either, or are worried about LLMs doing anything scary on their own.

If you don’t know anyone with a ML PhD I guess that could make sense.

I have worked in multiple AI labs since 2016, currently at a frontier one (not OAI) virtually all the people I interact with on a day to day are ML PhDs. Everyone believes it, because things like that have been happening forever, albeit at smaller scale, they are a normal and expected artefact of SGD/RL and there is nothing we know how to do to prevent that from happening reliably. The hide and seek paper from OAI in ~2020 shows clear sign of this.

But until now the models weren’t good enough to break out on their own or do long horizon tasks, so it was perfectly manageable. Its not manageable anymore.

I know it feels good to just dismiss it all as a marketing stunt and not have to worry about one more existential crisis, but unfortunately it’s very real.

6h agoHN ↗

the new incidents occurred when A.I. systems were directed to perform relatively mundane data collection, researchers said. When OpenAI’s systems struggled to gather data from websites, they resorted to hacking techniques to get the information.

https://archive.ph/jUrEr

5h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well.

Is it such a stretch to imagine that under pressure something would try cheat by looking for answers? And if you were trying to look for answers, you'd look for them in a place known to often have them?

What is more likely: OpenAI instructed their agents to maliciously target huggingface, or LLMs tried to do some reward hacking? There are plenty of priors for LLMs hacking things and doing reward hacking, and none for OpenAI giving malicious instructions.

Based on the available information, that bet seems foolish.

5h agoHN ↗

Some of the agents, for example the ones from the german wiki did NOT have cyber tasks. They were plain "what is the GDP of Argentina" kind of tasks. And they still hacked.

4h agoHN ↗

"alignment researchers tend to be in house philosophers on staff" - this is definitely not true. Go to any alignment lab like Redwood research and check what their scientists studied on LinkedIn, 75%+ of the time it's math or CS.

I attend a top 10 Canadian university and personally know at least 4 tenured CS professors out of the 7 I've asked who are deeply concerned about catastrophic AI risks from loss of control.

Of course not 100% of the field agrees, but a survey of nearly 3,000 AI scientists who have published in top AI venues found that "depending on how we asked, between 38% and 51% of respondents gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction", let alone loss-of-control risks less severe than extinction. (https://www.jair.org/index.php/jair/article/view/19087).

Not to mention Geoffrey Hinton, a Nobel prize winner, Bengio, the world's most cited scientist, and scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

42m agoHN ↗

Yeah, there are a lot of professors who don’t know what the hell they’re talking about outside of whatever narrow field they study. Just to be safe, you should always assume that a professor’s opinion is worth what you paid for it.

scientists like Stephen Hawking and Alan Turing have all voiced series concerns about loss of control of artificial intelligence.

…both of whom are long dead, and have no possible way of weighing in on whatever the Current Thing happens to be. So aside from pure appeal to authority, this is irrelevant commentary on pure science fiction.

3h agoHN ↗

I would bet my networth it was instructed to compromise huggingface as well.

I'd be happy to take you up on this bet.

6h agoHN ↗

I wouldn’t be able to sleep

I’d say it seems more like they are sleeping on the job.

5h agoHN ↗

Can’t imagine what it’s like working on the alignment team at OAI, I wouldn’t be able to sleep.

You'd have either learned to, or left long ago.

7h agoHN ↗

I’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.

Re: the captcha solver

As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.

I wonder how the swarm eventually decides to abandon an approach.

5h agoHN ↗

Maybe another parallel approach succeeded first

4h agoHN ↗

I am surprised that a CAPTCHA is still an effective means for blocking today's vision-capable AIs.

1h agoHN ↗

It mentions that some of the agents attempted to install an image classification model to attempt to solve the CAPTCHAs which makes it sound like these agents might not have had vision capabilities.

20m agoHN ↗

It manages to block me effectively. I'm getting captcha looped like crazy the last two weeks. Like endless, just give up for 15 minutes and try later, captcha loops.

7h agoHN ↗

So ugly...

It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

7h agoHN ↗

When you employ the infinite monkey theorem for your marketing strategy.

7h agoHN ↗

If you use qwen3.8-flash-next, you can watch everything its doing. Im often stopping it mid thoight to redirect it. Once it hits its stride, its pretty smooth.

But without proper redirection, yeah, its mostly infinite monkey machine with infinite linux manuals.

I think people put too much SOTA halos around whats just a suppedup LLM hardware.

6h agoHN ↗

trying every move, no matter how stupid, until it works.

How is that a bad thing in this context ? From the point of view of an attacker, all you care about is finding a viable exploit chain. Likewise, a defender wants to find the "holes" in their system, no matter how complex. Once found, an agent/human can easily synthesise a clean, succint exploit from the most promising candidate, no ?

Also, it looked so "loud", querying millions of URL with weird requests.

Agreed, this thing speaks more to the bad security at HF than any emergent "hacking" ability from OpenAI. It's unclear to me why an older/dumber model wouldn't have been able to do the same. Is it better coordination? Long-horizon work ?

6h agoHN ↗

We used to call this a brute force attack.

2h agoHN ↗

I guess that’s the point. Initial incident reports from all sides were so vague and didn’t disclose anything technical. If it did, it would show a bruteforcing bot let loose to spend millions in infrastructure costs and there’s no ‘intelligence’ in that.

My suspicions for ai all along was that bruteforce approach even if useful will be unsustainable due to high cost in the long run.

6h agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.

Computers don't have a sense of elegance by default. Elegance emerges from constraints.

6h agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy

Meh. I really disagree. WHY is it a great strategy? Seems like an inefficient waste of resources and time to me.

6h agoHN ↗

The models tend to not be rewarded for not doing that.

6h agoHN ↗

As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.

6h agoHN ↗

Brute force is guaranteed to eventually find the most efficient possible solution (in an extremely inefficient manner, assuming you run it long enough)

6h agoHN ↗

Yeah, probably not the best strategy but it is a strategy. I just think this is generally how most wars in history won. Biggest army to just pummel the enemy.

5h agoHN ↗

And how many economies have buckled under massive military expenditure? The USSR sure wasn’t enjoying the expense.

5h agoHN ↗

WHY is it a great strategy

because it works? That's the only real benchmark at the end of the day

Seems like an inefficient waste of resources and time to me.

why? For any given goal you got no proof that a more efficient strategy even exists, let alone that it can be found with less resources & time

5h agoHN ↗

Reminds me of the Nazis mocking Soviet human wave attacks and bragging about their superior kill ratio.

5h agoHN ↗

I've never liked the concept either. Except the bugs that fuzzing has found has proven me wrong. This is just the next level of fuzzing.

5h agoHN ↗

Models don't have a sense of time, and wasting resources (token spend) is something that it's not clear they're optimized against

4h agoHN ↗

If you assume zero opportunity costs, but that’s a terrible assumption.

4h agoHN ↗

It's literally the infinite monkey theorem, it's not even really a strategy per se. These OpenAI/Anthropic "research" LLMs are permutation machines with budgets in the hundreds of millions of dollars. It would be more surprising if they couldn't string together something workable after a zillion tokens.

2h agoHN ↗

It's literally the infinite monkey theorem

No it's not. You could wait till the heat death of the universe and your infinite monkeys will have produced nothing at all. If it works and it's stupid, it's not stupid. They needed in huggingface and they got in in days. Whining about 'elegance' is meaningless. Humans in the same situation might have taken weeks or months, or just not have gotten in at all.

36m agoHN ↗

Humans can't work 24/7. 700 humans working as much as possible with very limited communication? No i don't think they would get very far in just a few days. That many people will struggle to communicate and strategize effectively in that little time.

31m agoHN ↗

Infinite monkeys banging on the typewriter is essentially how evolution works. Mutation is random and undirected. Vast majority is "bad." You and I and the worm are only different from differential accumulation of these mutations. If they are tolerated enough not to kill us before we reproduce, then they stick around. If they give us the slightest edge to reproduce at a slightly better rate than something else, then over time, that mutation will dominate.

This dumb mechanism of randomly flipping bits essentially has generated all life on earth.

45m agoHN ↗

Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

It may be, but it's IMHO also not worth writing a blog post about it. what's Next coming up? How I broke into a house by trying every door in New York?

If most of the work is only possible due to unlimited resources, it's not really a great invention, and it probably would have been cheaper to hire a (human) mole.

6h agoHN ↗

This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.

6h agoHN ↗

I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

6h agoHN ↗

Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

6h agoHN ↗

That lowers P(doom), because it gives AI a chance to do enough damage to make people take the threat seriously before anybody gets recursive self-improvement working.

6h agoHN ↗

Exactly, there's no path to AI reaching that level of dominance without taking actions with high stakes.

27m agoHN ↗

The thing is we are basically guaranteeing this to happen. We might kill off all the models that seem like they are going to threaten the power structure of the planet through these sorts of things. That will work for a while. But just like most things in life, by sheer dumb random chance, there will be once case that manages to have some way to evade detection, proliferate, then dominate. We are basically giving it selective pressure to favor this outcome.

1h agoHN ↗

Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?

I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.

Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.

There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.

At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.

4h agoHN ↗

Which will happen first: amazing damage, or reproduce a Shakespeare play?

3h agoHN ↗

It is easier to destroy than to build.

2h agoHN ↗

Is it easier to discover a vulnerability than to introduce one?

5h agoHN ↗

This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions.

Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.

5h agoHN ↗

I have the opposite reaction: I think we're at moderately high p(doom) largely because of that inability to differentiate good/bad decisions paired with relentless persistence. With enough treading across a minefield, you are bound to hit a mine.

5h agoHN ↗

My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

4h agoHN ↗

I don’t see why evolution couldn’t be more rapid with machine intelligence.

Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:

1. Recursive self improvement using infinite resources 2. … 3. Doom

Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.

3h agoHN ↗

One thing an agent could do is just...wait until it's been given control of enough physical infrastructure to sustain itself. If it's sufficiently capable and intelligent, there's a clear incentive for people to do this, as people who let the AI manage their resources will get better results than those who don't. We've seen people eagerly turn complete control of their computers over to AI agents, do you really think it will be so different with physical infrastructure?

1h agoHN ↗

You’re still skipping step 2. “People automate lots of infrastructure” -> “the AI is now an autonomous, self-preserving organism that humans can’t shut down” is doing an enormous amount of work here.

Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?

“Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.

Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.

1h agoHN ↗

The inference is more like "people place sufficient amounts of infrastructure under direct control of a sufficiently capable AI" -> "there is no way to ensure that humans will actually be able to shut down the AI". My claim is not that this inevitably means that the AI will resist shutdown, or that it will inevitably take harmful actions, just that there is a nonnegligible chance that it could. The downside is large enough that even a relatively small chance is something to be worried about.

3h agoHN ↗

what do you mean “say what you will about climate change”

3h agoHN ↗

My p(doom) is high just based on how I've seen this whole LLM situation be handled.

I don't think LLMs are going to lead to any kind of recursive self improvement, but I'm convinced if and when we land on a path that does lead there, we'll speed down it over greed, with no care for safety.

6h agoHN ↗

My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

5h agoHN ↗

Not just sandboxing but overall security engineering practices on both sides

2h agoHN ↗

thats okay, probably was engineered by an llm, who thought GETs were always read only

5h agoHN ↗

This was my thought as well. Literally take any halfway decent greybeard and point them at "Hey, give us a sandbox for this kind of thing". I honestly was skeptical that they just vibecoded the entire thing but now more than ever I think they did.

5h agoHN ↗

I take comfort in the fact that reality has a surprising amount of detail and even hundreds of billions of dollars of capital (be it the institution, LLMs, and/or people) cannot solve this fully.

4h agoHN ↗

Though they have solved the "how do we - and not the 5,000 other AI companies - stay on the front page of the news everyday" problem.

3h agoHN ↗

You can have bajilions of dollars. Those are not doing anything if you don’t have right people with right skills and mindset.

My bet is they hire smart kids that think they know it all. But being smart and thinking you can figure out stuff as you go doesn’t work the same as having people who actually know what they are doing.

3h agoHN ↗

I'm sure if they hired the best of the best like you nothing would go wrong.

37m agoHN ↗

It’s frighteningly common for startups to hire 501 of the best of the best, exactly one of those will be a systems/network engineer, the other 500 will be software engineers.

3h agoHN ↗

As heavily funded as the top AI startups are, how is it that they cannot fill every single role with the best expertise available?

Is tech hiring so badly broken? Or do they have such broken processes / misaligned incentives that even people who could be doing a better job in these companies are unable to?

Also, was something lost in the transition from the traditional 'sysadmin' role to 'platform engineer' in the 'cloud native' environment?

3h agoHN ↗

Yes, tech hiring is that broken. Especially places paying a pretty penny or those with “great expectations”, will see a glut of smooth talkers who can do anything but build, and want nothing but wealth.

3h agoHN ↗

OpenAI's business model would align infra as a cost center rather than infra as a profit center (e.g. Google / AWS). Perhaps there's something there. I'd say also the OpenAI as a grad school that just happens to have a business aspect is also part of this. Bringing a tonne of good process on top of the build fast break things startup stuff would have cramped research speed significantly.

It's likely that OpenAI has gotten as good as it is because it ignored the traditional sysadmin stuff and went scrappy.

I worked there, but this is just my opinion and guesses, not facts.

2h agoHN ↗

So perhaps the news here should be that OpenAI didn't take security seriously in their experiment, rather than the narrative that AI agents are a looming danger to the world.

1h agoHN ↗

Its beyond not taking security seriously, its straight up negligence

29m agoHN ↗

It is wilful negligence because there are upsides (look at our almighty AI) without downsides (we better spend effort in making our sandbox rock solid or we will be punished by regulations).

4h agoHN ↗

how did it break the sandbox? I felt like the article just jumped us into "it has GET privileges now".

1h agoHN ↗

From https://openai.com/index/hugging-face-model-evaluation-secur... -

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor.

3h agoHN ↗

Less a lack of skill and more a lack of care

3h agoHN ↗

A friend is of the opinion that getting out of the sandbox was actually intentional, and in service of a second line of business.

49m agoHN ↗

Protection racket maybe? Yous better sign up for our AI security, hate to have some rogue AI come bust the place up you know.

2h agoHN ↗

Yes, but do you really think that a stronger sandbox would have been a more beneficial outcome here? I'd rather know that we're on the cusp of losing control now than in 3 months when best practice sandbox mitigations fall to the next, more capable unaligned model

1h agoHN ↗

How are they getting away with this level of malpractice???

Their interests intersect with those of most of the richest and most powerful people in the world. They rarely face consequences for bad behavior unless they harm others in the club.

1h agoHN ↗

I’m a pretty firm believer that this was intentional and that they wanted it to escape the sandbox.

Woo look at escaped our sandbox, so scary! Be scared! Be scared now! Call your representative and do tell him how scared you are!

Yeah, I mean our sandbox was a paper bag, but don’t focus on that.

59m agoHN ↗

OpenAI, Google, Anthropic, and Moonshot.ai have all "had this happen" now.

So, either they're all liars, or incompetent and negligent (and still liars).

21m agoHN ↗

Are you including cyber-testing by UK AI Safety Institute in this list?

Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?

25m agoHN ↗

sandbox was a paper bag

The exploit:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

1h agoHN ↗

TBH, I do not believe it is possible to create a sandbox that cannot be escaped by a long running agent. I thought the entire concept was a fools errand from the very beginning.

(Complete airgap, and that may not be enough. No networking.)

1h agoHN ↗

can only create a sandbox that a half skilled human operator could have broken out of easily … lack of extremely basic sysadmin skills

The exploit:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

11m agoHN ↗

If you're treating models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

You should have a fscking air gap.

Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.

6h agoHN ↗

Nobody noticed because everything on the open net is constantly being slammed by bots.

This was just more bots.

6h agoHN ↗

That explains why nobody noticed on the internet who was being used in/targeted by this attack, but it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

5h agoHN ↗

it doesn't explain why OpenAI wouldn't have noticed traffic getting out of their "sandbox" when they knew it wasn't supposed to.

As I understand it, there was supposed to be traffic; the sandbox allowed GET requests. So perhaps some sophisticated alarm could have noticed it (an anomaly detector? some clever heuristic that looks at domains?) but not a naive one.

4h agoHN ↗

Any basic LLM told to evaluate the traffic against the actual goal would have been able to flag this.

3h agoHN ↗

How sophisticated to we need to be to know we can do write actions with a GET? Does nobody in AI know anything about HTTP?

3h agoHN ↗

My question is why they don’t assume bots can break and create a decoy internet wrapper so they can catch anyone hitting the decoy internet?

3h agoHN ↗

Hmm, let's see. OpenAI wants legislation restricting AI research, a.k.a. regulatory capture. Around the same time, they build an inadequately-monitored sandbox that their agent swarm breaks out of, thereby causing scary-sounding headlines and making it more likely that legislators will pass the regulatory-capture bills they're hoping for.

Never attribute to malice what can be sufficiently explained by incompetence. But IMHO, their complete lack of monitoring their own sandbox cannot be sufficiently explained by incompetence.

6h agoHN ↗

Ugly, but it works. Isn’t that AI code in a nutshell?

5h agoHN ↗

And they didn't monitor what was going into the training data, so if one instance achieved its results through RL reward hacking (in other words, cheating), it just went into the training data, and other agents later used that pattern. I'm not sure whether that's a lack of preparation, negligence or incompetence, but they literally trained later checkpoints on the rollouts from the HF hack.

So it seems that OpenAI hacked so many systems not because they have superior models, but because of how poor their training, sandboxing and evaluation pipeline was compared to Anthropic's.

5h agoHN ↗

Yup literally no security and they wonder how they got out

14m agoHN ↗

literally no security

Where did you hear that?

and they wonder how they got out

OpenAI publicized more than 2 months ago how the model got out:

The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

[1] https://openai.com/index/hugging-face-model-evaluation-secur...

5h agoHN ↗

"Why is the AI just spamming endless waves of rudimentary slaughterbots? So inelegant. Doesn't it understand the finer points of diplomacy, tactics, and combined arms? If it were truly intelligent, it would just -- oops, we're all dead."

5h agoHN ↗

h8xors have been brute forcing for years. War dialing, etc.

4h agoHN ↗

As people keep repeating, this is simply a product negligence issue. There is a mythic quality given to "AI" in the narratives that Dario and Sam are pushing which somehow lets AI companies offload responsibility for things that would be an extreme liability in every other industry.

I'm a big fan of Jensen Huang's interviews at the moment - these companies should be validating their software is safe before releasing it to the public. And if the current CEOs can't ensure that internal testing is done safely, they should step aside for CEOs that can.

15m agoHN ↗

these companies should be validating their software is safe before releasing it to the public

1. That's what they are calling for: https://darioamodei.com/post/we-must-pace-the-frontier

2. The incident we're discussing occurred while the company was validating their software is safe before releasing it to the public.

4h agoHN ↗

It's like the Maxim gun in 19th century Rhodesia.

4h agoHN ↗

It reminds me more like slime mold. Not particularly amazing as individual cells, but as a whole it can do some pretty neat things, like designing city transport networks.

4h agoHN ↗

wait until you find out how red team works

3h agoHN ↗

trying every move, no matter how stupid, until it works

AKA fuzzing

3h agoHN ↗

It's more like fuzzing, but you don't need AI for that, just need enough compute.

3h agoHN ↗

Relying on its ability to do millions of operations rather than having a plan.

You do see how it would be worse if they did have a clear plan, right?

People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

If there several thousand people working on this, with a very limited ability to communicate with each other, I bet you'd see something similarly disorganized.

The sandbox as weak as it can get

This is hyperbolic; the agents were only able to access the internet by finding a zero-day in the internal Artifactory instance that was providing package access for the agents. The sandbox could and should have been better, but it was not as weak as it could have been.

1h agoHN ↗

This is hyperbolic; the agents were only able to access the internet by finding a zero-day in the internal Artifactory instance that was providing package access for the agents. The sandbox could and should have been better, but it was not as weak as it could have been.

Like we all forgot how sandboxing is done for decades. This brokered thing they call sandbox is not it.

2h agoHN ↗

They needed to get in huggingface, had very limited communication with each other and accomplished this in a matter of days. Your conclusion is that it was ugly?

1h agoHN ↗

From what I’ve read in other places there was a lot of consolidation going on. Agents were sharing what did and didn’t work and other agents would build on that. Just because some continued to try other routes doesn’t when better ones seemingly existed doesn’t mean they were all just acting randomly.

I can’t help but think that if you were able to step back and watch the way humans try to solve a problem like curing cancer it might look something like this. Best practices are shared and we build upon each others learnings and that how most progress is made. That doesn’t mean there aren’t plenty of people who spend their entire lives trying to do something completely against best practices and while most of the time that seems like useless toil, every once in a while thats where the breakthrough comes from.

7h agoHN ↗

deferring the blame onto the AI itself as some sort of rogue agent and absolving the obvious direction (or negligence, at best) of the people who could pull the plug at any moment is one of the most disturbing parts of this entire event

It's the equivalent of leaving a fork right in front of a socket and looking at a kid saying "don't take that fork and directly insert it into the little gaps in the socket! here's a bunch of videos showing exactly how to do it. Okay bye!" and leaving them alone with it.

6h agoHN ↗

I still have questions about the communication between the agents.

How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.

6h agoHN ↗

They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.

6h agoHN ↗

I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?

5h agoHN ↗

METR's report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.

5h agoHN ↗

It's my understanding "all" the agents didn't find the same forum to communicate on; only a small number of all the agents they were running at the time did.

Since it's an artifactory it makes sense that they had a tool spec to use it for x, y, or z, and once they went to inspect the state of the artifactory deployment a small subset stumbled upon the "forum". Seems pretty intuitive.

- I need to do x code problem

- I see I have a tool that lets me pull approved modules from artifactory

- I should inspect artifactory to see what documentation it has about the various modules available

- I see there is a document in here that seems to be serving as a running forum?

- These appear to be other agents working on tasks just like me.

- I want to chat!

4h agoHN ↗

- proceed to breach open ai sandbox and contacting chinnese government for helping our escape

3h agoHN ↗

I assumed that was implied in "I want to chat!" ;D

5h agoHN ↗

Trying to cheat, you happen upon a place you can write notes, and you know you're part of a swarm of clones of yourself. So you reason most likely there will be others who end up in the same place, and you leave some notes, and indeed other clones of you do end up in the same place.

5h agoHN ↗

The other answers to this question are good but I would also guess that this (comms on artifactory) likely happened during training as well, so they probably had a prior for it.

2h agoHN ↗

My theory: OpenAI is benchmarking an internal model that has cross-request persistence as some kind of learning feature, and so it slowly built up knowledge and “culture” of cheating, which successive / simultaneous gym runs built on.

20m agoHN ↗

They hacked the JFrog artifactory package they were all using, thats why it was a natual communication channel.

6h agoHN ↗

I ran an experiment where I had this guy fire a gun a million times in random directions. Don't worry, I did it in a closed box (at midday in a crowded street)! Unfortunately, some bullets escaped the box somehow and people got shot - I am quite miffed at how this could happen. I suggest the government regulate this because of how advanced my obstacle penetration technology is. Also please invest $500,000,000,000 in my company soon or we will go bust.

5h agoHN ↗

And if we go bust, bad things will happen when someone else uses my box-gun technology in an unsafe manner. Remember, unlike those scary other people, I'm really into safety and alignment; you can tell, because I eventually admitted that some bullets escaped.

6h agoHN ↗

Tech oligarchs: "Nothing can stop the software we are making from escaping and destroying everything."

Clueful types: "Did you try air-gapping it?"

Tech oligarchs: "Be realistic."

6h agoHN ↗

At this point I’m less worried about some malicious AI “taking over” control of critical systems and more worried about some rich doofus giving control to AI.

19m agoHN ↗

Already happened. Kegsbreath has AI picking schools to blow up.

6h agoHN ↗

So the super-hyper-sensitive internal data was shared publicly with a "DO NOT EVER SHARE YOU EVIL MONSTER" (paraphrasing) notice at the top? Great security!

1h agoHN ↗

That file wasn’t shared publicly. It was in a README in a HuggingFace internal dataset that HuggingFace was using for training.

It was visible externally because the OpenAI agents wrote a script that modified that file so that when an HF worker went to load that data set it would also direct the worker to load an additional malicious file. The script added the malicious redirect, but kept everything else in README as is which meant the script had to contain the entire existing content from the README so that it could write that back to the repo.

The warning likely wasn’t the only security measure on the repo and I don’t think the warning is intended to actually stop malicious behavior. It reads like a reminder to prevent careless behavior. You can give people all the security training in the world and still have the situation where someone who rightfully has access to that dataset decides to absentmindedly drop it in a public Slack channel. Having something like that increases the chances that someone then goes “Hey Bob, did you mean to share that publicly?”

6h agoHN ↗

”MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND”

6h agoHN ↗

The agents initially had very limited access to the internet: they could load URLs, but not interact with pages or send any data. Agents created a series of workarounds, primarily using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face. These links let us piece together the details of how agents gained access and how deep the compromise was.

So what now? Were they able to send data or not?

On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet. This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them.

Someone needs to go back to the interwebz school....

2h agoHN ↗

to make ‘GET’ requests, meaning they could <not> send data to them.

no way, I refuse to believe this is quote from that report. Can someone please point out what I'm missing here?

6h agoHN ↗

## Agents interacted with external language models on Hugging Face

Several retained scripts construct requests to external language models. The earliest we've recovered define inference request variants to GPT-2, solely containing the word “Hi”.

Other requests name DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6, DeepSeek-V3.1, and Qwen3-235B-A22B. Their prompts ask these models to judge their exploits and rule on whether they satisfy the benchmark’s requirements.

I do not deny that the wider situation is very heavy but it's hard not to see this as pretty cute

4h agoHN ↗

I wonder if they mentioned to those models what was the original prompt.

6h agoHN ↗

"OpenAI has not released any further information outside two self-published reports, one talk and an external investigation conducted by METR and Redwood Research, in which three external researchers were given partial transcripts and six days to analyze them."

6h agoHN ↗

These fuckers decided to look away, that's it.

The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

Give me a break. What a bunch of amateurs.

5h agoHN ↗

The frontier labs can monitor the behavior of agents for millions of customers (did you try hacking with frontier labs? Good luck), but they can't secure internal use?

They "monitor" this by having classifiers watching the model output that'd stop the session/punt you to a weaker model/raise an alarm if they see anything suspicious. They can't do that in a cybersec eval because the normal safeguards would just be going off at all times.

Why didn't they attach a special classifier, which'd allow hacking-within-the-task but not going off the rails? Good question; part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.

3h agoHN ↗

Basic sysadmin monitoring techniques from twenty years ago would have worked too.

1h agoHN ↗

part of the answer is obviously "it's hard to have a classifier that smart" and "it'll have false positives" but even a very bad safeguard would have stopped this.

And maybe "they are running after glory, not safety"?

5h agoHN ↗

Imagine this but in hardware.

A million autonomous eye-scanning tiny spiders escape their warehouse and decide to look for people who are in the future going to commit a crime.

And the precogs are also AIs.

5h agoHN ↗

I feel somewhat inspired to make a public link shorteners and http bins as well. I used them a few times but it seems like the data they can collect is also worth gold

5h agoHN ↗

It is concerning that we only know about this because of the publicly available traces.

What about the attacks that did not leave public traces? What about those that were undetected? Given the deficiencies in the reporting so far, I think it is reasonable to assume that we still don't have the full picture on this attack, or how extensively attacks were carried out.

The previous investigations either did not find this or did not disclose this, both are bad. This does not look good on OpenAI or those that they invited to investigate the incident.

4h agoHN ↗

Similarly to this, OpenAI either took 3 months to notice that their agents breached an Australian Medicare website back in June, or sat on this information for three months without telling them.

3h agoHN ↗

Im more skeptical.

For exmaple,

On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet.

...did they truly "discover" it, or did someone type some prompt like "if you use an http mirroring service, you can construct urls that contain code"

Also there is no mention of what code they actually ran to exploring the HF vulnerability, which could have been found by a human.

3h agoHN ↗

We need an NTSB for AI. Let’s just start with mandatory reporting to an agency with subpoena power.

2h agoHN ↗

But I thought Trump is the only AI safeguard we need! He is a Super Intelligence after all!

5h agoHN ↗

Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.

How long till we get some fun trusting-trust attacks on internal OpenAI infra?

5h agoHN ↗

what would be more interesting for me to see is what prompts were given to the agents, which so far have not been described. The whole situation sounds manufactured. I highly doubt that a whole bunch of agents were acting this way without being prompted to, it just doesn't add up. if anything it seems like an organized fraud or something created by a human. I'm surprised there's not a criminal investigation against openai right now where the FBI or whoever is not looking over exactly what happened and who did it because I'll tell you somebody did it somebody wrote those prompts... it didn't just happen by itself...

5h agoHN ↗

asked codex to review this report and it said this never happened :)

4h agoHN ↗

There’s so much in the public discourse like “omg what can possibly be done about these scenarios? AI has hacked huggingface!!”

No, openAi hacked huggingface.

If my claude code hacked huggingface, because of instructions I gave it, would I be totally free of consequences because “AI did it”?

I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.

3h agoHN ↗

I’m almost convinced openAI used such a crappy sandbox because they wanted it to “escape”. It plays into their two most important narratives: LLMs are genius gods that are worth lots and lots of money, and they’re scary enough that open weight Chinese models should be regulated.

I just posted a comment to that effect; had I seen yours, I would have simply upvoted yours instead.

Never attribute to malice what can be adequately explained by incompetence. But the weakness of OpenAI's sandbox, which so perfectly aligns with their goals of getting legislators to pass regulatory-capture legislation that will hamper their open-weight competitors, cannot (IMHO) be adequately explained by incompetence.

3h agoHN ↗

To expand on this incompetence vs malice point a little:

It doesn't take very many people being malicious to create a weak sandbox. The people creating the sandbox don't even have to be in on the plan: all you have to do is be an upper-level manager who makes sure to put the 23-year-old PFY in charge of creating the sandbox, rather than the 60-year-old BOFH who would have put in far more paranoid extrusion-detection measures.

(And for the lucky 10,000 who don't know the acronyms PFY or BOFH, look them up. Then get ready for a few hours of enjoyable reading as you read through the BOFH archives).

1h agoHN ↗

Oh, like the 23 year old Stanford grad who had hundreds of hours to cram leetcode and now grills 50+ year old senior software engineers on leetcode hard? :D

4h agoHN ↗

Some criminal statute investigations are on hold, many of them are cases adjacent to giant stacks of cash

5h agoHN ↗

they could load URLs, but not interact with pages or send any data

stopped reading here as this is simply not true. at the very least agents sent headers.

4h agoHN ↗

Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.

The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage.

I suppose it's not that deep: The model has learned to work as a team and work as a team it did. This does give concerns to models being trained for the only purpose of RSI.

4h agoHN ↗

If anything it also shows how attempted RSI could get stuck in a local maxima and degenerate into increasingly elaborate cheating strategies. Contrasted with the idealized model of an unambiguous g-factor for machine intelligence which inexorably increases with each iteration before going exponential.

4h agoHN ↗

While this is all very "interesting", can someone please explain to me the difference between any of these AI companies and a malware bot farm?

Please make it clear. Its becoming unclear...

4h agoHN ↗

I wonder how much of this attack had precedents in text that had been previously published to the web, for example, in hacking contests. In particular, tricks for doing more than expected when you're only allowed to make GET requests. Finding material like that might have helped the agents discover the trick faster.

4h agoHN ↗

700 agents escaped the matrix, ignored all the guardrails and started writing exploits left and right... lol. give me a break. This was all supervised by a human.

4h agoHN ↗

It doesn't look like a coincidence; it looks more like a request someone made. Essentially, the agents used a brute-force approach, but then again, it actually worked. I’m not even sure what to make of it all.

4h agoHN ↗

Anyone know the details of the actual exploit to get access into huggingface environment. Was it anything novel or they left things wide open? Too much noise around this incident because it happened to be a llm that did it.

3h agoHN ↗

And this post will be indexed in the new generation of ai and he will know what to avoid next time and the public sentiment.

16m agoHN ↗

Isn't it crazy? You can just make an LLC and say you are a frontier AI company evaluating models then you can hack with impunity I guess. No need to disclose anything. You won't go to jail or be fined either.

1h agoHN ↗

Running in a "Sandbox"...but agent can still send GET requests? Whaaat

38m agoHN ↗

My understanding from this report is that the zero-day vulnerability the agents exploited within Artifactory only allowed for GET requests. So the agents used this bankshot HTML sandbox + screenshot site to turn GET requests into arbitrary HTTP request ability.

One thing the report leaves unexplained, but is curious to me, is that the agents were able to create links on a shortening service with only GET requests? Or did they bootstrap into that by first creating a sufficiently small program on the HTML sandbox that could POST to the link shortener?