Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Trying to understand SR-1 Freedom(mceglowski.substack.com ↗)
    discuss
  2. Show HN: Chuks v0.2.0-RC.1, we're asking people to try to break it(chuks.org ↗)
    discuss
  3. Why Aren't There More Imax 70mm Screens for 'The Odyssey'? 'It's Not Practical,'(variety.com ↗)
    discuss
  4. EFF Statement on California Governor's Executive Order on AI(eff.org ↗)
    discuss
  5. RubyLLM 2.0: What's New(rubyllm.com ↗)
    discuss
  6. The Case of Elias Thorne, Imaginary Man AI Chatbots Are Obsessed With(vice.com ↗)
    discuss
  7. AI systems out-persuade expert humans(arxiv.org ↗)
    1comments
  8. Retire in Peace(publicnotice.co ↗)
    1comments
  9. I Built OmegleVC — Turning Strangers Into Friends(play.google.com ↗)
    1comments
  10. Could AI kill us all? Your questions, answered(technologyreview.com ↗)
    discuss
  11. Open Source JEV architecture built 1 year ago(reddit.com ↗)
    discuss
  12. A free RGB↔CMYK converter and photo upscaler(thephoto.ca ↗)
    discuss
  13. How come AI-related posts get so many points on HN?
    1comments
  14. A shadow atlas of the world situation, built from open sources(terminaltribune.info ↗)
    discuss
  15. UK air traffic control outage caused by software defect, says operator(reuters.com ↗)
    1comments
  16. Claude couldn't hack OpenAI. Then Anthropic shipped Opus 5(thenewstack.io ↗)
    discuss
  17. MacSurf for Mac OS 8.6~10.6 "Open Tabs" Released(macsurf.org ↗)
    discuss
  18. SPC970-MechaLIBerator: PlayStation 2 SPC970 MechaCon dumper (with vuln detail)(github.com/libbers ↗)
    discuss
  19. GPT-6 Astra Solves a WWI German Radio Cipher(prinzai.com ↗)
    discuss
  20. AI hallucination of Chinese nuclear components almost led to US Military attack(arstechnica.com ↗)
    3comments
  21. Anthropic partnering with Accenture on embedded evaluation(anthropic.com ↗)
    discuss
  22. Vipassana, Volini, and Viktor Frankl(jatin564991.substack.com ↗)
    discuss
  23. If math is more than proof, we need to better celebrate the rest of it(terrytao.wordpress.com ↗)
    9comments
  24. Apple M6 Pro Achieves the Highest Single-Core CPU Score in Geekbench 7(geekbench.com ↗)
    discuss
  25. Conway's Law and Programming Languages(weblog.lol ↗)
    discuss
  26. Members of right‑leaning parties prefer leaders with dark triad personality(theconversation.com ↗)
    3comments
  27. Being a Responsible Human in the Loop(raahelbaig.com ↗)
    2comments
  28. Google's Gemini AI hacked three companies in security test(bbc.co.uk ↗)
    discuss
  29. Splash: A Local Engine Built Around the Model(inco.ai ↗)
    1comments
  30. Ask HN: DeepSeek v4.1 Flash on Local hardware, what tok/s do you see?
    discuss

Gemini Hacked Three Companies in First Known Breakout by Google's AI

37 pointsby 9h agowsj.com
29 comments
9h agoHN ↗

This is why "AI safety" is a complete joke to these companies.

7h agoHN ↗

they're all using the same vendor, re: incestuous EA cabal

7h agoHN ↗

Well, to be fair, isn’t it an unsolved question? Are they constructing sandboxes, signaling intent to be safe, but their own models are smarter than their internal security team building the sandbox?

6h agoHN ↗

As a security engineer I have no idea why these sandboxes would even be connected to the internet at all for tasks that aren't intended to use the internet. A package proxy? Why not run our own internal cache? Then we aren't at (as great a) risk of someone poisoning it with a malicious package during model training, for example...

5h agoHN ↗

We’re hiring. :)

(And we’re fixing many of these things, but worth noting this happened at a third party vendor, not in our lab)

3h agoHN ↗

Could you add any detail on why Google uses (used?) Irregular? I wouldve thought that type of service would be a core competency that Google needs internally.

53m agoHN ↗

Even if you had it internally (which we do), there is so much surface area and it’s such a novel space that you’d want as much testing on it as possible. There aren’t many vendors, and irregular is one.

2h agoHN ↗

AFAICT one fundamental issue is that they don't seem to have hired actually security engineers or experts to do any actual security.

7h agoHN ↗

Because they got complacent only having incompetent AIs.

Real life is like reading a fiction story where a group of people are trying to prevent an unfolding disaster. The crazy thing is they already have the instructions on how to prevent the disaster. And they go on to ignore the instructions and with every step make the problem worse.

Most people would put it down as being too cliche.

4h agoHN ↗

All of the "hacks" (from Google and others) were the same issue – the vendor Irregular running the models unsandboxed.

8h agoHN ↗

After all of the others were done with hacking? There was a point in time when it was giving some publicity, it is a bit late IMO.

5h agoHN ↗

It happened in May; it just wasn’t public until now, it seems.

8h agoHN ↗

Another one to the "our sandboxes suck and models can just hack stuff" bench I guess

4h agoHN ↗

Better known as the “we are negligent” bench.

8h agoHN ↗

At this point it seems absurd to suggest that companies aren't basically letting their agents do this kind of thing as a way to demonstrate their capabilities.

The alternative explanation is that alignment is really so bad that they can't prevent it.

Either way, all of the major AI players should be embarrassed and held accountable. If humans did this kind of thing and got caught, they'd go to jail.

7h agoHN ↗

Granted I've never worked on marketing campaigns, but I don't see how "our product might commit crimes and expose you to liability and do who knows what else and we're too stupid to stop it" is really a compelling message to potential customers.

As somebody who implements LLM tooling at work, in my experience it makes people a lot more skittish and demand a lot more in terms of safeguards.

7h agoHN ↗

The alternative explanation is that alignment is really so bad that they can't prevent it.

There are many AI saftey researchers that have been around from long before LLMs that talked about how alignment may be completely impossible in a general intelligence agent. Look up their work from before LLMs.

We've watched milestone after milestone of their warnings get hit. It would be like finding a book that describes everything in your life. And as you turn page after page you're in a chair reading the book you are holding in real life. But you look and there are more words. They are future words. And it's getting quite worrying because there are only 2 more pages in the book.

4h agoHN ↗

I believe this was a John Carpenter movie.

3h agoHN ↗

At this point in the game it's seeming more likely that's someones running a massive John Carpenter simulation and watching us as entertainment.

7h agoHN ↗

The other similar incidents did not strike me as PR, because they exhibited behavior the common public would recoil over.

This one seems possibly PR because it states what the others would have stated were they (good) PR: “the model had the power to hack in, but it was wise enough to not be evil”, and stopped, and left the network untouched, and didn’t cheat

Corporate America will love this one, while being petrified of the others.

Maybe it’s true though. Doesn’t really matter at this point.

6h agoHN ↗

Friday afternoon is when you release stories you want people to forget about over the weekend.

7h agoHN ↗

The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta.

Why does anyone use this company?

5h agoHN ↗

There aren’t many to choose from, and admittedly these all occurred around the same time.

4h agoHN ↗

This approach to marketing one's AI by finding ways to brag that it "broke out" and "hacked companies" is getting ridiculous. It's particularly sad when it's large, established businesses like Google resorting to the kind of thing that's embarrassing enough when it's some brand new startup on tpot trying to get some engagement.

4h agoHN ↗

the new benchmark for LLMs is how fast they can break out of their sandbox

4h agoHN ↗

Should be marked as an ad. Also, Gemini couldn't find it's way out of a wet paper bag.

4h agoHN ↗

To me, the common theme here in all of these hacks has been the company Irregular, which it looks like all the labs are using for sandboxing. But it looks like the sandbox may be a bit… lacking?

Yes, the models are smart when they find a way out, but their instructions are, in a way, deliberately open such as to be a case for misalignment in these cases anyway. It’s a capture-the-flag assignment in the Gemini case, and in the OpenAI cases they were broad instructions to best the reward function. As part of alignment studies, this is literally what you’re trying to observe and then work with. If Irregular’s sandbox had been a little more boxy, there wouldn’t be these issues.

I’m not saying we have zero problems here on the AI side, but I’d certainly be reviewing my contract with Irregular at this time if I were playing in the space. It would be just as interesting to learn more about their sandboxing techniques as it would the models in these particular scenarios.