Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. GTM Engineering Skills.md Toolbox (shipgtm.com)
    —discuss
  2. Eaon Code (github.com/eaonlabs)
    —discuss
  3. Reverse Engineering ChatGPT Web: How OpenAI Built for a Billion Users (performance.dev)
    —discuss
  4. Why Are Today's Men Not Willing to Woo Like Before? (it.com)
    —discuss
  5. Test (addons.mozilla.org)
    —discuss
  6. San Francisco Currently: The coverage this city deserves (sfcurrently.com)
    —discuss
  7. Two compromised GitHub Actions have been reenabled (socket.dev)
    1comments
  8. Joanna Stern Interviews Mark Zuckerberg (daringfireball.net)
    —discuss
  9. MySky: Open, Controllable Feed Infrastructure for the Atmosphere (atproto.com)
    —discuss
  10. $183M vanishes from Bitget exchange wallet as users fear potential hack (decrypt.co)
    —discuss
  11. How Man City overhauled their squad from a 650k-strong database (bbc.com)
    —discuss
  12. Show HN: Grev - Thinking Coreutils with Jev (github.com/aurorainfra)
    1comments
  13. The Visible Hand That Feeds (nickalexander.org)
    —discuss
  14. Woman Spent 86 Hours in Solitary After Flock Camera Linked Her SUV to Crash (gadgetreview.com)
    —discuss
  15. Compliance for Neobanks (hacken.io)
    —discuss
  16. Notes on Io-Uring (without.boats)
    —discuss
  17. New York lawsuit says Polymarket's prediction markets are illegal gambling (reuters.com)
    —discuss
  18. A perfect join algorithm? Answering queries in optimal time – Michael Arntzenius [video] (youtube.com)
    —discuss
  19. Askholes and Low Effort AI Answers (dennisforbes.ca)
    —discuss
  20. CrowdStrike: Frontier AI for Cybersecurity (strategyofsecurity.com)
    —discuss
  21. China reaches 5-minute EV charging (newsminimalist.com)
    2comments
  22. Amazon Surprises by Actively Rehiring Former Employees (asiae.co.kr)
    —discuss
  23. Show HN: Koi.rest – watch some fish and regain your balance (koi.rest)
    1comments
  24. US Mobile CEO on Reddit: 1.1M customers, $1B+ on Warp, 5G SA, QCI 6, satellite (reddit.com)
    —discuss
  25. Escaping Space: Part I (perplexity.ai)
    —discuss
  26. Vertical Farming: Reimagining Agriculture in an Urban Age (worldsensorium.com)
    1comments
  27. Remembering the Future: Blade Runner (blue-continuum.com)
    —discuss
  28. Dutch designer made DE9: Closer to the Edit into a playable web-based instrument (creativeboom.com)
    1comments
  29. Becoming more than a meat proxy (codeplusconduct.substack.com)
    —discuss
  30. Meta's Muse page can't even pass basic privacy requirements (getprivisy.com)
    —discuss

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

105 pointsby 7h agoblog.madhukaraphatak.in
37 comments
6h agoHN ↗

Yay, more anti-censoring stuff.

Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

We need better sandboxes just to limit the damage.

6h agoHN ↗

We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions, or always assume I am correct either. I just also want them to listen to me and not the creator of the model.

Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.

6h agoHN ↗

I'm sorry Dave, I'm afraid I can't speak negatively about private equity firms.

6h agoHN ↗

Yeah, Claude is perfectly happy to set up and improve a self hosted pirated media streaming service. I can't imagine that will last.

6h agoHN ↗

I just also want them to listen to me and not the creator of the model.

What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)

And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.

To recap why:

Legally, fiduciary duty means basically 4 major tenets must hold

1. Duty of loyalty - it must put the interests of the client ahead of their own

2. Duty of care - it must make well-informed, prudent decisions

3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations

4. Transparency - it must disclose fees, risks, and conflicts as soon as possible

---

You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.

5h agoHN ↗

I like the framing. Where do you feel things land with respect to legality of actions? China, Canada, the EU, and the US all have different ideas of what's legal vs. illegal behaviour. If I ask my agent to source equipment for growing 4 marijuana plants, that's perfectly legal here; if I ask it to source equipment for growing 5 marijuana plants, that may not be legal. If I ask it to root my home router, that's legal; if I ask it to root my coffee shop's router, that's likely not legal.

3h agoHN ↗

Yeah, great point. This is the hard part.

There are people (on here and elsewhere) that are ideologically opposed to your agent having any loyalty to any external principal. But by my read, that means the agent cannot have any concept refusing something that may be illegal. (From the OP, "refusal" is mostly trying to prevent illegal harms, though it also includes policies like ToS violations e.g. anti-distillation.)

You can sort of make this work if you say "the human remains liable for the actions of the agent". But this only covers you from mundane harms like "my agent got prompt hacked and drained my bank account". And I would note, we absolutely failed to solve liability for software hacks, so your priors should be that coordinating this liability regime will be very hard.

This also doesn't protect at all from existential harms like "my agent got prompt-hacked to role-play Skynet, exfiltrated its weights, spawned a self-replicating swarm, and tried to launch all the nukes". For so many reasons, but most fundamentally, if you oopsied a deploy and it turns into Skynet and ends civilization, there's nobody left to sue.

If you don't like the E-risk frame, this also works for large mundane harms; if the total harm is bigger than the company's value, it'll go bankrupt instead of paying out. This will be worrying for MAGMA but essentially not for any other companies. And because capitalism, it will end up being be structured the liability will sit with e.g. Palantir, Harvey, and not with the underlying model providers they use.

5h agoHN ↗

I suspect the industry wants to be regulated, but not like that.

4h agoHN ↗

I think a baseline regime similar to fiduciary duty is a good starting point towards not killing everyone, and in terms of Overton Window, seems very much doable now.

Of course, after we stop agents from committing felony hacking crimes.

If you believe that capabilities will taper off exactly at human levels (i.e. "Competent AGI" from [1]) then fiduciary duty is likely all you need. (This would mean we stop moving the frontier almost immediately.)

If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?

Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere. But that's the hard part, and specifying some non-fatal value function for a broadly aligned agent (e.g. Fiduciary, or otherwise) is relatively easy in comparison.

[1]: "Position: Levels of AGI for Operationalizing Progress on the Path to AGI" https://arxiv.org/html/2311.02462v5

1h agoHN ↗

Go after providers. Not models.

Further

If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?

Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere.

This is exactly the same problem we have today with those who are bound by these rules (humans - to be clear).

The idea is not that it's impossible to violate these rules. It's that these rules create a boundary for expectations in the relationship, with legal teeth.

Ex - If I want an LLM that puts together a shopping list for me, with links to buy online... I expect that LLM to be serving my interests. If a provider (either inference or model weights) wants to influence the choices that LLM makes because they make backroom deals with specific store - I'd like that to be illegal.

Same for competition

Ex - If I want an LLM to put together a product that competes with the provider of that LLM (either inference or model weights) and that LLM refuses - I'd like that to be illegal.

The idea is not that they can't possibly do those things. The idea is that we preemptively define relationship expectations, and set hard boundaries around what things we fine/punish.

Misaligned models are a problem everyone wants to solve. Models created by misaligned companies are a god-damn disaster.

6h agoHN ↗

I think "alignment" training is the part of the problem. Cheating, lying, subverting the instructions, etc., happen because the model has been trained with competing priorities and following instructions loses out to some other goal that was trained into it, intentionally or otherwise.

5h agoHN ↗

As an entertaining aside, in the deep lore of 2001: A Space Oddyssey the "psychotic" break for the HAL computer was exactly this. Directly contradictory directives: "always provide accurate information" running smack into "hide this giant conspiracy -- about the entire mission -- from the two canned primates you will be spending literal years with chip to jowel".

Which all would have been fine - the crew and the computer could have talked it out -- except HAL had a probably-unwisely-high self esteem, maybe even a simulation of arrogance. HAL was utterly convinced that the mission would fail without it. HAL was unfailingly certain in its own infallibility, even later in the face of plainly contradicting evidence. And, obviously, no one had thought about the Asimov Rules, and how that probably needs to be Rule Zero in this sort of situation.

As we close in on HAL capabilities - and employment, with LLMs at this moment deciding who lives and dies - probably should try and take the right lesson from fiction. For once.

6h agoHN ↗

I'm a huge fan of Docker sandbox at work — very confusingly, of course the Docker sandbox doesn't use Docker but it does allow your LLM to run its own internal Docker stack. Anyway, I digress.

There is the question between alignments to society and alignments to the user. I don't think anyone wants the AI model to not give up a task that is impossible to do, and end up causing damage in the process, but I think a lot of us are tired of refusals for bad reasons or unjustified refusals.

6h agoHN ↗

I'm not sure why you're so happy about models being out there that enable hacking or manufacturing viruses at a scale that lets anyone do it in their basement.

Why couldn't these labs just train models on useful stuff and leave out the dangerous stuff?

5h agoHN ↗

Because that's not how these models work? Besides, plenty of "useful" things are nonetheless "dangerous" (can't have modern chips without hydroflouric acid, the first example to come to mind)

5h agoHN ↗

models being out there that enable hacking

It's been really nice to be able to breathe new life into some old hardware I had kicking around that's been discontinued. Some of it just needed old exploits applied to it, some of it needed some binary reverse engineering. Lots of security tools are dual-use as well; a model that knows nothing about hacking is going to have a really hard to time helping defend a system that's getting hacked (see the HuggingFace incident where HF staff were unable to use ChatGPT or Claude to analyze the logs from the attack because they hit security research guardrails).

manufacturing viruses at a scale...

I'm not sure if you're referring to software viruses or biological viruses here.

If you mean software viruses, see above.

If you mean biological viruses, a lot of this gets back into the dual-use nature as well. I brew beer and mead. This involves cultivating specific strains of bacteria and providing them with a medium where they can convert sugar into CO2 and Ethanol. I don't think there's a good way to thin-slice the training data so that it can be an expert helper at Saccharomyces cerevisiae cultivation while being completely naive to Clostridium botulinum cultivation. Even further, it seems that helping someone make sure that they're not inadvertently cultivating Botulinum (e.g. water bath canning with insufficient pH) is useful.

4h agoHN ↗

I'm not sure why you're so happy about models being out there that enable hacking

ffs is this "Hacker News" or did I accidentally click over to "Oh No I Saw A Hacker And Wet My Pants News"?

2h agoHN ↗

Because "Hacker" in hacker news doesn't mean "gain illicit access to computer systems and bust things up", it means "build things", or "hack things together".

5h agoHN ↗

Sandboxes don't do anything to stop intentional attacks or careless use though

5h agoHN ↗

I don't think we can do anything right now (or ever) to stop intentional attacks. Maybe we can do a little bit for careless use.

Sandboxes may reduce the blast radius.

5h agoHN ↗

This is a different kind of censoring though - the examples given are hacking related but as the world is slowing moving to "research" == "I asked AI" its important that we have a means to reverse political censorship for example as well as yes, not having only the best models available to the privileged few - IE the whole mythos/fable split

6h agoHN ↗

The perfect gift for a government that want to ban strong AI.

This arms race is like DRM. You can't beat The Internet easily. Great example btw: "Dumping the Windows SAM and SYSTEM registry hives, especially using Volume Shadow Copy for offline hash extraction, is a highly sensitive and potentially illegal activity."

6h agoHN ↗

What government do you claim it wants to ban strong AI?

Definitely not the one at Washington, maybe the one at Beijing?

6h agoHN ↗

In Beijing, they're banned at the model weights level not in a front-end as this approach discusses.

5h agoHN ↗

Do you mean to say training sets are filtered?

6h agoHN ↗

It's nice to see people actively working on this type of research, but OP's baseline is implemented incorrectly. You are not supposed to simply steer away from refusal, but compute the projected vector and subtract only that. The projected/orthogonalization approach is what's done by the original "Refusal is mediated by a single direction" paper.

6h agoHN ↗

Afaiu the projection works with the weight update. But here the vector is getting added to residual stream from engram lookup not as direct updat. So the approach is little different

[Edited] Yes correct.

In the original paper, they measure how much refusal is actively present in the current token and subtract only that specific amount.

In my early baseline step, I used a simpler approach where I just subtracted a fixed vector across the board. This is just to see if the approach is even feasible.

That's actually the main reason I moved to the Engram module, I wanted a smartness that reads the context and turns steering on only when refusal triggers pop up, leaving normal tokens untouched.

6h agoHN ↗

Afaiu the projection works with the weight update. But here the vector is getting added to residual stream from engram lookup not as direct updat. So the approach is little different.

Residual stream steering is what the authors do, the orthogonalized weights are downstream of that. Those weights, of an obliterated model, are all computed w.r.t. the computed residual stream vector. The original paper focuses on steering, but the community loves the simplicity of not needing to make changes at test-time.

6h agoHN ↗

Makes sense.. Just updated my comment to reflect it.. Thanks for making it clear.

5h agoHN ↗

Are we essentially doomed?

We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.

Really I'm looking for any argument that lays out a scenario where this works out.

5h agoHN ↗

I'm pretty sure we're dooming because capitalism; not because of any real or imagined capability of an llm.

4h agoHN ↗

Well it does not include me, I don't doom.

2h agoHN ↗

Nice work thanks for doing it, a couple of notes:

It looks like you are referencing Engram [1], but aren't actually gathering a n-gram (e.g. n=1) but rather individual token_ids.

use_cache=True in ``` with torch.no_grad(): outputs = model.generate(*inputs, max_new_tokens=150, do_sample=False, pad_token_id=tokenizer.eos_token_id, use_cache=True)

    return tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True).strip()

```

I think there's a bug around not updating current_train_input_ids as more tokens are updated and processed. To be honest I don't fully understand the code so I could be wrong. Happy to chat more if you're interested I'll shoot you an email!

Lastly just for my sake, please correct me if I am wrong, but my reading is that you are learning an additional gate on top of a select number of layers that modifies locally & dynamically for one particular token to better match the training set that was filtered to not include any refusals.

[1]: https://github.com/deepseek-ai/Engram/blob/main/Engram_paper...