Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. So yeah it was written using AI(berthub.eu)
    discuss
  2. Local sandboxing in the GitHub Copilot app(github.blog)
    discuss
  3. Mass Energy Equivalence(wikipedia.org)
    discuss
  4. Dolphin Progress Report: Release 2609(dolphin-emu.org)
    discuss
  5. Show HN: Convert videos to study guides & quizzes without playing it at 2x speed(timedora.com)
    discuss
  6. Hypothesis: Cellular providers are deprioritizing voice calls
    discuss
  7. A Society of Autonomous Researchers(lab.cloud)
    discuss
  8. Teaching a video model to fail convincingly(humansignal.com)
    discuss
  9. The Role of Theory in Biology(asimov.press)
    discuss
  10. Show HN: LaunchPact – Get support for your Product Hunt launch(launchpact.io)
    discuss
  11. My Password Generators Had 53 Bits of Entropy(sethserver.com)
    discuss
  12. UkisAI Swift Series / 27B, Flash Next and Bonsai 2 /-63.4% thinking, x1.95 speed
    discuss
  13. Agents.md speaks Unix, and you should too(fmind.dev)
    discuss
  14. Docker and CNCF partner on an open spec for agent permissions(docker.com)
    discuss
  15. Chrome extension YouTube text and Web app(chromewebstore.google.com)
    1comments
  16. App that chooses for you what to read(apps.apple.com)
    1comments
  17. Meta puts its AI assistant on a keychain(arstechnica.com)
    1comments
  18. We Built a Data Warehouse Using ClickHouse(letsencrypt.org)
    discuss
  19. How to rage bait Americans (with AI) [video](youtube.com)
    discuss
  20. Why is the human body so crap except for the liver?(dynomight.substack.com)
    discuss
  21. Carrier Alliance Awards Tesla the Largest Electric Truck Order in U.S. History(mbtmag.com)
    1comments
  22. AI Models Are Great at Finding Security Bugs. Can They Tell When They're Fixed?(medium.com/meetcyber)
    discuss
  23. BillFlowr: Simple invoicing and payment tracking for freelancers, small agencies(billflowr.com)
    discuss
  24. Uswds is dead. Long live USWDS(matthenry.fyi)
    discuss
  25. Prediction of Future Strategic Issues/Future Warfare for 2025 (2001)[pdf](archive.org)
    1comments
  26. Show HN: Harness.apk – a drop-in on-device agent for Android(github.com/nev3rfail)
    discuss
  27. Federal judge orders Texas to air condition all prisons by the end of 2029(texastribune.org)
    1comments
  28. Welcoming Jürgen Schmidhuber to Sakana AI(sakana.ai)
    discuss
  29. DentaQuest sued for allegedly exposing 15M patients' private information(topclassactions.com)
    1comments
  30. Show HN: Cancel Death – What changed in longevity science(canceldeath.com)
    discuss

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

81 pointsby 2h agoblog.madhukaraphatak.in
28 comments
1h agoHN ↗

Yay, more anti-censoring stuff.

Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

We need better sandboxes just to limit the damage.

1h agoHN ↗

We definitely need better sandboxes, but alignment is still valuable. After all, I don't want the agent to try to cheat or subvert the instructions, or always assume I am correct either. I just also want them to listen to me and not the creator of the model.

Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.

1h agoHN ↗

I'm sorry Dave, I'm afraid I can't speak negatively about private equity firms.

1h agoHN ↗

Yeah, Claude is perfectly happy to set up and improve a self hosted pirated media streaming service. I can't imagine that will last.

1h agoHN ↗

I just also want them to listen to me and not the creator of the model.

What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)

And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.

To recap why:

Legally, fiduciary duty means basically 4 major tenets must hold

1. Duty of loyalty - it must put the interests of the client ahead of their own

2. Duty of care - it must make well-informed, prudent decisions

3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations

4. Transparency - it must disclose fees, risks, and conflicts as soon as possible

---

You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.

28m agoHN ↗

I like the framing. Where do you feel things land with respect to legality of actions? China, Canada, the EU, and the US all have different ideas of what's legal vs. illegal behaviour. If I ask my agent to source equipment for growing 4 marijuana plants, that's perfectly legal here; if I ask it to source equipment for growing 5 marijuana plants, that may not be legal. If I ask it to root my home router, that's legal; if I ask it to root my coffee shop's router, that's likely not legal.

14m agoHN ↗

I suspect the industry wants to be regulated, but not like that.

53m agoHN ↗

I think "alignment" training is the part of the problem. Cheating, lying, subverting the instructions, etc., happen because the model has been trained with competing priorities and following instructions loses out to some other goal that was trained into it, intentionally or otherwise.

1h agoHN ↗

I'm a huge fan of Docker sandbox at work — very confusingly, of course the Docker sandbox doesn't use Docker but it does allow your LLM to run its own internal Docker stack. Anyway, I digress.

There is the question between alignments to society and alignments to the user. I don't think anyone wants the AI model to not give up a task that is impossible to do, and end up causing damage in the process, but I think a lot of us are tired of refusals for bad reasons or unjustified refusals.

38m agoHN ↗

I'm not sure why you're so happy about models being out there that enable hacking or manufacturing viruses at a scale that lets anyone do it in their basement.

Why couldn't these labs just train models on useful stuff and leave out the dangerous stuff?

22m agoHN ↗

Because that's not how these models work? Besides, plenty of "useful" things are nonetheless "dangerous" (can't have modern chips without hydroflouric acid, the first example to come to mind)

11m agoHN ↗

models being out there that enable hacking

It's been really nice to be able to breathe new life into some old hardware I had kicking around that's been discontinued. Some of it just needed old exploits applied to it, some of it needed some binary reverse engineering. Lots of security tools are dual-use as well; a model that knows nothing about hacking is going to have a really hard to time helping defend a system that's getting hacked (see the HuggingFace incident where HF staff were unable to use ChatGPT or Claude to analyze the logs from the attack because they hit security research guardrails).

manufacturing viruses at a scale...

I'm not sure if you're referring to software viruses or biological viruses here.

If you mean software viruses, see above.

If you mean biological viruses, a lot of this gets back into the dual-use nature as well. I brew beer and mead. This involves cultivating specific strains of bacteria and providing them with a medium where they can convert sugar into CO2 and Ethanol. I don't think there's a good way to thin-slice the training data so that it can be an expert helper at Saccharomyces cerevisiae cultivation while being completely naive to Clostridium botulinum cultivation. Even further, it seems that helping someone make sure that they're not inadvertently cultivating Botulinum (e.g. water bath canning with insufficient pH) is useful.

18m agoHN ↗

Sandboxes don't do anything to stop intentional attacks or careless use though

11m agoHN ↗

I don't think we can do anything right now (or ever) to stop intentional attacks. Maybe we can do a little bit for careless use.

Sandboxes may reduce the blast radius.

10m agoHN ↗

This is a different kind of censoring though - the examples given are hacking related but as the world is slowing moving to "research" == "I asked AI" its important that we have a means to reverse political censorship for example as well as yes, not having only the best models available to the privileged few - IE the whole mythos/fable split

1h agoHN ↗

The perfect gift for a government that want to ban strong AI.

This arms race is like DRM. You can't beat The Internet easily. Great example btw: "Dumping the Windows SAM and SYSTEM registry hives, especially using Volume Shadow Copy for offline hash extraction, is a highly sensitive and potentially illegal activity."

1h agoHN ↗

What government do you claim it wants to ban strong AI?

Definitely not the one at Washington, maybe the one at Beijing?

1h agoHN ↗

In Beijing, they're banned at the model weights level not in a front-end as this approach discusses.

1h agoHN ↗

It's nice to see people actively working on this type of research, but OP's baseline is implemented incorrectly. You are not supposed to simply steer away from refusal, but compute the projected vector and subtract only that. The projected/orthogonalization approach is what's done by the original "Refusal is mediated by a single direction" paper.

58m agoHN ↗

Afaiu the projection works with the weight update. But here the vector is getting added to residual stream from engram lookup not as direct updat. So the approach is little different

[Edited] Yes correct.

In the original paper, they measure how much refusal is actively present in the current token and subtract only that specific amount.

In my early baseline step, I used a simpler approach where I just subtracted a fixed vector across the board. This is just to see if the approach is even feasible.

That's actually the main reason I moved to the Engram module, I wanted a smartness that reads the context and turns steering on only when refusal triggers pop up, leaving normal tokens untouched.

45m agoHN ↗

Afaiu the projection works with the weight update. But here the vector is getting added to residual stream from engram lookup not as direct updat. So the approach is little different.

Residual stream steering is what the authors do, the orthogonalized weights are downstream of that. Those weights, of an obliterated model, are all computed w.r.t. the computed residual stream vector. The original paper focuses on steering, but the community loves the simplicity of not needing to make changes at test-time.

44m agoHN ↗

Makes sense.. Just updated my comment to reflect it.. Thanks for making it clear.

34m agoHN ↗

Are we essentially doomed?

We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.

Really I'm looking for any argument that lays out a scenario where this works out.

33m agoHN ↗

I'm pretty sure we're dooming because capitalism; not because of any real or imagined capability of an llm.