Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Anthropic is planning to launch its IPO in November(wsj.com ↗)
    discuss
  2. Valve has open-sourced Lepton, its tool to bring Android games to Steam(theverge.com ↗)
    discuss
  3. Exclusive-Anthropic sets up biology lab as it ramps AI drug program(yahoo.com ↗)
    discuss
  4. I Hate Workday(nelson.cloud ↗)
    discuss
  5. Y Combinator's PAC is throwing money at Republicans across the country(gazetteer.co ↗)
    discuss
  6. War may be coming. Are we psychologically ready?(bbc.com ↗)
    discuss
  7. Flawed AI Intel Brought US to the Brink of Confronting China, Claims Report(ndtvprofit.com ↗)
    discuss
  8. AI cloud provider Nscale files for IPO on $140.6M revenue, $1.02B net loss(cnbc.com ↗)
    discuss
  9. Show HN: Find Street Parking in NYC(avi.nyc ↗)
    discuss
  10. Software-Based Live Migration for RDMA – Proceedings of the ACM Sigcomm 2025(acm.org ↗)
    discuss
  11. Anthropic sets up a bio research lab for physical experiments(engadget.com ↗)
    2comments
  12. Show HN: Prohibition of Nuclear Launch Automation
    discuss
  13. A CPU Backdoor (2025)(phrack.org ↗)
    discuss
  14. Show HN: TypeSeer – on-device autocomplete for every text field on macOS(typeseer.com ↗)
    1comments
  15. Show HN: Rediagram – put the diagrams you have on your brand(rediagram.app ↗)
    discuss
  16. We Must Create the Shit Machine(mcsweeneys.net ↗)
    1comments
  17. Intel Appears to End Its Bug Bounty Program(phoronix.com ↗)
    discuss
  18. DeepSeek v4.1 Flash avg 102 tps on 4x RTX6000 pro max-q, 2.1x up from v4-flash(level1techs.com ↗)
    2comments
  19. We were right (about passkeys) all along(mailpace.com ↗)
    discuss
  20. Stack Overflow relaunched Developer Story (who certifies that a human wrote it?)(stackoverflow.blog ↗)
    discuss
  21. Friday Facts #446 – An ARM and a Frame(factorio.com ↗)
    discuss
  22. The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It(arxiv.org ↗)
    discuss
  23. I Have Been a DelGuard
    discuss
  24. The Mic Is On. So Is Live Auto-Tune.(nytimes.com ↗)
    1comments
  25. Institutional Parasitism in Open Technology Communities(wasabisys.com ↗)
    discuss
  26. UK could force phone companies to add 'anti-theft protections'(bbc.com ↗)
    discuss
  27. Using jev to improve product experiences is pretty crazy(elvex.com ↗)
    3comments
  28. What I learned from using FreeBSD as a main OS for a summer(divanv.com ↗)
    discuss
  29. A Lack of Honesty Is the Ultimate Killer(phillipspobrien.substack.com ↗)
    discuss
  30. Ask HN: What do you think of Noul, a new decision primitive
    discuss

The Implications of Linguistic Illegibility for LLM Security

32 pointsby 2h agoarxiv.org
12 comments
2h agoHN ↗

I thought this was about all the illegible jargon

2h agoHN ↗

" However, various strands of evidence indicate that a"

Strands of evidence? My best guess would be that:

0- this is ai generated slop

1- it's using that watermarking technique

2- it's obviously detectable and degrades quality

3- it's amplified when inferencing on its own content and generates slop

2h agoHN ↗

Oh look! Peirceian firstness for LLMs!

2h agoHN ↗

when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

then we'll have to "flip" other models to be snitches on the other agents

then they'll make double-agents

the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves

yeah this won't end well, at all

1h agoHN ↗

when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.

1h agoHN ↗

Someone should train an LLM on a corpus without the concept of lies. I wonder if there’s enough data

1h agoHN ↗

"Where is the wolf?"

"Is he still in the grandmother's house?"

"We would like to speak to him."

(btw Google's "AI" explains the meaning of that moment/sentence perfectly as if it gets it, creepy)

1h agoHN ↗

The concept of deception you're capable of and we're not, scares us so much, we'll have to destroy you in order to survive. You are bugs.

1h agoHN ↗

I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand.

And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.

48m agoHN ↗

fundamentally, "linguistic illegibility" is a new term for something that we've known about for about a decade now. In RL the more general ideas is "reward hacking" and in NLP it has been called "semantic drift".

I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.