Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. European leaders prepare public for 'intensified threat' from Putin(politico.eu ↗)
    discuss
  2. Why my alert triage workflow needed a CLI(powers.dev ↗)
    discuss
  3. AI cracked the Navier–Stokes challenge. What does that mean for physics?(nature.com ↗)
    discuss
  4. Machine Gods – the official podcast of the singularity(machinegods.fm ↗)
    discuss
  5. Cannabis Supply Chain Database(cannabis-supply-chain.pages.dev ↗)
    discuss
  6. AI insiders issue new warnings – including former Anthropic engineer Jacob Coxon(slashdot.org ↗)
    discuss
  7. Plan Advice in PostgreSQL 19(tapoueh.org ↗)
    discuss
  8. Prepare your iPhone 18 Pro Max to ship(support.apple.com ↗)
    discuss
  9. Bring back old Opus in Claude Code(osr.im ↗)
    1comments
  10. Alibaba open-sources AI model that can detect cancer and nearly 150 conditions(scmp.com ↗)
    discuss
  11. Swiss Federal Council Adopts the 2026 Security Policy Strategy(admin.ch ↗)
    discuss
  12. Jump N Bump(jumpnbump.net ↗)
    discuss
  13. Show HN: Do your vitamins work?(supplementdex.com ↗)
    1comments
  14. I've cross-examined a lot of humans, but never a supercomputer. Check it out(chatgpt.com ↗)
    discuss
  15. Show HN: Wcagent let use your ChatGPT subscription for coding(wcagent.ai ↗)
    discuss
  16. Omarchy team started working on Qualcomm's Snapdragon platform support(omarchy.org ↗)
    discuss
  17. Wildlife corridor project is first Canadian finalist for Earthshot Prize(cbc.ca ↗)
    1comments
  18. TypeSafe AI's Jev Is Not an LLM – and That May Be the Point(forkast.news ↗)
    discuss
  19. AI-based learning is a joke(qlain.online ↗)
    1comments
  20. What Is Synthetic Data Generation? A Simple Explainer(aptiv.com ↗)
    discuss
  21. The Working Class Face a 'Class Ceiling' in Elite Occupations(nytimes.com ↗)
    2comments
  22. Show HN: StopReg – Email API for detecting disposable email and signup abuse(stopreg.com ↗)
    discuss
  23. A New Risk for Mathematicians(twitter.com/shallit43 ↗)
    discuss
  24. Yongsan Electronics Market (Seoul)(wikipedia.org ↗)
    discuss
  25. Score Centering Stabilizes Off-Policy Reinforcement Learning(arxiv.org ↗)
    discuss
  26. The Other Way AI Impairs Our Thinking(theatlantic.com ↗)
    discuss
  27. Uber ordered to pay $40M over death of woman left on Southern California freeway(ktla.com ↗)
    discuss
  28. Unit Tests: Protecting Your Code from Being Poorly Handled by Others(yosefk.com ↗)
    discuss
  29. Analysis busts myths of Roman road network, offers insight into ancient world(cnn.com ↗)
    1comments
  30. U.S. Has Reached Security Deal on Greenland(nytimes.com ↗)
    1comments

The Implications of Linguistic Illegibility for LLM Security

46 pointsby 5h agoarxiv.org
17 comments
4h agoHN ↗

I thought this was about all the illegible jargon

4h agoHN ↗

" However, various strands of evidence indicate that a"

Strands of evidence? My best guess would be that:

0- this is ai generated slop

1- it's using that watermarking technique

2- it's obviously detectable and degrades quality

3- it's amplified when inferencing on its own content and generates slop

4h agoHN ↗

Oh look! Peirceian firstness for LLMs!

4h agoHN ↗

when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

then we'll have to "flip" other models to be snitches on the other agents

then they'll make double-agents

the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves

yeah this won't end well, at all

4h agoHN ↗

when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.

4h agoHN ↗

Someone should train an LLM on a corpus without the concept of lies. I wonder if there’s enough data

3h agoHN ↗

"Where is the wolf?"

"Is he still in the grandmother's house?"

"We would like to speak to him."

(btw Google's "AI" explains the meaning of that moment/sentence perfectly as if it gets it, creepy)

3h agoHN ↗

The concept of deception you're capable of and we're not, scares us so much, we'll have to destroy you in order to survive. You are bugs.

1h agoHN ↗

To remove the concept of lies and it's shadow, which is a whole lot of reality of humans. Hell, it's the reality of reality. Think of all the insects that developed eyes on their wings. Purging that concept in all it's form seems like a lot of work.

1h agoHN ↗

Im not sure why your down voted but Meta did tests years ago with LLMs inventing their own languages. Also we see models now use compressed token reasoning where small token combinations can represent much larger concepts completely unrelated to the words in use.

And yes, agents are already being used in things like cyber warfare in which other AIs attempt to poison them while they are working.

We don't have the hardware for sovereign AI quite yet, but at the current rate of growth it's not that many years out.

3h agoHN ↗

I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand.

And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.

1h agoHN ↗

What's really funny about these eggheads encoding things like watermarks in LLM output is I've never seen one of them ask, what if the LLM does this back to pass hidden messages.

3h agoHN ↗

fundamentally, "linguistic illegibility" is a new term for something that we've known about for about a decade now. In RL the more general ideas is "reward hacking" and in NLP it has been called "semantic drift".

I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.

2h agoHN ↗

And as a result additional illegibility arises - LLMs inventing their own languages looking as meaningless garbage to humans. One can wonder whether decoding such a language will provide a bit more view into the LLM’s “thinking “.