Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Breaking Up with Google Play: Why Conversations Is Now Free (gultsch.de)
    105comments
  2. Understanding the Impact of LLM Watermarking on AI Agent Behavior (lasso.security)
    22comments
  3. Fifteen years later, the Apple Cards origin story (lexontech.org)
    17comments
  4. Revealing the details of how OpenAI agents hacked Hugging Face (swarmtraces.org)
    359comments
  5. We're gonna need a lot more mathematicians (terrytao.wordpress.com)
    293comments
  6. Plan mode is dead (aymannadeem.com)
    371comments
  7. Modern Object Pascal Introduction for Programmers – Castle Game Engine (castle-engine.io)
    7comments
  8. ASML currently sells no chipmaking machines in Europe, executive says (nltimes.nl)
    70comments
  9. Ollaya – Ollama for open-source, Jev-style decision models (ollaya.dev)
    126comments
  10. Floci: Locally emulating any cloud service (floci.io)
    11comments
  11. A single function Jev-like wrapper for LLMs, including vision models (allanrbo.blogspot.com)
    30comments
  12. Is your Postgres migration safe or not safe? (safenotsafe.dev)
    20comments
  13. 16GB iPod Nano 3G Upgrade (tuckerosman.com)
    8comments
  14. Show HN: Jev Plays Pokémon Red (jev-pokemon.vercel.app)
    91comments
  15. Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML (jointhefreeworld.org)
    10comments
  16. What even is an OS now? (sockpuppet.org)
    325comments
  17. Calculating atmospheric drag on satellites for a Cubesat [pdf] (osti.gov)
    5comments
  18. Ask HN: Who's still keeping a DOS machine up because the business depends on it?
    163comments
  19. Gravity seems holographic. What does that mean for reality? (quantamagazine.org)
    186comments
  20. The Murky History of Soviet-Born Tetris (mitpress.mit.edu)
    15comments
  21. Jury finds Facebook liable for deceiving users in Cambridge Analytica case (cbsnews.com)
    80comments
  22. Scientists build most accurate atomic clock (phys.org)
    17comments
  23. HomelabFest will be in St. Louis in September 2027 (homelabfest.org)
    38comments
  24. One Month Without AI (bustikiller.com)
    126comments
  25. The Copilot+ PC brand is dead (windowscentral.com)
    22comments
  26. Fourier Analysis: Drawing Llamas with Circles (adekau.github.io)
    6comments
  27. The far side of the Moon provides clues to a previous magnetic field (ethz.ch)
    1comments
  28. Excel now supports multiple values in a single cell (techcommunity.microsoft.com)
    152comments
  29. From Thin Air to Bootable Images: The Tine Build System (amutable.com)
    8comments
  30. Lab on a Contact Lens Can Measure Stress Through Serotonin (ieee.org)
    20comments

Understanding the Impact of LLM Watermarking on AI Agent Behavior

40 pointsby 1h agolasso.security
18 comments
48m agoHN ↗

Watermarking sounds like a good idea, but it's not. Token drift from watermarking will degrade the quality of outputs and could allow clever people to circumvent guardrails.

You should assume all text is AI generated. If you want to "test" someone at school or during an interview, have them write with a pencil and paper.

13m agoHN ↗

Changing a random seed could either improve or degrade the output. In theory, better and worse outputs should be equally probable, depending on your luck.

44m agoHN ↗

This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.

38m agoHN ↗

Model companies are doing this for themselves anyways, it’s so they don’t feed generated content back into the slopper and collapse the model. From that angle it over time contributes to better model quality.

6m agoHN ↗

Also - it forms a cartel.

Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).

So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).

This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.

21m agoHN ↗

The article has a pretty decent summary of the watermarking algo though. This reads as a pretty dogmatic statement in comparison.

In your analogy: What if seed 42 specifically causes poor quality behaviour (in some contexts specifically). Normally, these quality differences will be washed out because the seed is random, now it is no longer random, so shouldnt we check into specific behaviour under this specific seed?

15m agoHN ↗

That's not true. Watermarks are messing with the next token generation probabilities based on some random seed. The quality is neccesarily lower, the difference is simply too small to notice, typically.

7m agoHN ↗

No, the probability distribution is the same. Watermarking changes the rng sequence used to pick from that distribution.

10m agoHN ↗

“When implemented correctly” is probably what people are complaining about.

Opus 5 started adding a bunch of comments to code, even when instructed not to, and for very simple changes where the comment itself was longer than the code change. Was that so that there are enough tokens outputted for watermarking? Many people suspected so.

6m agoHN ↗

So that's where that nonsense comes from...

44m agoHN ↗

Am I missing something, or did they actually completely misunderstand how this technology works?

15m agoHN ↗

It’s hard to tell because the writing quality is garbage.

Second paragraph:

Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token.

This is a stretch. True, but barely. The LLM is making slightly different choices near the end of the token generation process.

At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection.

Claim support, if it appears, is pages later.

At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it.

?

Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools.

Wtf. Non-sequitor. Where does this come from?

Such a watermarking procedure can therefore affect both what the model says and what an agent does.

Duh? In the literal sense of outputting different tokens.

We call this behavioral effect sampling drift.

I think they should have used an LLM for writing help, or paid more for the one they used.

29m agoHN ↗

In _1984_ the Big Brother regime has the idea that by controlling language you can influence what is possible to think, and thus becomes a key tool of political repression.

Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.

Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.

All of these attempts to control language are fundamentally misguided at best, often have severe unintended consequences, and are genuinely immoral at worse.

15m agoHN ↗

I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

14m agoHN ↗

This article reads like it was written at least partly by AI to me. Specifically it reads like an article written by AI with edits made by a human further prompting the AI.

Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.

Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.

Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".

Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.

The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.