for some reason people seem to have Anthropic Derangement Syndrome over a pretty straightforward combination of cryptographic primitives and LLM sampling techniques. so as someone who isn't an employee, let me try to explain what's going on with watermarking step by step.
- LLMs work by probabilistically autoregressively sampling tokens. meaning: at every token, the model weights don't output a single next token, but a probability distribution over *all possible tokens*, a bunch of little numbers that sum to 1
- the "temperature" sampling setting affects how this distribution is constructed. at 1, it's just the "ground truth" / whatever the model thinks. shifting it above 1 will make the distribution much more fat-tailed, lower probability tokens will be higher probability etc. shifting it towards 0 makes it closer to deterministic, making the highest probability tokens much more likely to be selected, and at 0 always just selecting the single most probable option
- modern reasoning models almost exclusively use temperature 1, and apis often no longer even expose it as a customizable setting, so "deterministic generation" isn't common
- when you "pick something randomly" on a computer (not just LLMs), almost always it's actually *pseudorandom*, eg using a complex algorithm that based on an initial seed number, generates a chain of numbers that has nice cryptographically provable properties: the distribution of the numbers has no pattern, no information content, and can't be predicted better than chance by any method except having the initial seed number and in fact running the same pseudorandom algorithm. that means you *can* deterministically generate the same "randomness" if you have the same seed, but no one else can tell the difference between that and real randomness
- often these pesudorandom number generators are initially seeded either with something like time, or for more security using a hardware randomness source, something that samples physical temperature or the like on-chip. those sources are too slow and expensive otherwise use for all randomness
- pre watermarking, when Anthropic generated Claude tokens it would use pseudorandom generators to select those tokens from the LLM distribution, with the generators seeded in an ~unknown but generic way. very likely just some system default which is one of the above, but this is completely opaque to the end user. as mentioned, *tokens are already selected randomly*, it is not the case that LLMs always pick the most likely next token, that in fact is very undesirable and leads to much lower quality generations
- post-watermarking, the only thing that changes is *how the seed is selected*. now, it's always *seeded* with a deterministic hash of the prior tokens plus the secret key. the selection is still *psuedorandom*, with exactly the same properties described above. it's still the case that cryptographically, at each given token, if you look at the full token probability distribution the LLM outputs and see which one the sampler selects, its actual choices are indistinguishable from having used a source of physical randomness *unless you know the seed*. it doesn't cut out certain words from the vocabulary, it doesn't "affect phrasing", any more than the existing system of random selection already does
- the only difference is now, it's possible for Anthropic to take a sequence of words, run each token through Claude to get the LLM probabilities per token, then seed the same pseudorandom number generator in the same way with their secret keys, and *check whether the token selected matches the one the pesudorandom generator would have selected when seeded in that way*. the property they're taking advantage of is that PRNGs are in fact deterministic, so if this exact setup was used to generate the tokens then all the selections will match exactly, and this will be wildly wildly implausible / virtually zero probability on any meaningful sequence of words
- what this will not do: it won't (and can't) differentiate between very very very overdetermined content. for example if you prompt an LLM with "What is 1 + 1? output nothing besides the numerical integer answer", then the token distribution is: 2 with probability 99.9999%, and then every other token in the universe with negligible probability. no matter how you seed your PRNG (which again, all LLMs already use for sampling), all the probability mass is on 2! the randomness is used for selection weighted by probability, and there aren't any other probable choices here
- but it really doesn't require very many token choices for this to come up, because the vast majority of English is *not* overdetermined. you can personally inspect the LLM output distributions in various playgrounds and see how many often quite close to equal probability tokens are sitting near the top. as mentioned these are *already* being selected between, never just selecting the single most probable, so within a sentence or two of normal output there will be overwhelming (but undetectable without the private key) watermark evidence
- also, they're using Google's SynthID algorithm for this, and *gemini already does exactly this and has for like a year and a half*. basically all text you've gotten from gemini has already gone through exactly this process!
hopefully that clears things up somewhat!
Show more