AI · Technical notes
Technical foundations of text watermarking
How SynthID-Text creates a statistical signal through token selection, and how a detector distinguishes that signal from chance.
Suppose a language model can continue a sentence with either “fast” or “quick”. Both may be suitable, but only one will appear in the output. Text watermarking uses this freedom of selection to introduce a reproducible statistical preference. Across a sufficiently long passage, the preference can provide evidence that the text was generated with a particular watermark configuration.
The intuition is to divide the vocabulary into two groups using the recent context and a watermark key. For one watermark layer, tokens receive either score 0 or score 1. Tournament sampling favors the score-one group, shifting the probability distribution toward its tokens. The groups change with the context; they do not describe the meaning or quality of the words.
For detection, only the resulting text is observed. Using the matching configuration, the detector recomputes which group each observed token belongs to and collects evidence across the sequence. A hypothesis test checks whether the preference for the favored groups is stronger than expected without the matching watermark. Unwatermarked generation does not use these groups to select tokens, although chance alone can produce an apparent preference.
SynthID-Text applies this idea across multiple layers, each with its own group assignment. Later layers can favor a different candidate, so the final token need not have score 1 in every layer. The detector combines evidence across usable positions and layers.
This article develops that idea for Google DeepMind’s SynthID-Text. It assumes familiarity with discrete probabilities and expectations; the hypothesis test is introduced explicitly. Other members of the SynthID family use different mechanisms for images, audio, and video. DeepMind’s SynthID overview.
1. From model probabilities to token selection
A tokenizer maps text to a sequence of token identifiers. A token may represent a word, a word fragment, or punctuation. For clarity, the examples below treat each displayed word as a single token in an artificial four-token vocabulary. The numerical probabilities are illustrative model outputs.
At position , the language model assigns a probability to each candidate token, conditional on the preceding sequence. Let denote the probability of selecting token . Ordinary generation samples , appends the selected token, and repeats with the updated context.
A watermark modifies this selection procedure. It transforms into a distribution using a watermark configuration and the recent context. The model supplies the candidate probabilities; the watermark introduces a preference within those choices. Once the next token has been emitted, the procedure starts again for the new context.
Generation:
preceding text → model probabilities → watermark sampling → next token
Detection:
observed text + matching configuration → reconstructed scores → test
Generation requires the language model to compute candidate probabilities. Detection reconstructs the scores of tokens that actually occurred; it does not need to reconstruct the full probability distribution or run the original model. This distinction is central to SynthID-Text’s design. Dathathri et al., 2024.
Where do the model probabilities come from?
Let denote the vocabulary and the model’s logit, an unnormalized numerical preference for candidate . With temperature , the softmax operation gives
The denominator ensures that probabilities sum to one. A smaller temperature concentrates probability on candidates with larger logits. A sampler may also restrict the candidate set, for example to the most probable tokens. In the remainder of the article, denotes the distribution after these operations. The watermark analysis starts from this distribution and does not depend on how the model computed its logits.
2. A worked example of one tournament layer
Consider one generation step. We omit the position index and write for the original distribution. A score function assigns each candidate a binary value :
| Candidate token | Original probability p(v) | Score g(v) |
|---|---|---|
| fast | 0.40 | 1 |
| quick | 0.30 | 0 |
| rapid | 0.20 | 1 |
| swift | 0.10 | 0 |
The score assignment is fixed for this context and configuration. It does not mean that “fast” is generally preferred to “quick”. A different context can produce a different assignment.
A binary tournament performs three operations:
- Sample two candidates independently from , with replacement.
- If their scores differ, select the candidate with score 1.
- If the scores are equal, select either candidate with probability .
For example, “fast” defeats “quick”. A match between “fast” and “rapid” is resolved randomly. Drawing “quick” twice still produces “quick”: score-zero tokens are not forbidden.
The total preference for score 1
Let denote the probability that an ordinary draw has score 1:
Only a match between two score-zero candidates produces a score-zero winner. Since the draws are independent, this event has probability . Therefore,
where is the selected output token. The tournament raises the total probability of score 1 from to . This increase creates evidence that can accumulate across generation steps.
The probability of an individual token
Let denote the probability that the tournament outputs token . This differs from , which is the probability of drawing as a candidate before the tournament selects a winner.
The output can be in either of two ways:
- The first candidate is , and the tournament selects the first candidate.
- The second candidate is , and the tournament selects the second candidate.
Consider a token with score 1. The first candidate is with probability . Its opponent has score 0 with probability , in which case the first candidate wins with certainty. Otherwise, the opponent has score 1, and the first candidate wins the tie with probability . Therefore,
The second candidate is drawn from the same distribution and follows the same selection rules, so its corresponding probability has the same value. The tournament selects exactly one candidate position, so these two events cannot both occur. We can therefore add their probabilities:
This also works when both candidates are . The tie-break selects one position or the other, each with probability , but the output token is either way. Adding the two selection probabilities counts this outcome once, not twice.
For a token with score 0, either position can only win when the opponent also has score 0 and the tie-break selects that position. Each position contributes , giving
Combining the two cases gives
For “fast”, . For “quick”, . The other probabilities become and , respectively, and all four still sum to one. In general, normalization follows from . This update matches the binary-layer calculation in DeepMind’s reference implementation.
One tournament layer
Exact four-token exampleTotal probability of score 1: 60.0% → 84.0%. The vertical axis shows token probability.
Explore Figure 1. First compare the two bars for each token. Then switch the score assignment: the preference follows the scores, not the words. Finally, enable the low-entropy distribution. Here “low entropy” means that almost all probability is concentrated on one candidate. Even when that candidate receives score 0, it remains likely because both tournament draws usually select it.
3. Multiple layers and reproducible scores
SynthID-Text extends this procedure to multiple layers, each with a different pseudorandom score function. A conceptual binary tournament with layers starts with candidate draws. Pairwise winners advance until a single output remains. SynthID-Text paper, Fig. 2.
For two layers, suppose the candidates are “fast”, “quick”, “rapid”, and “fast”. The repeated occurrence of “fast” is valid because sampling is performed with replacement. Under the score assignment above, the first pair selects “fast”. The second pair is a tie; suppose it selects “rapid”. In the second layer, a new function assigns and , so “rapid” is emitted.
Tournament sampling, step by step
18-second animationInspect the same tournament as a static diagram
Two layers, four candidates
Illustrative bracketDraw independently from p
Layer 1 · use g₁
Layer 2 · use g₂
The two matches in the first layer produce independent winners with the same updated distribution. Consequently, the next layer can be analyzed by applying the single-layer calculation to that distribution. There is no need to construct an exponentially large candidate pool in an implementation. Starting from , compute
then draw the output token from . The superscript denotes the layer, not a power. Each must be recomputed using the distribution entering that layer. A token can win a tie with score 0, so the final winner need not score 1 in every layer.
Pseudorandom during generation, reproducible during detection
A pseudorandom score is deterministic once its inputs are fixed. Conceptually, those inputs are the recent token sequence, the candidate token, and the configuration for the layer. The same inputs produce the same score. The key supplies the reference needed to reconstruct the assignment; a score is not freshly chosen by an unrelated random draw during detection.
The public implementation uses recent token identifiers and layer-specific keys. Its n-gram length controls the context required for score computation. The tokenizer and watermark configuration must match between generation and detection. DeepMind and Hugging Face, configuration guide.
Detection reconstructs these assignments from the finished text. Section 5 follows that process from individual tokens to a statistical decision.
4. Why conditional bias can preserve marginal probabilities
Figure 1 changes token probabilities. How can such a mechanism also be described as non-distortionary? The answer depends on what is held fixed and what is averaged over.
For a fixed score assignment, generally differs from . Now consider an idealized ensemble in which every distinct token receives an independent fair score bit. Holding fixed and averaging over these assignments gives
Thus, the marginal token probabilities are preserved in this ensemble. The dependency remains visible to someone who knows the score assignment used for each selection. Preserving a marginal distribution does not imply independence between the output and the scores.
A two-token example makes this distinction explicit. Let both tokens have probability . The four equally likely assignments are , , , and . The first token’s output probabilities are, respectively, , , , and . Their average is still , although each mixed assignment introduces a preference.
This is an average over the idealized score-generation randomness. It is not a guarantee that repeated generation with one fixed key and one repeated context reproduces . The paper distinguishes token-level and sequence-level non-distortion and introduces additional handling for context reuse. Section “Preserving the quality of generative text”.
Why concentrated distributions carry less signal
If one token has probability 1, both candidate draws are always that token. The tournament cannot affect the output, regardless of its score. More generally, frequent duplicate draws reduce the selection freedom available for embedding the watermark.
Derive the signal strength from the collision probability
Let denote the probability that two independent draws select the same token. Under independent fair score bits, and
Consequently, . Substituting into the winning-score probability yields
For equally likely tokens, , so the expected winning score approaches as increases. For a deterministic distribution, , and the expectation is . The derivation accounts for identical scores on repeated occurrences of the same token.
5. From the observed text to a detection decision
During generation, the watermark favors tokens from groups determined by the context and key. Detection asks whether this preference left a measurable trace in the finished text. It reconstructs the group membership of the tokens that occurred, counts the evidence, and compares it with what could occur without the matching watermark.
We first work through a small example with one layer. Its scores are invented to make the calculation transparent. The hypothesis test uses an idealized probability model; afterward, we distinguish that model from a practical SynthID detector.
From text to a detection decision
1 min 40 sec · Captioned videoStep 1: Recompute a score for each observed token
Suppose the detector receives this sentence:
The small team built a fast model that runs well on ordinary office computers.
For this example, treat each word as one token and ignore punctuation. Let the score function use the two preceding tokens, the current token, and one fixed key . The first two words provide context, leaving 12 positions to score.
Write
where is the token at position and is its preceding context. A score of 1 means that this token belongs to the group favored by this layer in that context. Here is an illustrative result:
| Preceding context cₜ | Observed token xₜ | Computed score yₜ |
|---|---|---|
| The small | team | 1 |
| small team | built | 1 |
| team built | a | 0 |
| built a | fast | 1 |
| a fast | model | 1 |
| fast model | that | 1 |
| model that | runs | 1 |
| that runs | well | 1 |
| runs well | on | 1 |
| well on | ordinary | 1 |
| on ordinary | office | 1 |
| ordinary office | computers | 1 |
For the first row, the detector evaluates . It then moves one position forward and evaluates . The same operation produces every remaining row.
The text contains no visible scores. The detector computes them from the text and configuration. For a fixed passage and configuration, repeating the calculation gives exactly the same values. There is no tournament during detection, and the detector does not recover the candidates that lost during generation. It only evaluates the tokens present in the passage.
Summarize the result by the number of usable scores and the number of ones :
Eleven of the twelve observed tokens belong to their respective favored groups. The remaining question is whether that could reasonably happen without watermarking.
Step 2: Define the reference case without a matching watermark
The numbers and refer to different random experiments. Distinguishing them explains both the null model and its limits.
First, hold the context and score assignment fixed, and draw a token. In our four-token example, the favored group contains “fast” and “rapid”. Their original probabilities sum to
An unwatermarked draw therefore lands in this group with probability . One tournament layer raises its output probability to . The value still denotes the original mass, ; is the mass after watermarking. If we repeatedly sampled this same context with this same assignment, the appropriate unwatermarked baseline would be , not .
Now hold the model probabilities fixed and vary the score assignment. Under the idealized model, each token independently receives score 1 with probability . Some assignments favor likely tokens and others favor unlikely ones. Reversing the scores in our example would favor “quick” and “swift”, whose probabilities sum to . The two complementary assignments have masses and , averaging to .
More generally, averaging over all idealized assignments gives
Thus, comes from the fair assignment of scores, even when word probabilities are highly unequal. It does not require either group to contain exactly half the vocabulary or half the probability mass for a particular assignment.
For detection, imagine fixing the entire unwatermarked passage first, then assigning independent fair bits to its distinct context-token pairs. Each token already in the passage has probability of receiving score 1. Under this idealization, our twelve distinct pairs give twelve independent fair bits. Watermarked generation breaks this independence between text and assignments because its choices favor score-one candidates.
In an actual detector, the key is fixed. It recomputes the assignments; it does not try new keys or flip coins. As the context changes along the passage, the score function receives different inputs. Modeling its outputs on unwatermarked text as independent fair bits is the approximation that connects the idealized experiment to the detector. An average score of alone does not establish independence, and the absence of watermarking does not guarantee that every fixed configuration follows this model exactly.
With that assumption explicit, define the null hypothesis as generation without the matching watermark, modeled here by independent fair scores. The alternative hypothesis models a preference for score 1:
A Bernoulli variable is 1 with the specified probability and 0 otherwise. We assume independence under both hypotheses for this worked example. The capital describes a hypothetical score before its value is known; the lowercase is a value already computed in our table.
The calculation below is exact under this probability model. For a deployed detector, the false-positive rate must be checked with its fixed configuration on representative unwatermarked passages, accounting for usable length and dependencies. Calibration sets a threshold from the resulting reference distribution; it does not force each context’s probability mass to become . DeepMind’s guidance on detector thresholds.
Step 3: Calculate how unusual the observed count is
Under , the total number of ones is
The expected count is . A count above six is not sufficient evidence by itself, since chance produces variation. We ask how often the null model produces at least the observed count of eleven.
There are equally likely binary score sequences under this model. Exactly twelve contain eleven ones: the single zero can occupy any of twelve positions. One sequence contains twelve ones. Hence,
This is the one-sided p-value: approximately . It includes twelve ones because that outcome would provide still stronger evidence in the same direction. We test for an excess of ones because generation favors that group.
The p-value answers a conditional question: assuming and the stated model, how often would the count be this large or larger? It does not assign a probability to the claim that the passage is unwatermarked. That would reverse the conditioning.
Step 4: Apply a decision rule chosen in advance
Choose a false-positive rate of at most , or , before inspecting the result. A false positive occurs when this test flags text generated under .
For a count-based test, select the smallest threshold such that
For twelve scores, the threshold is . We already calculated that eleven or more ones occur with probability . Lowering the threshold to ten would give
which exceeds the permitted . Because counts are discrete, the achievable false-positive rate can be smaller than the chosen limit.
Our observed count is , so the test flags the passage. With , it would not flag it at this threshold, even though ten is above the expected count of six. A result below the threshold means that the test has insufficient evidence; it does not prove that no watermark is present.
The complete calculation is now visible: the text and configuration yield twelve scores; their sum is eleven; the null model assigns a probability of about to a count at least that large; the result crosses the detection threshold for the preselected limit. This is a demonstration using invented scores, not evidence that this short sentence would be detectable by SynthID.
Step 5: Understand what longer passages change
The same test works for another number of independent scores by recalculating the threshold. For , the null count has mean 100 and standard deviation . At a false-positive limit of , the threshold becomes . An observed count of gives
Here, a fraction of suffices. Longer sequences let a modest preference accumulate into stronger evidence. The 11-of-12 result above was deliberately extreme to keep the first calculation small.
We can also ask how often watermarked text crosses the threshold. This probability is the detection power. To compute it in our model, we must specify the alternative probability . For example, means that each score is 1 with probability under . It does not mean that the passage has a probability of being watermarked.
The null-model p-value and threshold require no choice of . The alternative is needed to calculate power and compare how reliably different signal strengths are detected.
Statistical watermark detection
Independent Bernoulli modelThe vertical dashed line is the threshold. The shaded region contains counts that trigger a detection. The triangle marks the observed passage.
Threshold reached. Computed fraction: 0.600. One-sided p-value: 0.00284.
The p-value is a null-model tail probability, not a probability that this passage is watermarked. Changing n preserves the approximate computed fraction.
Explore Figure 3 in this order. Start with and the computed count . The triangle is the result for one passage, and the shaded region contains counts that trigger detection. Move only the count slider to cross the threshold of 117. This changes the decision for that passage, while the reference distributions stay fixed.
Next, change the alternative probability . The solid distribution moves and detection power changes; the null distribution and threshold stay fixed. Finally, increase while keeping fixed. The fractions concentrate more tightly around their respective means, making the preference easier to distinguish. These curves describe repeated hypothetical outcomes, rather than uncertainty about the scores already computed for one passage.
What does the score-randomization control represent?
The control replaces each score with an independent fair bit with probability . Under the alternative, the resulting probability of score 1 is
At , the two hypotheses produce the same score distribution, so detection power equals the false-positive rate. This is a model of lost score information. It is not a percentage of edited words: a real edit can change tokenization and the contexts of several subsequent tokens.
From the small test to a practical SynthID detector
SynthID reconstructs a score for each usable token under each layer, producing a matrix of scores. The same underlying idea applies: combine evidence of the preference and compare it with a reference for text without the matching watermark. The public reference implementation includes mean, weighted-mean, and trained Bayesian scoring methods. DeepMind’s detector documentation.
The simple binomial calculation above does not automatically apply to every entry of that matrix. Repeated contexts can repeat the same evidence, and scores can be dependent across positions and layers. Positions without sufficient preceding context and repeated contexts are excluded using masks. Counting tokens times layers as independent trials can overstate the evidence when those independence assumptions fail.
A practical detector therefore needs a scoring method and threshold appropriate to its configuration and usable sequence length. Calibrating a threshold means checking how often representative unwatermarked text exceeds it; testing on matching watermarked text then measures detection power. DeepMind discusses length-dependent thresholds, while its Bayesian detector requires training. Reference detector guidance.
The small example establishes the logic of detection. The reconstruction, masking, aggregation, and calibration determine whether that logic gives reliable decisions in a real implementation.
Which inputs and parameters are needed to implement detection?
The detector needs the passage and the configuration used to assign scores, followed by a statistical decision rule:
| Required item | Purpose |
|---|---|
| Matching tokenizer | Recover the token identifiers used by the score function. |
| Layer keys and context length | Reconstruct each token’s score under each layer. |
| Exact score-function configuration | Match the hashing procedure and any lookup-table parameters. |
| Masking rules | Determine which positions contribute usable evidence. |
| Scoring method and calibrated threshold | Aggregate the evidence and decide whether to flag the passage. A trained detector also needs its learned parameters. |
The public Hugging Face implementation uses keys, ngram_len, sampling_table_seed, and sampling_table_size for score reconstruction. An ngram_len of 5 includes four preceding tokens and the token being scored. Repetition masking also uses context_history_size; end-token handling depends on the tokenizer. These names belong to that implementation. Score-function source.
Detection does not require the original model weights, its probability distributions and , the per-context mass , or the tournament candidates. All scored token-context pairs come from the passage itself. A passage extracted without its original prompt can still provide usable positions after its initial context window.
The key alone is insufficient if tokenization or hashing differs. The public Transformers implementation documents a hashing difference from the Gemini app, so its example configuration does not automatically detect Gemini’s watermark. Implementation note.
6. Interpreting the result
A positive result provides statistical evidence of a matching watermark configuration. It does not identify the person who submitted the prompt, establish the truth of the content, or prove that the passage is unedited. A negative result may reflect an absent watermark, insufficient evidence, or degradation of the signal.
Short passages provide few observations. Highly constrained continuations provide little selection freedom. Paraphrasing and translation can modify both selected tokens and their contexts. These are distinct reasons for reduced detection capability. Official SynthID-Text limitations.
The false-positive rate also differs from the fraction of detected passages that are incorrectly flagged. The latter depends on how often matching watermarks occur in the population being examined. Consider 10,000 passages, of which 1% contain a matching watermark. At 90% detection power and a 1% false-positive rate:
- 90 of the 100 watermarked passages are detected.
- 99 of the 9,900 unmarked passages are incorrectly flagged.
- Of the 189 flagged passages, only about 48% contain a matching watermark.
Equivalently, with prevalence , detection power , and false-positive rate , Bayes’ theorem gives
A small false-positive rate can therefore coexist with a limited positive predictive value. The figures in this example describe an illustrative population, not measured SynthID performance.
The complete mechanism can now be followed from generation to detection: the model supplies alternative tokens, tournament sampling introduces a dependency on reproducible scores, and a detector evaluates the accumulated evidence against a calibrated reference. The distinction between those three steps is essential for interpreting what a watermark result establishes.
References and scope
This article describes the publicly documented binary tournament method. It does not specify undisclosed details of the current Gemini deployment. Figures and worked examples are original explanatory constructions. The detection experiment uses explicitly simplified assumptions.
- Dathathri et al. (2024), Scalable watermarking for identifying large language model outputs. Research paper and technical supplement.
- Google DeepMind’s SynthID-Text reference implementation. Generation, score reconstruction, masking, and detectors. The README notes that the reference hash does not guarantee cryptographic security.
- Google DeepMind and Hugging Face, Introducing SynthID Text. Configuration, integration, training, and limitations.
- Google DeepMind, SynthID. The broader watermarking family.