David Knichel, PhD
← All notes

AI · Technical notes

Technical foundations of text watermarking

How SynthID-Text creates a statistical signal through token selection, and how a detector distinguishes that signal from chance.

Suppose a language model can continue a sentence with either “fast” or “quick”. Both may be suitable, but only one will appear in the output. Text watermarking uses this freedom of selection to introduce a reproducible statistical preference. Across a sufficiently long passage, the preference can provide evidence that the text was generated with a particular watermark configuration.

The intuition is to divide the vocabulary into two groups using the recent context and a watermark key. For one watermark layer, tokens receive either score 0 or score 1. Tournament sampling favors the score-one group, shifting the probability distribution toward its tokens. The groups change with the context; they do not describe the meaning or quality of the words.

For detection, only the resulting text is observed. Using the matching configuration, the detector recomputes which group each observed token belongs to and collects evidence across the sequence. A hypothesis test checks whether the preference for the favored groups is stronger than expected without the matching watermark. Unwatermarked generation does not use these groups to select tokens, although chance alone can produce an apparent preference.

SynthID-Text applies this idea across multiple layers, each with its own group assignment. Later layers can favor a different candidate, so the final token need not have score 1 in every layer. The detector combines evidence across usable positions and layers.

This article develops that idea for Google DeepMind’s SynthID-Text. It assumes familiarity with discrete probabilities and expectations; the hypothesis test is introduced explicitly. Other members of the SynthID family use different mechanisms for images, audio, and video. DeepMind’s SynthID overview.

1. From model probabilities to token selection

A tokenizer maps text to a sequence of token identifiers. A token may represent a word, a word fragment, or punctuation. For clarity, the examples below treat each displayed word as a single token in an artificial four-token vocabulary. The numerical probabilities are illustrative model outputs.

At position tt, the language model assigns a probability to each candidate token, conditional on the preceding sequence. Let pt(v)p_t(v) denote the probability of selecting token vv. Ordinary generation samples Xt∼ptX_t\sim p_t, appends the selected token, and repeats with the updated context.

A watermark modifies this selection procedure. It transforms ptp_t into a distribution qtq_t using a watermark configuration and the recent context. The model supplies the candidate probabilities; the watermark introduces a preference within those choices. Once the next token has been emitted, the procedure starts again for the new context.

Generation:
  preceding text → model probabilities → watermark sampling → next token

Detection:
  observed text + matching configuration → reconstructed scores → test

Generation requires the language model to compute candidate probabilities. Detection reconstructs the scores of tokens that actually occurred; it does not need to reconstruct the full probability distribution or run the original model. This distinction is central to SynthID-Text’s design. Dathathri et al., 2024.

Where do the model probabilities come from?

Let VV denote the vocabulary and zt(v)z_t(v) the model’s logit, an unnormalized numerical preference for candidate vv. With temperature τ>0\tau>0, the softmax operation gives

pt(v)=exp⁡(zt(v)/τ)∑u∈Vexp⁡(zt(u)/τ).p_t(v)=\frac{\exp(z_t(v)/\tau)}{\sum_{u\in V}\exp(z_t(u)/\tau)}.

The denominator ensures that probabilities sum to one. A smaller temperature concentrates probability on candidates with larger logits. A sampler may also restrict the candidate set, for example to the most probable tokens. In the remainder of the article, ptp_t denotes the distribution after these operations. The watermark analysis starts from this distribution and does not depend on how the model computed its logits.

2. A worked example of one tournament layer

Consider one generation step. We omit the position index tt and write pp for the original distribution. A score function assigns each candidate a binary value g(v)∈{0,1}g(v)\in\{0,1\}:

Candidate tokenOriginal probability p(v)Score g(v)
fast0.401
quick0.300
rapid0.201
swift0.100

The score assignment is fixed for this context and configuration. It does not mean that “fast” is generally preferred to “quick”. A different context can produce a different assignment.

A binary tournament performs three operations:

  1. Sample two candidates independently from pp, with replacement.
  2. If their scores differ, select the candidate with score 1.
  3. If the scores are equal, select either candidate with probability 1/21/2.

For example, “fast” defeats “quick”. A match between “fast” and “rapid” is resolved randomly. Drawing “quick” twice still produces “quick”: score-zero tokens are not forbidden.

The total preference for score 1

Let aa denote the probability that an ordinary draw has score 1:

a=∑v∈Vp(v)g(v)=0.40+0.20=0.60.a=\sum_{v\in V}p(v)g(v)=0.40+0.20=0.60.

Only a match between two score-zero candidates produces a score-zero winner. Since the draws are independent, this event has probability (1−a)2(1-a)^2. Therefore,

Pr⁡(g(X)=1)=1−(1−a)2=2a−a2=0.84,\Pr(g(X)=1)=1-(1-a)^2=2a-a^2=0.84,

where XX is the selected output token. The tournament raises the total probability of score 1 from 0.600.60 to 0.840.84. This increase creates evidence that can accumulate across generation steps.

The probability of an individual token

Let q(v)q(v) denote the probability that the tournament outputs token vv. This differs from p(v)p(v), which is the probability of drawing vv as a candidate before the tournament selects a winner.

The output can be vv in either of two ways:

  1. The first candidate is vv, and the tournament selects the first candidate.
  2. The second candidate is vv, and the tournament selects the second candidate.

Consider a token vv with score 1. The first candidate is vv with probability p(v)p(v). Its opponent has score 0 with probability 1−a1-a, in which case the first candidate wins with certainty. Otherwise, the opponent has score 1, and the first candidate wins the tie with probability 1/21/2. Therefore,

Pr⁡(first candidate is v and is selected)=p(v)((1−a)+a2).\Pr(\text{first candidate is }v\text{ and is selected}) =p(v)\left((1-a)+\frac{a}{2}\right).

The second candidate is drawn from the same distribution and follows the same selection rules, so its corresponding probability has the same value. The tournament selects exactly one candidate position, so these two events cannot both occur. We can therefore add their probabilities:

q(v)=p(v)((1−a)+a2)+p(v)((1−a)+a2)=p(v)(2−a).\begin{aligned} q(v) &=p(v)\left((1-a)+\frac{a}{2}\right) +p(v)\left((1-a)+\frac{a}{2}\right)\\ &=p(v)(2-a). \end{aligned}

This also works when both candidates are vv. The tie-break selects one position or the other, each with probability 1/21/2, but the output token is vv either way. Adding the two selection probabilities counts this outcome once, not twice.

For a token vv with score 0, either position can only win when the opponent also has score 0 and the tie-break selects that position. Each position contributes p(v)(1−a)/2p(v)(1-a)/2, giving

q(v)=p(v)1−a2+p(v)1−a2=p(v)(1−a).q(v)=p(v)\frac{1-a}{2}+p(v)\frac{1-a}{2}=p(v)(1-a).

Combining the two cases gives

q(v)=p(v)(1+g(v)−a).\boxed{q(v)=p(v)\bigl(1+g(v)-a\bigr).}

For “fast”, q=0.40×1.40=0.56q=0.40\times1.40=0.56. For “quick”, q=0.30×0.40=0.12q=0.30\times0.40=0.12. The other probabilities become 0.280.28 and 0.040.04, respectively, and all four still sum to one. In general, normalization follows from ∑vq(v)=1+a−a=1\sum_vq(v)=1+a-a=1. This update matches the binary-layer calculation in DeepMind’s reference implementation.

One tournament layer

Exact four-token example
Original pWatermarked q
Token probabilities before and after one tournament layer, with fixed context scores00.51fastg = 1quickg = 0rapidg = 1swiftg = 0
fast: 40.0% → 56.0%quick: 30.0% → 12.0%rapid: 20.0% → 28.0%swift: 10.0% → 4.0%

Total probability of score 1: 60.0% → 84.0%. The vertical axis shows token probability.

Figure 1. Scores are fixed for a given context and key. Dark bars show the conditional bias. The switch uses two illustrative assignments, not a real keyed hash. A concentrated distribution limits the available selection freedom.

Explore Figure 1. First compare the two bars for each token. Then switch the score assignment: the preference follows the scores, not the words. Finally, enable the low-entropy distribution. Here “low entropy” means that almost all probability is concentrated on one candidate. Even when that candidate receives score 0, it remains likely because both tournament draws usually select it.

3. Multiple layers and reproducible scores

SynthID-Text extends this procedure to multiple layers, each with a different pseudorandom score function. A conceptual binary tournament with mm layers starts with 2m2^m candidate draws. Pairwise winners advance until a single output remains. SynthID-Text paper, Fig. 2.

For two layers, suppose the candidates are “fast”, “quick”, “rapid”, and “fast”. The repeated occurrence of “fast” is valid because sampling is performed with replacement. Under the score assignment above, the first pair selects “fast”. The second pair is a tie; suppose it selects “rapid”. In the second layer, a new function assigns g2(fast)=0g_2(\text{fast})=0 and g2(rapid)=1g_2(\text{rapid})=1, so “rapid” is emitted.

Tournament sampling, step by step

18-second animation
Candidate sampling → layer 1 → layer 2 → token emission → score reconstruction
Animation 1. A two-layer example of binary tournament sampling. Scores and tie outcomes are illustrative. Use Play or select a stage to inspect the process. Detection subsequently aggregates reconstructed scores over multiple positions.
Inspect the same tournament as a static diagram

Two layers, four candidates

Illustrative bracket

Draw independently from p

fast · g₁ = 1
quick · g₁ = 0
rapid · g₁ = 1
fast · g₁ = 1

Layer 1 · use g₁

fast defeats quick
rapid wins tie with fast

Layer 2 · use g₂

fast · g₂ = 0
rapid · g₂ = 1
Emit rapid
Figure 2. Candidate tokens may repeat. Ties are random. Each layer uses a different score function; a winner need not have a 1 in every layer.

The two matches in the first layer produce independent winners with the same updated distribution. Consequently, the next layer can be analyzed by applying the single-layer calculation to that distribution. There is no need to construct an exponentially large candidate pool in an implementation. Starting from p(0)=pp^{(0)}=p, compute

aℓ=∑vp(ℓ−1)(v)gℓ(v),p(ℓ)(v)=p(ℓ−1)(v)(1+gℓ(v)−aℓ),\begin{aligned} a_\ell&=\sum_v p^{(\ell-1)}(v)g_\ell(v),\\ p^{(\ell)}(v)&=p^{(\ell-1)}(v)\bigl(1+g_\ell(v)-a_\ell\bigr), \end{aligned}

then draw the output token from p(m)p^{(m)}. The superscript denotes the layer, not a power. Each aℓa_\ell must be recomputed using the distribution entering that layer. A token can win a tie with score 0, so the final winner need not score 1 in every layer.

Pseudorandom during generation, reproducible during detection

A pseudorandom score is deterministic once its inputs are fixed. Conceptually, those inputs are the recent token sequence, the candidate token, and the configuration for the layer. The same inputs produce the same score. The key supplies the reference needed to reconstruct the assignment; a score is not freshly chosen by an unrelated random draw during detection.

The public implementation uses recent token identifiers and layer-specific keys. Its n-gram length controls the context required for score computation. The tokenizer and watermark configuration must match between generation and detection. DeepMind and Hugging Face, configuration guide.

Detection reconstructs these assignments from the finished text. Section 5 follows that process from individual tokens to a statistical decision.

4. Why conditional bias can preserve marginal probabilities

Figure 1 changes token probabilities. How can such a mechanism also be described as non-distortionary? The answer depends on what is held fixed and what is averaged over.

For a fixed score assignment, qq generally differs from pp. Now consider an idealized ensemble in which every distinct token receives an independent fair score bit. Holding pp fixed and averaging over these assignments gives

Eg[q(v)]=p(v)(1+E[g(v)]−∑up(u)E[g(u)])=p(v)(1+12−12)=p(v).\begin{aligned} \mathbb{E}_g[q(v)] &=p(v)\left(1+\mathbb{E}[g(v)]-\sum_u p(u)\mathbb{E}[g(u)]\right)\\ &=p(v)(1+\tfrac12-\tfrac12)=p(v). \end{aligned}

Thus, the marginal token probabilities are preserved in this ensemble. The dependency remains visible to someone who knows the score assignment used for each selection. Preserving a marginal distribution does not imply independence between the output and the scores.

A two-token example makes this distinction explicit. Let both tokens have probability 1/21/2. The four equally likely assignments are (0,0)(0,0), (0,1)(0,1), (1,0)(1,0), and (1,1)(1,1). The first token’s output probabilities are, respectively, 1/21/2, 1/41/4, 3/43/4, and 1/21/2. Their average is still 1/21/2, although each mixed assignment introduces a preference.

This is an average over the idealized score-generation randomness. It is not a guarantee that repeated generation with one fixed key and one repeated context reproduces pp. The paper distinguishes token-level and sequence-level non-distortion and introduces additional handling for context reuse. Section “Preserving the quality of generative text”.

Why concentrated distributions carry less signal

If one token has probability 1, both candidate draws are always that token. The tournament cannot affect the output, regardless of its score. More generally, frequent duplicate draws reduce the selection freedom available for embedding the watermark.

Derive the signal strength from the collision probability

Let C=∑vp(v)2C=\sum_vp(v)^2 denote the probability that two independent draws select the same token. Under independent fair score bits, E[a]=1/2\mathbb{E}[a]=1/2 and

Var⁡(a)=∑vp(v)2Var⁡(g(v))=C4.\operatorname{Var}(a)=\sum_vp(v)^2\operatorname{Var}(g(v))=\frac{C}{4}.

Consequently, E[a2]=Var⁡(a)+E[a]2=(1+C)/4\mathbb{E}[a^2]=\operatorname{Var}(a)+\mathbb{E}[a]^2=(1+C)/4. Substituting into the winning-score probability yields

Eg[2a−a2]=34−14C.\mathbb{E}_g[2a-a^2]=\frac34-\frac14C.

For KK equally likely tokens, C=1/KC=1/K, so the expected winning score approaches 0.750.75 as KK increases. For a deterministic distribution, C=1C=1, and the expectation is 0.50.5. The derivation accounts for identical scores on repeated occurrences of the same token.

5. From the observed text to a detection decision

During generation, the watermark favors tokens from groups determined by the context and key. Detection asks whether this preference left a measurable trace in the finished text. It reconstructs the group membership of the tokens that occurred, counts the evidence, and compares it with what could occur without the matching watermark.

We first work through a small example with one layer. Its scores are invented to make the calculation transparent. The hypothesis test uses an idealized probability model; afterward, we distinguish that model from a practical SynthID detector.

From text to a detection decision

1 min 40 sec · Captioned video
Video 2. A one-layer example with invented scores and an idealized binomial test. All explanations appear on screen; there is no audio. The final chapter distinguishes the toy model from practical calibration. The steps below provide the full text explanation. Download MP4.

Step 1: Recompute a score for each observed token

Suppose the detector receives this sentence:

The small team built a fast model that runs well on ordinary office computers.

For this example, treat each word as one token and ignore punctuation. Let the score function use the two preceding tokens, the current token, and one fixed key kk. The first two words provide context, leaving 12 positions to score.

Write

yt=g(xt,ct;k)∈{0,1},y_t=g(x_t,c_t;k)\in\{0,1\},

where xtx_t is the token at position tt and ctc_t is its preceding context. A score of 1 means that this token belongs to the group favored by this layer in that context. Here is an illustrative result:

Preceding context cₜObserved token xₜComputed score yₜ
The smallteam1
small teambuilt1
team builta0
built afast1
a fastmodel1
fast modelthat1
model thatruns1
that runswell1
runs wellon1
well onordinary1
on ordinaryoffice1
ordinary officecomputers1

For the first row, the detector evaluates g(team,The small;k)=1g(\text{team},\text{The small};k)=1. It then moves one position forward and evaluates g(built,small team;k)=1g(\text{built},\text{small team};k)=1. The same operation produces every remaining row.

The text contains no visible scores. The detector computes them from the text and configuration. For a fixed passage and configuration, repeating the calculation gives exactly the same values. There is no tournament during detection, and the detector does not recover the candidates that lost during generation. It only evaluates the tokens present in the passage.

Summarize the result by the number of usable scores nn and the number of ones ss:

n=12,s=∑tyt=11.n=12,\qquad s=\sum_t y_t=11.

Eleven of the twelve observed tokens belong to their respective favored groups. The remaining question is whether that could reasonably happen without watermarking.

Step 2: Define the reference case without a matching watermark

The numbers 0.60.6 and 0.50.5 refer to different random experiments. Distinguishing them explains both the null model and its limits.

First, hold the context and score assignment fixed, and draw a token. In our four-token example, the favored group contains “fast” and “rapid”. Their original probabilities sum to

a=p(fast)+p(rapid)=0.4+0.2=0.6.a=p(\text{fast})+p(\text{rapid})=0.4+0.2=0.6.

An unwatermarked draw therefore lands in this group with probability 0.60.6. One tournament layer raises its output probability to 2a−a2=0.842a-a^2=0.84. The value aa still denotes the original mass, 0.60.6; 0.840.84 is the mass after watermarking. If we repeatedly sampled this same context with this same assignment, the appropriate unwatermarked baseline would be 0.60.6, not 0.50.5.

Now hold the model probabilities fixed and vary the score assignment. Under the idealized model, each token independently receives score 1 with probability 1/21/2. Some assignments favor likely tokens and others favor unlikely ones. Reversing the scores in our example would favor “quick” and “swift”, whose probabilities sum to 0.3+0.1=0.40.3+0.1=0.4. The two complementary assignments have masses 0.60.6 and 0.40.4, averaging to 0.50.5.

More generally, averaging over all idealized assignments gives

Eg[a]=Eg ⁣[∑vp(v)g(v)]=∑vp(v) 12=12.\mathbb{E}_g[a] =\mathbb{E}_g\!\left[\sum_v p(v)g(v)\right] =\sum_v p(v)\,\frac12 =\frac12.

Thus, 0.50.5 comes from the fair assignment of scores, even when word probabilities are highly unequal. It does not require either group to contain exactly half the vocabulary or half the probability mass for a particular assignment.

For detection, imagine fixing the entire unwatermarked passage first, then assigning independent fair bits to its distinct context-token pairs. Each token already in the passage has probability 1/21/2 of receiving score 1. Under this idealization, our twelve distinct pairs give twelve independent fair bits. Watermarked generation breaks this independence between text and assignments because its choices favor score-one candidates.

In an actual detector, the key is fixed. It recomputes the assignments; it does not try new keys or flip coins. As the context changes along the passage, the score function receives different inputs. Modeling its outputs on unwatermarked text as independent fair bits is the approximation that connects the idealized experiment to the detector. An average score of 0.50.5 alone does not establish independence, and the absence of watermarking does not guarantee that every fixed configuration follows this model exactly.

With that assumption explicit, define the null hypothesis H0H_0 as generation without the matching watermark, modeled here by independent fair scores. The alternative hypothesis H1H_1 models a preference for score 1:

H0:Yi∼Bernoulli⁡(0.5),H1:Yi∼Bernoulli⁡(p1),p1>0.5.\begin{aligned} H_0 &: Y_i\sim\operatorname{Bernoulli}(0.5),\\ H_1 &: Y_i\sim\operatorname{Bernoulli}(p_1),\qquad p_1>0.5. \end{aligned}

A Bernoulli variable is 1 with the specified probability and 0 otherwise. We assume independence under both hypotheses for this worked example. The capital YiY_i describes a hypothetical score before its value is known; the lowercase yiy_i is a value already computed in our table.

The calculation below is exact under this probability model. For a deployed detector, the false-positive rate must be checked with its fixed configuration on representative unwatermarked passages, accounting for usable length and dependencies. Calibration sets a threshold from the resulting reference distribution; it does not force each context’s probability mass to become 0.50.5. DeepMind’s guidance on detector thresholds.

Step 3: Calculate how unusual the observed count is

Under H0H_0, the total number of ones is

S=∑i=112Yi∼Binomial⁡(12,0.5).S=\sum_{i=1}^{12}Y_i\sim\operatorname{Binomial}(12,0.5).

The expected count is 12⋅0.5=612\cdot0.5=6. A count above six is not sufficient evidence by itself, since chance produces variation. We ask how often the null model produces at least the observed count of eleven.

There are 212=40962^{12}=4096 equally likely binary score sequences under this model. Exactly twelve contain eleven ones: the single zero can occupy any of twelve positions. One sequence contains twelve ones. Hence,

Pr⁡H0(S≥11)=Pr⁡H0(S=11)+Pr⁡H0(S=12)=(1211)+(1212)212=12+14096≈0.00317.\begin{aligned} \Pr_{H_0}(S\ge11) &=\Pr_{H_0}(S=11)+\Pr_{H_0}(S=12)\\ &=\frac{\binom{12}{11}+\binom{12}{12}}{2^{12}}\\ &=\frac{12+1}{4096}\approx0.00317. \end{aligned}

This is the one-sided p-value: approximately 0.317%0.317\%. It includes twelve ones because that outcome would provide still stronger evidence in the same direction. We test for an excess of ones because generation favors that group.

The p-value answers a conditional question: assuming H0H_0 and the stated model, how often would the count be this large or larger? It does not assign a 0.317%0.317\% probability to the claim that the passage is unwatermarked. That would reverse the conditioning.

Step 4: Apply a decision rule chosen in advance

Choose a false-positive rate of at most α=0.01\alpha=0.01, or 1%1\%, before inspecting the result. A false positive occurs when this test flags text generated under H0H_0.

For a count-based test, select the smallest threshold kk such that

Pr⁡H0(S≥k)≤α.\Pr_{H_0}(S\ge k)\le\alpha.

For twelve scores, the threshold is k=11k=11. We already calculated that eleven or more ones occur with probability 13/4096≈0.317%13/4096\approx0.317\%. Lowering the threshold to ten would give

Pr⁡H0(S≥10)=66+12+14096≈1.929%,\Pr_{H_0}(S\ge10) =\frac{66+12+1}{4096} \approx1.929\%,

which exceeds the permitted 1%1\%. Because counts are discrete, the achievable false-positive rate can be smaller than the chosen limit.

Our observed count is s=11s=11, so the test flags the passage. With s=10s=10, it would not flag it at this threshold, even though ten is above the expected count of six. A result below the threshold means that the test has insufficient evidence; it does not prove that no watermark is present.

The complete calculation is now visible: the text and configuration yield twelve scores; their sum is eleven; the null model assigns a probability of about 0.317%0.317\% to a count at least that large; the result crosses the detection threshold for the preselected 1%1\% limit. This is a demonstration using invented scores, not evidence that this short sentence would be detectable by SynthID.

Step 5: Understand what longer passages change

The same test works for another number of independent scores by recalculating the threshold. For n=200n=200, the null count has mean 100 and standard deviation 200/4≈7.07\sqrt{200/4}\approx7.07. At a false-positive limit of 1%1\%, the threshold becomes k=117k=117. An observed count of s=120s=120 gives

Pr⁡H0(S≥120)=∑j=120200(200j)2−200≈0.00284.\Pr_{H_0}(S\ge120) =\sum_{j=120}^{200}\binom{200}{j}2^{-200} \approx0.00284.

Here, a fraction of 120/200=0.6120/200=0.6 suffices. Longer sequences let a modest preference accumulate into stronger evidence. The 11-of-12 result above was deliberately extreme to keep the first calculation small.

We can also ask how often watermarked text crosses the threshold. This probability is the detection power. To compute it in our model, we must specify the alternative probability p1p_1. For example, p1=0.6p_1=0.6 means that each score is 1 with probability 0.60.6 under H1H_1. It does not mean that the passage has a 60%60\% probability of being watermarked.

The null-model p-value and threshold require no choice of p1p_1. The alternative is needed to calculate power and compare how reliably different signal strengths are detected.

Statistical watermark detection

Independent Bernoulli model
H₀: no matching watermarkH₁: matching watermark
Binomial probability mass distributions over computed fractions of score-one tokens. Detection power 69.4 percent.0.30.40.50.60.70.80.9Probability mass (vertical scale adapts)Computed fraction S / n
Detect at S ≥ 117False positives 0.97%Detection power 69.4%

The vertical dashed line is the threshold. The shaded region contains counts that trigger a detection. The triangle marks the observed passage.

Threshold reached. Computed fraction: 0.600. One-sided p-value: 0.00284.

The p-value is a null-model tail probability, not a probability that this passage is watermarked. Changing n preserves the approximate computed fraction.

Figure 3. Exact binomial probabilities, with adjacent points joined for readability. The shaded area is the detection region, chosen for a false-positive rate of at most 1%. These results refer to the illustrative statistical model, not measured SynthID performance. Randomization independently replaces score observations with fair coins; it does not simulate real text editing.

Explore Figure 3 in this order. Start with n=200n=200 and the computed count s=120s=120. The triangle is the result for one passage, and the shaded region contains counts that trigger detection. Move only the count slider to cross the threshold of 117. This changes the decision for that passage, while the reference distributions stay fixed.

Next, change the alternative probability p1p_1. The solid distribution moves and detection power changes; the null distribution and threshold stay fixed. Finally, increase nn while keeping p1p_1 fixed. The fractions concentrate more tightly around their respective means, making the preference easier to distinguish. These curves describe repeated hypothetical outcomes, rather than uncertainty about the scores already computed for one passage.

What does the score-randomization control represent?

The control replaces each score with an independent fair bit with probability rr. Under the alternative, the resulting probability of score 1 is

peffective=(1−r)p1+r⋅0.5.p_{\mathrm{effective}}=(1-r)p_1+r\cdot0.5.

At r=1r=1, the two hypotheses produce the same score distribution, so detection power equals the false-positive rate. This is a model of lost score information. It is not a percentage of edited words: a real edit can change tokenization and the contexts of several subsequent tokens.

From the small test to a practical SynthID detector

SynthID reconstructs a score for each usable token under each layer, producing a matrix of scores. The same underlying idea applies: combine evidence of the preference and compare it with a reference for text without the matching watermark. The public reference implementation includes mean, weighted-mean, and trained Bayesian scoring methods. DeepMind’s detector documentation.

The simple binomial calculation above does not automatically apply to every entry of that matrix. Repeated contexts can repeat the same evidence, and scores can be dependent across positions and layers. Positions without sufficient preceding context and repeated contexts are excluded using masks. Counting tokens times layers as independent trials can overstate the evidence when those independence assumptions fail.

A practical detector therefore needs a scoring method and threshold appropriate to its configuration and usable sequence length. Calibrating a threshold means checking how often representative unwatermarked text exceeds it; testing on matching watermarked text then measures detection power. DeepMind discusses length-dependent thresholds, while its Bayesian detector requires training. Reference detector guidance.

The small example establishes the logic of detection. The reconstruction, masking, aggregation, and calibration determine whether that logic gives reliable decisions in a real implementation.

Which inputs and parameters are needed to implement detection?

The detector needs the passage and the configuration used to assign scores, followed by a statistical decision rule:

Required itemPurpose
Matching tokenizerRecover the token identifiers used by the score function.
Layer keys and context lengthReconstruct each token’s score under each layer.
Exact score-function configurationMatch the hashing procedure and any lookup-table parameters.
Masking rulesDetermine which positions contribute usable evidence.
Scoring method and calibrated thresholdAggregate the evidence and decide whether to flag the passage. A trained detector also needs its learned parameters.

The public Hugging Face implementation uses keys, ngram_len, sampling_table_seed, and sampling_table_size for score reconstruction. An ngram_len of 5 includes four preceding tokens and the token being scored. Repetition masking also uses context_history_size; end-token handling depends on the tokenizer. These names belong to that implementation. Score-function source.

Detection does not require the original model weights, its probability distributions pp and qq, the per-context mass aa, or the tournament candidates. All scored token-context pairs come from the passage itself. A passage extracted without its original prompt can still provide usable positions after its initial context window.

The key alone is insufficient if tokenization or hashing differs. The public Transformers implementation documents a hashing difference from the Gemini app, so its example configuration does not automatically detect Gemini’s watermark. Implementation note.

6. Interpreting the result

A positive result provides statistical evidence of a matching watermark configuration. It does not identify the person who submitted the prompt, establish the truth of the content, or prove that the passage is unedited. A negative result may reflect an absent watermark, insufficient evidence, or degradation of the signal.

Short passages provide few observations. Highly constrained continuations provide little selection freedom. Paraphrasing and translation can modify both selected tokens and their contexts. These are distinct reasons for reduced detection capability. Official SynthID-Text limitations.

The false-positive rate also differs from the fraction of detected passages that are incorrectly flagged. The latter depends on how often matching watermarks occur in the population being examined. Consider 10,000 passages, of which 1% contain a matching watermark. At 90% detection power and a 1% false-positive rate:

  • 90 of the 100 watermarked passages are detected.
  • 99 of the 9,900 unmarked passages are incorrectly flagged.
  • Of the 189 flagged passages, only about 48% contain a matching watermark.

Equivalently, with prevalence π\pi, detection power β\beta, and false-positive rate α\alpha, Bayes’ theorem gives

Pr⁡(H1∣flag)=βπβπ+α(1−π).\Pr(H_1\mid\text{flag})= \frac{\beta\pi}{\beta\pi+\alpha(1-\pi)}.

A small false-positive rate can therefore coexist with a limited positive predictive value. The figures in this example describe an illustrative population, not measured SynthID performance.

The complete mechanism can now be followed from generation to detection: the model supplies alternative tokens, tournament sampling introduces a dependency on reproducible scores, and a detector evaluates the accumulated evidence against a calibrated reference. The distinction between those three steps is essential for interpreting what a watermark result establishes.

References and scope

This article describes the publicly documented binary tournament method. It does not specify undisclosed details of the current Gemini deployment. Figures and worked examples are original explanatory constructions. The detection experiment uses explicitly simplified assumptions.

← Back to all notesDiscuss this note ↗