All articles
SynthIDAI detectionAI watermarkingAI contentSEOAEOGEO

How Google SynthID detects AI-generated content: the science, formulas and limitations

RankNexus·October 9, 2026 9 min read
How Google SynthID detects AI-generated content: the science, formulas and limitations

Quick answer: Google SynthID detects compatible AI-generated text by looking for an invisible statistical watermark embedded during token selection. Its detector uses the matching watermark configuration and calibrated statistical scoring. No watermark found does not mean a human wrote the text.

What is Google SynthID?

AI can now write an article, produce a convincing voiceover, create a realistic image or generate a video from a short prompt. That raises a simple question: how can we tell where digital content came from?

Google DeepMind developed SynthID to help answer that question. SynthID is a watermarking technology designed to identify content generated by supported AI systems. Rather than judging whether language sounds robotic, it embeds an invisible, verifiable signal during generation.

For text, the signal is statistical: a compatible language model subtly changes how it selects tokens. For images and videos, watermark signals are built into the generated visual data. For audio, the signals are designed to stay inaudible.

On October 7, 2026, Google announced expanded access to SynthID Detector for supported image, video and audio verification. Google reported more than 180 billion images and videos watermarked, and around 240,000 years of audio. These are watermarking scale figures, not accuracy rates. The announcement does not establish that unrestricted text detection is available on the public site.

How SynthID Text works

A language model generates text one token at a time. Tokens can be words, pieces of words or punctuation. At each step it calculates probabilities for the possible next tokens.

Take the sentence "The future of online discovery depends on..." The next token might be relevance, trust, context or search. A SynthID-enabled generator uses private-key-dependent functions to influence how these candidates are selected. The text reads normally, yet over many token selections a detectable statistical pattern emerges.

The detector needs the corresponding watermark scheme and configuration to recognise that pattern. It is not checking for a banned-word list, perfect grammar or a stock set of AI phrases.

GENERATIONVERIFICATION published, shared, edited Promptuser request Model scorescandidate tokens Private keys setg-values (0 or 1) Tournament tiltsthe token choice Watermarkedtext Text to checkfrom any source Detector withthe matching keys Mean g-valuescore S(x) Calibratedthreshold Detected Not detected Inconclusive Without the matching keys and configuration, a detector cannot test for the watermark at all.
SynthID Text in two halves: a keyed bias added while the text is generated, and a keyed statistical test run later.

The mathematics, explained step by step

Stage A: base token probability

P(x_i) = exp(z_i / T) / Σ_j exp(z_j / T)

Here z_i is the model's logit for token i, and T is the sampling temperature. Suppose the base candidate probabilities are:

TokenBase probability
Search0.40
Content0.30
Visibility0.20
Ranking0.10

Stage B: keyed pseudorandom g-values

The watermark derives context-dependent values using private keys. A simplified layer uses g_i ∈ {0,1}. For illustration only:

Tokeng-value
Search1
Content0
Visibility1
Ranking0

These assignments come from the keyed computation, not from the meaning of the words.

Stage C: tournament probability adjustment

An equivalent single-layer probability update from the published reference method is:

G    = Σ_i p_i · g_i
p'_i = p_i · (1 + g_i − G)

For the hypothetical example, G = 0.40 + 0.20 = 0.60. The updated probabilities become:

TokenBeforeAfter
Search40%56%
Content30%12%
Visibility20%28%
Ranking10%4%
Illustrative example: one tournament layer, G = 0.60 BeforeAfter the keyed tilt 0%20%40%60% 40%30%20%10% 56%12%28%4% SearchContentVisibilityRanking g = 1g = 0g = 1g = 0
The keyed tilt moves probability toward tokens with g = 1. Both columns still sum to 100%. Illustrative numbers, not Google production output.

The probabilities still sum to 100%. This shows the mathematical bias; it is not an example of actual Google production output. Multiple watermark layers and context controls make real deployments more complex.

Stage D: mean watermark detection score

S(x) = [1 / (mN)] · Σ_t Σ_l g_l(x_t, r_t)

N is the number of analysed token positions, m is the number of layers, and r_t is the context-dependent seed. Under a simplified unwatermarked null model the expected g-value is around 0.5; stronger signals may indicate watermarking. But a score of 0.60 is not "60% AI-written".

For masked positions, a more appropriate expression is:

S_masked = [Σ_t Σ_l M_t · g_l(x_t, r_t)] / [m · Σ_t M_t]

M_t excludes positions, such as specified repeated contexts, that should not contribute independent evidence.

Stage E: Bayesian interpretation

                 P(G | W) · P(W)
P(W | G) = ─────────────────────────────────────────
           P(G | W) · P(W) + P(G | not W) · P(not W)

Here W is the watermark hypothesis and G is the observed evidence. The result depends on the model's likelihood estimates, the prior and calibration. It is not a universal percentage of machine authorship. A good detector should allow an uncertain result.

Key configuration parameters

ParameterPurpose
keysSecret integers driving the watermark functions
ngram_lenToken context used to derive watermark information
sampling_table_sizePseudorandom lookup table capacity
sampling_table_seedReproducibility of pseudorandom sampling
context_history_sizeControl of repeated n-gram contexts
Watermark depthNumber of keyed layers
TokenizerCompatibility with the original token representation
Detector thresholdsCalibrated decisions and error trade-offs

Google's documentation recommends ngram_len = 5 as a useful default and a sampling_table_size of at least 2^16 (65,536). These recommendations do not reveal every Google production setting. Keys and deployment-specific thresholds are not public.

What can and cannot be detected

A detector may recognise text generated by a compatible watermarked system. It cannot be assumed to detect every ChatGPT, Claude, Gemini or other model output. A different model may use no watermark, a different watermark, or keys the detector does not have.

Three outcomes need three different labels:

ResultWhat it means
Watermark detectedEvidence of a compatible watermark signal.
Watermark not detectedNo compatible signal found. Origin is unresolved.
InconclusiveNot enough evidence for a reliable conclusion.

Detection is probabilistic, so false positives and false negatives are possible. Confidence generally benefits from more usable token positions, but no fixed word count guarantees a result. Highly predictable or short factual answers may offer weak watermarking opportunities.

Does editing change the result?

It can. Minor word changes, cropping and mild paraphrasing may preserve detectable evidence. Extensive rewriting or translation can reduce the signal because the token sequences change. There is no universal number of edits that guarantees a watermark disappears.

More importantly, watermark manipulation is not editorial improvement. Rewriting an inaccurate article does not make it credible unless the facts and substance are corrected. Our companion guide, how to humanise AI-generated content, covers what real editorial improvement looks like.

Images, video and audio

SynthID also supports compatible image, video and audio workflows. The signal is embedded in the media data itself, not merely in EXIF metadata or a filename. Google describes resilience to some ordinary transformations, including selected forms of compression, cropping and editing, but not to every possible operation.

When a mixed-media video contains camera footage, animation and AI-generated inserts, a watermark finding does not prove that every frame is AI-generated. Conclusions should stay limited to the detected evidence.

Does Google Search penalise AI content?

Google's published Search guidance does not say that AI-assisted content is automatically penalised. It stresses useful, reliable, people-first content. Producing low-value pages at scale mainly to manipulate rankings may violate Google's scaled content abuse policy.

SynthID watermark detection is not an established universal Google ranking factor. A watermark cannot tell you whether an article is original, accurate, well sourced or useful. Likewise, an undetected watermark does not imply strong SEO performance.

What this means for SEO, AEO and GEO

SEO helps important, relevant pages be found and understood in search. Answer Engine Optimization (AEO) emphasises clear responses to real questions. Generative Engine Optimization (GEO) emphasises distinctive, verifiable information that AI answer experiences may find useful to reference, which is why query fan-out matters so much.

Across all three, the right content workflow asks:

  • Are the important claims fact-checked?
  • Does the page answer a real user need?
  • Does it contain original perspective or evidence?
  • Are sources authoritative and correctly attributed?
  • Can users and crawlers reach the useful information?
  • Is the author or editor accountable for the final result?

No editing tactic guarantees inclusion in AI-generated answers. The objective is content worth referencing, not content engineered to game a detection score.

A responsible verification workflow

  1. Determine whether the tool measures an embedded watermark or only predicts linguistic style.
  2. Establish whether the suspected source model and its watermark scheme are supported.
  3. Check how many usable tokens, frames or audio segments are available.
  4. Preserve uncertainty, and never equate "not detected" with "human-written".
  5. Review accuracy, originality and relevance separately from provenance.
  6. Document tool versions, detection configurations and reviewer sign-off.

The RankNexus perspective

At RankNexus.ai, we view content quality as broader than whether AI helped write a page. A useful article needs accuracy, clarity, relevant examples, original information and a purpose beyond filling a publishing calendar.

Watermark detection helps with provenance. SEO, AEO and GEO evaluate different aspects of usefulness and discoverability. Folding those signals into a fictional "AI percentage" would mislead publishers.

The better question is simple: does this page deserve to be found?

Conclusion

SynthID adds verifiable evidence to the generation process instead of relying only on guesses about how AI-written content sounds. That is valuable, but it has real limits: unsupported systems, edited media, short text and unavailable keys can all constrain verification.

For publishers, the larger opportunity is unchanged. Publish information that is accurate, distinctive, understandable and genuinely useful. Measure provenance and content quality separately.

Next in this series: How to humanise AI-generated content in 2026: a complete SEO, AEO and GEO guide.

Sources and recommended reading

  1. Google: Google expands SynthID Detector for AI content (October 7, 2026)
  2. Google DeepMind: Watermarking AI-generated text and video with SynthID
  3. Google AI for Developers: SynthID tools for watermarking and detecting LLM-generated text
  4. Nature: Scalable watermarking for identifying large language model outputs
  5. GitHub: google-deepmind/synthid-text
  6. Google DeepMind: SynthID
  7. Google Search Central: Using generative AI content
  8. Google Search Central: Creating helpful, reliable, people-first content

Token probabilities in this article are illustrative, not test output or Google production settings.

Frequently asked questions

How does SynthID identify AI-written text?

Compatible generators embed keyed token-selection patterns, and calibrated detectors test for those patterns.

Does SynthID detect every ChatGPT article?

No. A compatible watermark and detector configuration must be available.

Does SynthID examine grammar and sentence length?

Not as its primary watermark mechanism. Its published text approach is keyed and statistical.

Is a negative result proof of human authorship?

No. It only indicates that no compatible watermark was detected under that test.

Is SynthID completely accurate?

No. Performance depends on calibration, text characteristics, implementation and transformations.

Can short text be detected?

Sometimes, but short or highly constrained text may provide limited evidence.

Can publishers implement SynthID Text?

Google provides open-source code for compatible watermarking and verification configurations, not Google's production keys.

Does a detected watermark cause an SEO penalty?

Google's public Search guidance does not establish a general ranking penalty based solely on SynthID watermark presence.

Want this done for your site?

Run a free audit and see exactly what to fix for Google and AI search.

Run a free audit