What are the three layers?
Character-level artifacts live in the string itself: a zero-width character, a nonstandard space, a Unicode tag. Statistical watermarks live in the pattern of token choices a model made while generating. File provenance sits beside the content in a file structure, often as a signed record. The word "invisible" fits all three, and that's about all they share.
Copy a document as plain text and you keep the characters and the word choices but lose the file metadata. A Unicode cleanup changes characters and leaves the words alone. A full rewrite changes the word choices and leaves the subject. What survives depends on the mechanism and on what was done to the text.
What does SynthID Text do?
Google's documentation describes SynthID Text as a watermarking component in the generation pipeline. A keyed pseudorandom function influences which tokens get sampled. The configuration includes keys and context-related settings, and detection uses a compatible detector. Google has published implementations that developers can run in their own systems, and the method appears in a Nature paper from Google DeepMind.
Publishing the code doesn't let an arbitrary third-party script authenticate any Gemini passage. A detector has to match the system being checked, and having source code for an algorithm is a long way from having a provider's production configuration and verification access. That's why the inspector never labels a narrow space "SynthID."
How does Claude's text watermark work?
Anthropic describes its text watermark as a version of the SynthID Text approach and says it adds no hidden characters. The model still chooses among near-equivalent words, but a key stands in for the plain random number, and only someone holding that key can check the pattern. Anthropic says it's introducing the watermark in future Claude models to comply with the EU AI Act, with older models covered over a transition period.
Its documentation separates watermark verification from a classifier that guesses whether prose looks AI-generated. Detection runs through an API in private preview, open to eligible organizations rather than the general public. Anthropic also notes that short passages carry little signal, that a complete rewrite removes the watermark, and that a hit can show Claude was involved without separating "Claude wrote this" from "Claude heavily edited this."
Two consequences follow. A missing signal doesn't establish human authorship, and a detected one doesn't mean the provider originated every idea or sentence. This site follows the same line: its classifier never invents a Claude model name, and its character inspector doesn't claim access to any private verification key.
How does a keyed statistical watermark work?
Picture a research system that uses a secret key and the previous word to split the possible next words into two groups. While generating, it leans slightly toward one group. A matching detector repeats the grouping and counts whether the observed word choices favor that group more than you'd expect under a defined null model. Nothing gets inserted between the letters.
That's an intuition for keyed statistical testing, not a specification of SynthID. Real systems make decisions about tokenizers, context windows, repeated-context handling and calibration, and changing any one of them can leave a plausible-looking detector incompatible with its generator.
A detector also needs enough informative observations. Repeating the same keyed context isn't necessarily independent evidence. The Python research lab that accompanies this project uses an explicitly synthetic pair-based scheme so those assumptions can be inspected, and its output isn't treated as a production provider detector.
What can a Unicode scan tell you?
Unicode includes characters that exist only to control rendering. A joiner can combine emoji components. A non-joiner can change how letters connect. Bidirectional controls affect the order in which text displays, and a variation selector picks a particular glyph. Ordinary writing systems depend on all of these.
A scanner can report exactly where each one sits and what it's called. From a code point alone, it can't tell whether someone inserted the character as a watermark, an editor added it during formatting, or the language requires it. Judging that takes context.
The sample text supplied with the watermark checker contains a deliberately encoded binary payload next to legitimate Unicode. It demonstrates the controls and doesn't show what any AI provider emits. The decoder tries a few explicit conventions, and failing to decode something is no evidence that nothing is encoded.
What does cleaning Unicode actually remove?
Judge selective cleanup at the character level: record which categories you picked, count the edits and compare the output. Look for words that joined unexpectedly, a language that stopped rendering correctly, and emoji that changed. Keep the original.
Canonical normalization can change the code-point sequence while leaving its conventional meaning the same. It helps with some text-processing jobs and it's no universal watermark eraser. Compatibility folding is broader and can remove distinctions that matter in some contexts.
For a statistical scheme, the useful question is what a matching detector reports before and after the change, under the same configuration. If that detector isn't available, the honest result is "not verified," even when every unusual character is gone.
Can a watermark tell you which model wrote the text?
Not necessarily. A provider-level mark doesn't always identify a model release. Several systems may share one marking configuration, and a document can pass through more than one system. Naming a specific model requires the verifier to output that distinction explicitly. A guess based on tone or a familiar phrase doesn't count.
The optional verifier interface on this site can report a provider, a scheme, a result, evidence and a model name when one is supplied. Without a model field, the display says "model not supplied." Without a configured verifier, the statistical check stays untested. A high AI score from a third-party classifier fills neither gap.
Which check should you run next?
Use local Unicode inspection when your question is about the string itself. Use an authorized statistical verifier for a specific generation-time mark. Use a provenance-aware file tool to read a signed file credential, and keep the original file, not just copied text, if provenance matters to you.
For experiments, state the scheme, configuration, tokenizer, detector and transformation before you report success or failure. If you only scanned characters, call it a character scan. Precise wording keeps the work reproducible and stops a narrow observation from growing into a claim about who wrote something.