Text watermark checker

Reveal hidden characters, map their positions, and clean a copy of your text.

0 words · 0 characters

PDF, DOCX, TXT or MD · Up to 10 MB · Or drag a file here

Inspect up to 100,000 characters.

Your analysis

Ready when you are.

Every character, in view.

Reveal hidden characters, explore their positions, and choose what to clean.

How to read your results

What does the inspector look for?

The inspector reads the characters in pasted text or text extracted from your uploaded document. It reports invisible formatting, nonstandard spaces, bidirectional controls, Unicode tags, joiners, variation selectors and combining marks. A limited mixed-script check can also flag lookalike characters inside otherwise Latin words. A character map shows code-point positions, and the inventory lists each character's category and surrounding context.

None of that is automatically a watermark. A nonbreaking space can come from typography. A joiner can be essential to a language or an emoji. A copied passage can hold ordinary formatting and an intentionally encoded payload at the same time, so where the text came from matters as much as what's in it.

Can this detect Google SynthID?

Not by itself. SynthID Text works through statistical choices made while the model generates text. It isn't a layer of invisible Unicode, so a character scrubber can't strip it out. Detection needs a compatible detector, plus the right configuration and tokenizer. Google's public implementation is useful for systems whose operators set it up, but it doesn't hand anyone Gemini's private production keys.

This site keeps local character inspection separate from optional remote watermark verification. If no authorized verifier is connected, the Google/SynthID and Claude checks show as "not tested." If a verifier does report a signal, the interface shows its scheme and evidence. An exact model name appears only when the verifier supplies one, and the site never guesses it from writing style. For the longer explanation, see the text watermark guide.

How do I clean invisible characters from text?

After a scan, pick the character families you want to clean. The preview lists every individual edit and leaves your original untouched in the input box. A zero-width space becomes an ordinary space so words don't join unexpectedly, and nonstandard spaces become ordinary spaces or line breaks. Optional NFC normalization composes canonically equivalent characters, for example an accented letter stored as a base letter plus a combining accent.

Some families start unselected. Joiners, bidi controls and variation selectors affect language and rendering, so they stay off by default. Unicode tags start selected, but they also belong to regional-flag emoji, so check the copy before you keep it. The tool doesn't blindly delete combining marks or swap lookalike letters. Your cleaned download is a character-level cleanup, and it isn't certified removal of a statistical watermark.

Can it decode a hidden message?

It tries two narrow encodings. One is byte-aligned runs of U+200B and U+200C read as binary UTF-8, tried with both bit polarities. The other is Unicode tag characters mapped to ASCII. A successful decode is a candidate payload, not an authenticated message, and other encodings may be present without being recognized.

The distribution map splits the passage into 64 character ranges. Regular intervals between findings can help you spot repeating patterns, but a tidy pattern may just be a formatting convention, and the counts can't establish intent on their own. Offsets are Unicode code points, which differ from both bytes and visual grapheme clusters.

Can I check a PDF or Word document?

Upload a text-based PDF, a Word DOCX file, or a TXT or Markdown document up to 10 MB. PDFs support up to 100 pages. Review the extracted text before inspecting it: a document format can alter spaces, line breaks or invisible characters during extraction.

This is a text inspection tool. It does not scan image watermarks, embedded file metadata, document signatures or image-only PDFs. For the most faithful character-level check, use the original plain text.