Try a small example with a real invisible character
Copy this short example into the inspector: redblue
An actual U+200B ZERO WIDTH SPACE separates "red" and "blue". Written visibly, the sequence is red[U+200B]blue. Those explanatory brackets are not part of the sample.
The inspector should report U+200B at offset 3, line 1, column 4. There are eight code points in the sample, including the invisible character. An ordinary search for "redblue" may behave differently across applications because each application decides how to handle formatting characters.
With Zyphh's default cleanup selection, the result is "red blue". The replacement is an ordinary space. It does not silently join the two pieces into "redblue".
Try "10 kg" without its quotation marks too. The gap is U+00A0 NO-BREAK SPACE. Compare it before and after cleanup. Both samples were constructed for this guide and have no claimed AI origin.
Recognize the common code points
The name and code point are more useful than a count of "suspicious characters". These commonly encountered characters have different jobs:
| Code point | Name | Usual purpose |
|---|---|---|
| U+200B | ZERO WIDTH SPACE | Supplies a break opportunity without a visible space |
| U+00A0 | NO-BREAK SPACE | Keeps adjacent text together on a line |
| U+2060 | WORD JOINER | Prevents a break without adding visible width |
| U+200C | ZERO WIDTH NON-JOINER | Requests separation in letter joining or ligatures |
| U+200D | ZERO WIDTH JOINER | Requests joined rendering |
| U+FEFF | ZERO WIDTH NO-BREAK SPACE | Commonly serves as a byte order mark in encoded files |
Context still matters for U+FEFF: a file signature and a character inside text need different handling. The Unicode layout controls specification explains these functions.
Repeated findings can come from formatting, language requirements or deliberate encoding. Their count does not distinguish those causes.
Read offsets as code points, not screen positions
Zyphh reports zero-based code-point offsets. The first stored character has offset 0. Lines and columns start at 1, and columns also count code points rather than the visible width of a glyph. The methodology records those units.
Consider this constructed string: A😀BC
Its U+200B sits after B at code-point offset 3 and column 4. The emoji counts as one code point here, although JavaScript represents it with two UTF-16 code units and UTF-8 uses four bytes. Check the units before using a position reported by another program.
Visible characters can also contain several code points. The accented ending of "café" contains an e followed by U+0301 COMBINING ACUTE ACCENT. A reader generally sees one accented letter. Unicode calls units used for this kind of text segmentation grapheme clusters. Unicode text segmentation describes the distinction.
Report the code point and surrounding text with the offset so someone else can locate it.
Keep language and emoji characters when they are needed
The Persian sample "میروم" contains U+200C. Deleting that character changes the joining behavior of the surrounding letters. Preserve it when the language requires it.
The sample "👩💻" contains U+1F469, U+200D and U+1F4BB. A supporting font can render that sequence as a woman technologist. Remove the joiner and the separate woman and computer symbols may appear instead. The heart "❤️" includes U+FE0F, a variation selector requesting emoji presentation.
Unicode also uses tag sequences for some flags. Tags are not automatically hidden messages. The Unicode emoji specification documents these sequences and presentation controls.
Zyphh leaves the joiner and variation-selector cleanup groups unselected by default. Unicode tags are selected by default, with a warning that removing them can change regional-flag emoji. Review every selected category when your text contains flags or writing you cannot confidently read.
What bidirectional controls mean
Bidirectional controls influence the display order of text containing right-to-left and left-to-right content. That includes ordinary combinations of Arabic or Hebrew with numbers, punctuation or English words. The stored order and displayed order are not always the same. Unicode's bidirectional algorithm defines the behavior.
Check U+200F RIGHT-TO-LEFT MARK and U+2067 RIGHT-TO-LEFT ISOLATE in context. Some controls have matching terminators; deleting one side can affect later text.
This guide spells those controls out instead of embedding active direction overrides in examples. Zyphh's context view uses visible code-point labels for categorized characters, which helps you examine them without relying only on their rendered appearance.
The bidi cleanup group starts unselected. If the original mixes writing directions, compare the cleaned text in the application where it will actually be read.
Choose cleanup categories for the problem you found
Open Preview selective cleanup after inspecting a passage. The default categories are invisible formatting, nonstandard spaces, Unicode tags and control characters. Review that selection against the actual findings before proceeding.
The cleanup rules are specific. U+200B becomes an ordinary space. Nonstandard spacing characters become ordinary spaces, while Unicode line and paragraph separators become newline characters. Other selected characters are removed. Tabs, carriage returns and line feeds are excluded from the control-character cleanup category.
Joiners, bidirectional controls, variation selectors and other formatting are opt-in. The tool does not offer blanket deletion of combining marks or automatic replacement of mixed-script lookalikes.
Preview the output before downloading it. In "redblue", a space replacement preserves a boundary. Removing a selected U+2060 between two letters simply removes that character; it does not insert a new space. Check both cases against what the text is supposed to say.
If the preview looks wrong, return to the character inspector with the original and change one category at a time.
Understand what NFC normalization changes
Canonical normalization is a separate checkbox. The sequence "café" ends with e plus U+0301; "café" ends with the single code point U+00E9. NFC can compose the first representation into the second. The text can look the same while its stored sequence and character count change.
NFC does not generally fold compatibility characters into simpler substitutes. For example, the circled digit "①" and the ligature "fi" remain in this tool's NFC output. NFKC is a broader normalization form with different tradeoffs. Unicode normalization describes the forms.
Zyphh applies NFC after selected category edits. Its change log lists the category edits, and a separate flag reports whether normalization changed the sequence. The edit count therefore does not include each separate normalization change.
Compare the exported source and result when exact text representation matters. A smaller finding count is not, by itself, a better document.
Treat decoded payloads as hypotheses
The inspector tries two limited decoding conventions. It reads qualifying runs of U+200B and U+200C as binary bytes, trying either character as zero, then requires valid UTF-8. It also maps qualifying Unicode tag runs to ASCII characters. It does not search every possible steganographic encoding.
A readable decode is a candidate payload. It is not an authenticated statement about where the text came from. The built-in sample deliberately encodes "ZYPHH" to demonstrate the feature; it does not imitate a verified signature from ChatGPT, Gemini or Claude.
The distribution chart and feature intervals locate repetition. A template or formatting process can create regular gaps too. No supported decode means only that these checks found none.
The text watermark guide explains why statistical watermarks require a different kind of verification. Cleaning Unicode does not certify that a generation-time signal is absent.
Check lookalikes without rewriting the alphabet
A visible letter can be unexpected too. In the constructed word "pаypal", the first a-looking letter is Cyrillic U+0430 rather than Latin U+0061. The spelling can look familiar while the stored letters differ.
Zyphh has a limited table of selected Cyrillic and Greek lookalikes. It flags a listed character when it appears inside an otherwise Latin-containing letter or number sequence. It is not a complete implementation of the Unicode security mechanisms.
Mixed scripts occur in names, quotations and multilingual writing. Check the intended spelling against its source before substituting a letter. The inspector leaves these edits to the reader.
Finding a lookalike also does not establish AI authorship. The AI detector accuracy guide explains the separate limits of classifying writing from its final form.
Inspect the source text before blaming the document
Text copied from a PDF or imported from a Word document may differ from the text the author typed. Extraction can affect reading order, spaces and line breaks. Compare the inspected passage with the original file, especially when it contains columns, tables or ligatures.
For a character-level investigation, original plain text gives you fewer extraction steps to account for. Preserve the document too if you need its editing history or metadata. The character inspector analyzes the extracted string; it does not authenticate the file.
The document checking guide describes the supported upload workflow and import limits. If two copies have different inventories, first compare how you obtained them.
A screenshot may show a rendering problem, but it does not preserve the underlying character sequence. Keep a text copy alongside any image you use to document the issue.
Keep an edit record you can understand later
Download the cleaned copy only after reviewing its wording, layout and emoji. Export the change log when you need to explain the transformation. The on-screen log shows up to 100 character edits; the JSON export contains the full log and source text.
Offsets in that log refer to the original input. Once characters are removed or combined, the same numerical position in the output may refer to a different character. Record the selected categories and whether NFC was enabled.
The export contains the source passage, so review it before sharing. Include the cleaned copy and explain why you made the edits.
If you also rewrite the wording, treat that as a separate change. The guide to paraphrasing without changing meaning covers the checks that a character inventory cannot perform, including facts, negation and citations.