Unicode Character Inspector – Find Invisible Characters

Paste a text to see what it is really made of: every code point with its name and encodings, and the invisible or look-alike characters that hide in it.

Loading tool…

About the Unicode Character Inspector

What looks like one character on screen can be several code points. The table lists the text by grapheme cluster (one user-perceived character, found with Intl.Segmenter) and, inside each cluster, by code point: the family emoji is three people joined by two zero-width joiners, é can be the single code point U+00E9 or e followed by the combining accent U+0301, and a flag is two regional indicator letters. For each code point you get the name, the general category (Lu, Mn, Cf…) and the block, the UTF-8 bytes, the UTF-16 code units, and the way to write it in HTML (&#x…;), JavaScript (\u{…}) and CSS.

Characters you cannot see are shown as labelled chips and counted in the summary: the zero-width space, joiner and non-joiner, the byte order mark, the soft hyphen, the no-break space and the other variants of the space, the line and paragraph separators, control characters, variation selectors and tag characters. Bidirectional controls such as U+202E (right-to-left override) are flagged because they can make source code or a file name read differently from what it is, the “Trojan Source” problem. Cyrillic and Greek letters that look like Latin ones are flagged when they appear inside a word that is otherwise Latin, as in a fake pаypal.com. Joiners and selectors that do their normal job, inside an emoji or in scripts such as Arabic and Devanagari, are marked as notes and not as warnings.

The cleaned text removes the invisible characters (keeping the ones emoji need), replaces the space variants with a normal space, optionally turns typographic quotes and dashes into ASCII, and applies a Unicode normalisation form: NFC composes letters and accents, NFD separates them, and NFKC and NFKD also replace compatibility characters such as fi, fullwidth letters and styled mathematical letters with plain ones. Names come from the Unicode Character Database for Latin, Greek, Cyrillic, punctuation, symbols and emoji, and from the naming rules for ideographs and Hangul; other characters show their block. The category comes from the Unicode data of your browser. Everything runs locally.

How to use it

  1. Paste the text, a file name, a URL or a line of code.
  2. Read the summary: the chips list the invisible and look-alike characters found.
  3. Look at the revealed view and the table for the position and the code point of each one.
  4. Choose what to clean and the normalisation form.
  5. Copy the cleaned text.

Frequently asked questions

What is a zero-width space and where does it come from?
U+200B is a character with no width that allows a line break. It arrives with text copied from web pages, chat applications, PDF files and word processors, and it makes two strings that look the same compare as different, breaks URLs, passwords and JSON keys, and causes “unexpected character” errors in code.
Why do the character count and the length in my code differ?
They count different things. String.length in JavaScript counts UTF-16 code units, so an emoji is 2 or more; a database column in bytes counts UTF-8 bytes; a reader counts grapheme clusters. The summary shows all four numbers.
What is the difference between NFC and NFKC?
NFC only chooses between equivalent ways of writing the same character, such as é as one code point. NFKC also replaces compatibility characters with their plain form: fi becomes fi, ① becomes 1 and fullwidth A becomes A. NFC is safe for any text; NFKC loses distinctions and is meant for identifiers and search.
Does the tool know the name of every character?
No. The complete list of names is too large for a web page. It has the names of about 6,000 characters (Latin, Greek, Cyrillic, punctuation, symbols, emoji, all invisible and control characters) and the rule-based names of CJK ideographs and Hangul syllables. For the rest the block, category and code point are shown.

Related tools