About the Unicode Character Inspector
What looks like one character on screen can be several code points. The table lists the text by grapheme cluster (one user-perceived character, found with Intl.Segmenter) and, inside each cluster, by code point: the family emoji is three people joined by two zero-width joiners, é can be the single code point U+00E9 or e followed by the combining accent U+0301, and a flag is two regional indicator letters. For each code point you get the name, the general category (Lu, Mn, Cf…) and the block, the UTF-8 bytes, the UTF-16 code units, and the way to write it in HTML (&#x…;), JavaScript (\u{…}) and CSS.
Characters you cannot see are shown as labelled chips and counted in the summary: the zero-width space, joiner and non-joiner, the byte order mark, the soft hyphen, the no-break space and the other variants of the space, the line and paragraph separators, control characters, variation selectors and tag characters. Bidirectional controls such as U+202E (right-to-left override) are flagged because they can make source code or a file name read differently from what it is, the “Trojan Source” problem. Cyrillic and Greek letters that look like Latin ones are flagged when they appear inside a word that is otherwise Latin, as in a fake pаypal.com. Joiners and selectors that do their normal job, inside an emoji or in scripts such as Arabic and Devanagari, are marked as notes and not as warnings.
The cleaned text removes the invisible characters (keeping the ones emoji need), replaces the space variants with a normal space, optionally turns typographic quotes and dashes into ASCII, and applies a Unicode normalisation form: NFC composes letters and accents, NFD separates them, and NFKC and NFKD also replace compatibility characters such as fi, fullwidth letters and styled mathematical letters with plain ones. Names come from the Unicode Character Database for Latin, Greek, Cyrillic, punctuation, symbols and emoji, and from the naming rules for ideographs and Hangul; other characters show their block. The category comes from the Unicode data of your browser. Everything runs locally.
How to use it
- Paste the text, a file name, a URL or a line of code.
- Read the summary: the chips list the invisible and look-alike characters found.
- Look at the revealed view and the table for the position and the code point of each one.
- Choose what to clean and the normalisation form.
- Copy the cleaned text.