Unicode Escape & Unescape

Escape non-ASCII characters for source code, JSON or HTML, or turn \u escapes and U+ code points back into readable text.

Loading tool…

About the Unicode Escape

Different languages write Unicode escapes differently. JSON, JavaScript and Java use \uXXXX, which holds only 16 bits, so characters above U+FFFF must be written as a UTF-16 surrogate pair: 🚀 (U+1F680) becomes \uD83D\uDE80. ES2015+, Rust and Swift accept \u{1F680}; Python and C use \U0001F680; HTML and XML use 🚀; documentation writes code points as U+1F680.

The encoder iterates over code points, not UTF-16 units, so astral characters are never split incorrectly. By default only non-ASCII characters are escaped, leaving readable ASCII; tick “Escape ASCII too” to escape everything. U+ format always lists every code point, separated by spaces.

The decoder recognises all of these formats in the same input, recombines surrogate pairs, and rejects out-of-range values such as \u{110000}. It is useful for reading escaped strings in logs, JSON produced with ensure_ascii=True, and Java .properties files.

How to use it

  1. Choose Escape or Unescape.
  2. Pick the escape format your language uses.
  3. Paste the text or escaped string.
  4. Copy the output.

Frequently asked questions

Why does an emoji become two \u escapes?
\uXXXX holds one UTF-16 code unit. Characters above U+FFFF need two units (a high and a low surrogate), for example \uD83D\uDE00 for 😀.
What is the difference between a code point and UTF-8 bytes?
A code point is the character's number in Unicode (U+00E9 for é). UTF-8 is one way to store it as bytes (C3 A9). This tool works with code points; use Text to Hex for bytes.
Can I use \u{…} in JSON?
No. JSON only supports \uXXXX, so astral characters must use surrogate pairs. The first format option produces valid JSON escapes.

Related tools