UtilityPilot.Find a tool

Developer Tools

Unicode Converter

Turn text into four-digit Unicode escapes or U+ code points, then decode either notation back to text. Inspect emoji and international text without evaluating code or changing Unicode normalization.

Processed in your browserPrivacy details ↗

256 KiB input · UTF-16 escapes or Unicode scalar code points

Adjust your options

No sign-up. No upload.

How to convert Unicode notation

  1. Choose Text to notation or Notation to text.
  2. Select UTF-16 escapes or U+ code points. Select the same notation when decoding a previous result.
  3. Paste your source and select Convert Unicode.
  4. Check the output and counts, then copy or download the text. Reset clears both panels.

Encoding transforms every character, including ordinary ASCII letters, spaces and line breaks. This makes invisible separators visible in a representation that can be inspected. Empty input produces empty output and zero counts; whitespace is meaningful input.

Code points, UTF-16 units and UTF-8 bytes

A Unicode code point identifies a value in the Unicode character space. U+0041 is the letter A; U+1F44B is the waving-hand emoji. Code-point notation is useful in documentation and debugging but is not itself a complete JSON string.

JavaScript strings use UTF-16 code units. Values outside the basic multilingual plane need two units, called a surrogate pair. Consequently, one emoji can have one code point, two UTF-16 units and four UTF-8 bytes. A visible symbol made from combining marks or joined emoji can contain several code points. These counts do not measure displayed width or user-perceived characters.

The results report counts for the decoded text, even when the output is a much longer sequence of escapes. For user-perceived character counts and word counts, use Word Counter’s text metrics. The definitions answer different questions.

Unicode escape examples

Input

Aé👋

Output

UTF-16: \u0041\u00e9\ud83d\udc4b
Code points: U+0041 U+00E9 U+1F44B

This input has three code points, four UTF-16 units and seven UTF-8 bytes. In code-point mode a space in the original text appears as U+0020. Whitespace between notation tokens is only a separator; to decode a space, include its U+0020 token.

Precomposed é and e followed by a combining acute accent look similar but produce different sequences. The converter does not normalize either spelling. Use that distinction when investigating why two apparently identical strings compare differently.

Decoding rules and common mistakes

Escape mode recognizes exactly four hexadecimal digits after each \u. Ordinary text can appear between escapes, but every literal backslash must itself be represented by \u005c. Short escapes such as \n, braced escapes and whole quoted JSON documents are outside this mode. Decoding happens once, so an escaped backslash followed by u0041 is not recursively turned into A.

Code-point mode accepts whitespace-separated U+ tokens with four to six hexadecimal digits. Values above U+10FFFF and surrogate code points U+D800 through U+DFFF are rejected. Escape mode accepts complete high/low surrogate pairs and rejects isolated surrogates rather than silently replacing them.

Input is limited to 256 KiB. No eval, script execution or character-name dataset is involved. If you need to inspect an entire JSON document containing escaped strings, JSON Viewer is the appropriate parser.

When Unicode conversion helps

Use escapes to reproduce an invisible character in a bug report, compare localization strings or see why an emoji affects a length calculation. Code points provide a stable way to identify characters when a font does not display them clearly. A sequence that decodes successfully is not evidence that the text is correct for a particular language or application.

Unicode escapes differ from Base64, which represents bytes, and URL percent encoding, which escapes URL component bytes. Use HTML entities for character references inside HTML. Use Regex Tester to examine JavaScript pattern behavior against the decoded string.

Processing and output safety

Conversion runs locally in a browser worker. The source and result are shown as plain text, not executed as JavaScript or interpreted as HTML. Input is not sent to a processing API, saved in a URL or automatically persisted. Copy and download create the copies you request; Reset clears the interface.

Frequently asked questions

Why does one emoji produce two escapes?

Four-digit escapes represent UTF-16 units. An emoji above U+FFFF needs a surrogate pair; U+ code-point mode represents it with one token.

Does this repair corrupted text?

No. It changes notation. It cannot infer the original encoding of already corrupted text or choose a normalization form for you.