About Unicode Lookup & Character Inspector
The Unicode Standard 15.1 character inspector and symbol lookup table. Paste or type any glyph, emoji, or text to decompose it into its exact hexadecimal code point, UTF-8 byte sequence, UTF-16 surrogate pairs, HTML entities, and language escape sequences (\u{...}). Includes a curated searchable table of popular Unicode symbols.
Key Capabilities & Features
- Real-time Unicode 15.1 character decomposition for single characters, emojis, and strings
- Decodes formal code point (U+XXXX), UTF-8 bytes, and UTF-16 surrogate code units
- Generates HTML decimal, HTML hex, and JavaScript/TypeScript string escape syntax
- Automated Unicode block category classification (Latin, CJK, Dingbats, Math, Arrows, etc.)
- Pre-indexed quick symbol palette for popular stars, hearts, checks, and arrows
How to Use Unicode Lookup & Character Inspector
Enter Character or String
Paste any symbol, emoji, or string into the inspector input field.
Inspect Architecture
Review the hex code point, byte sequences, and character block.
Search Symbols
Use the symbol search bar to find mathematical marks, stars, or currency symbols.
Copy Code Units
Copy JavaScript escape strings (\u{...}) or HTML entities with one click.
Privacy & In-Browser Execution Guarantee
100% Client-Side. Unicode character inspection executes via browser TextEncoder and String.codePointAt.
Frequently Asked Questions
What is a UTF-16 surrogate pair?
Characters outside the Basic Multilingual Plane (code points above U+FFFF, such as modern emojis) cannot fit into a single 16-bit code unit. UTF-16 represents them using two 16-bit code units called a high and low surrogate.
How do code points relate to UTF-8 bytes?
A Unicode code point is an abstract integer (from 0 to 0x10FFFF). UTF-8 is a variable-length binary encoding that serializes that integer into 1, 2, 3, or 4 bytes.