Emoji Splitter

Split emoji and text into visible characters (graphemes) and Unicode code points. See ZWJ parts, skin tones, variation selectors, flags and keycaps. Counts graphemes, code points, UTF-16 units and UTF-8 bytes, and copies code points as U+, JavaScript, HTML, Python or UTF-8.

About this tool

What does the Emoji Splitter do?

Many emoji that look like one character are really several Unicode code points. 👨‍👩‍👧 is three people joined by ZERO WIDTH JOINER characters, 🇯🇵 is two regional indicator letters, and 👍🏽 is a thumbs up followed by a skin tone modifier. The Emoji Splitter shows how any text is built, which helps when a string length looks wrong, a database column overflows, or an emoji is cut in half.

How to use

  1. Type or paste text into Input. The counts under the box update immediately.
  2. Click a character tile to see its breakdown: name, sequence type, ZWJ parts, Code points and Notations.
  3. Click the large character, a ZWJ part or a code point row to copy it. Use the copy buttons next to each notation to copy escapes.
  4. To copy the whole input as escapes, pick a Notation and click the copy button beside it.

Features

  • Grapheme splitting with the browser's Intl.Segmenter, the same rules used for cursor movement and text selection
  • Four counts: characters (graphemes), code points, UTF-16 units (JavaScript length) and UTF-8 bytes
  • Code point roles: Base, ZWJ, Skin tone, Hair, Variation selector, Regional indicator, Tag, Cancel tag, Keycap, Combining mark and Control
  • Sequence type such as ZWJ sequence, flag, tag sequence (e.g. the Scotland flag), keycap and skin tone sequence
  • ZWJ parts: e.g. 🏳️‍🌈 → 🏳️ + 🌈
  • Notations: Code points (U+1F44D), JavaScript (ES6) \u{1F44D}, JavaScript (UTF-16) surrogate pairs, HTML entity, Python and UTF-8 bytes
  • Invisible characters (ZWJ, variation selectors, tags) are shown as labeled placeholders

Examples

EmojiSequenceGraphemesCode pointsUTF-16 unitsUTF-8 bytes
👍🏽U+1F44D U+1F3FD1248
🇯🇵U+1F1EF U+1F1F51248
1️⃣U+0031 U+FE0F U+20E31337
👨‍👩‍👧U+1F468 U+200D U+1F469 U+200D U+1F46715818

FAQ

Why is "👨‍👩‍👧".length 8 in JavaScript?

JavaScript length counts UTF-16 code units. The family emoji has five code points (three people and two ZWJs), and each person needs two UTF-16 units, so the total is 8. Use Intl.Segmenter to count what users see as one character.

Why does an emoji show as several separate emoji?

Your device's font does not support that sequence, so it falls back to drawing the parts one by one. The code points are still correct.

Where do the names come from?

From the bundled emoji-datasource data (Unicode names and flag names). Code points not in that data are named by type, such as ZERO WIDTH JOINER or VARIATION SELECTOR-16; other characters may have no name.

Is my text sent anywhere?

No. Everything runs in your browser.