Split emoji and text into visible characters (graphemes) and Unicode code points. See ZWJ parts, skin tones, variation selectors, flags and keycaps. Counts graphemes, code points, UTF-16 units and UTF-8 bytes, and copies code points as U+, JavaScript, HTML, Python or UTF-8.
Many emoji that look like one character are really several Unicode code points. 👨👩👧 is three people joined by ZERO WIDTH JOINER characters, 🇯🇵 is two regional indicator letters, and 👍🏽 is a thumbs up followed by a skin tone modifier. The Emoji Splitter shows how any text is built, which helps when a string length looks wrong, a database column overflows, or an emoji is cut in half.
Intl.Segmenter, the same rules used for cursor movement and text selectionlength) and UTF-8 bytesU+1F44D), JavaScript (ES6) \u{1F44D}, JavaScript (UTF-16) surrogate pairs, HTML entity, Python and UTF-8 bytes| Emoji | Sequence | Graphemes | Code points | UTF-16 units | UTF-8 bytes |
|---|---|---|---|---|---|
| 👍🏽 | U+1F44D U+1F3FD | 1 | 2 | 4 | 8 |
| 🇯🇵 | U+1F1EF U+1F1F5 | 1 | 2 | 4 | 8 |
| 1️⃣ | U+0031 U+FE0F U+20E3 | 1 | 3 | 3 | 7 |
| 👨👩👧 | U+1F468 U+200D U+1F469 U+200D U+1F467 | 1 | 5 | 8 | 18 |
JavaScript length counts UTF-16 code units. The family emoji has five code points (three people and two ZWJs), and each person needs two UTF-16 units, so the total is 8. Use Intl.Segmenter to count what users see as one character.
Your device's font does not support that sequence, so it falls back to drawing the parts one by one. The code points are still correct.
From the bundled emoji-datasource data (Unicode names and flag names). Code points not in that data are named by type, such as ZERO WIDTH JOINER or VARIATION SELECTOR-16; other characters may have no name.
No. Everything runs in your browser.