Unicode blocks
Unicode divides its code space into named blocks, each a contiguous range reserved for one script or one family of symbols. Knowing which block a character lives in is what turns a mystery — why did this paste break, why does this regex miss it — into an answerable question.
The block names are identifiers, not descriptions: they are used verbatim in a regex as \p{Block=Cyrillic} and in every Unicode data file, so they stay in English in every language. What is translated is the sentence next to them.
This is a curated selection rather than all of the roughly 340 blocks — the ones that come up when working with text. The ASCII table covers the first block in full detail; this page starts where that one stops.
Latin
5Latin is spread over several blocks, which is why matching it by block never quite works.
| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+0000–U+007F | Basic Latin | ASCII, unchanged: Latin letters, digits, punctuation, control codes | A z 0 ? |
| U+0080–U+00FF | Latin-1 Supplement | Western European accents, plus £ § × ÷ and the non-breaking space | é ñ ü × |
| U+0100–U+017F | Latin Extended-A | Central and Eastern European accents — č, ł, ő, ș | č ł ő š |
| U+0180–U+024F | Latin Extended-B | Additional letters for African and phonetic orthographies | ƒ ǎ ȳ |
| U+1E00–U+1EFF | Latin Extended Additional | Vietnamese and other multi-accented Latin letters | ạ ế ṛ |
Combining marks and modifiers
4These attach to the character before them and have no width of their own.
| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+0300–U+036F | Combining Diacritical Marks | Accents applied to the preceding characterThe second, decomposed way to write é. Normalise to NFC to collapse it. | ◌́ ◌̈ ◌̌ |
| U+02B0–U+02FF | Spacing Modifier Letters | Superscript letters and IPA stress and tone marks | ʰ ˈ ˚ |
| U+FE00–U+FE0F | Variation Selectors | Invisible selectors choosing text or emoji presentationVS16 is what turns ✔ into an emoji. Invisible, but it counts. | |
| U+FE20–U+FE2F | Combining Half Marks | Halves of marks spanning two characters |
Other scripts
8| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+0370–U+03FF | Greek and Coptic | Greek, and the Coptic letters that share its block | α Ω π |
| U+0400–U+04FF | Cyrillic | Russian, Ukrainian, Serbian, Bulgarian and related alphabets | д Я ж |
| U+0590–U+05FF | Hebrew | Hebrew letters, vowel points and cantillation marks | א ש |
| U+0600–U+06FF | Arabic | Arabic letters, in their isolated forms | ا ب ي |
| U+0900–U+097F | Devanagari | Hindi, Sanskrit, Marathi and related languages | क ह |
| U+0E00–U+0E7F | Thai | Thai consonants, vowels and tone marks | ก ท |
| U+10A0–U+10FF | Georgian | Georgian Mkhedruli | ა ბ |
| U+1F00–U+1FFF | Greek Extended | Polytonic Greek — the accents of classical texts | ᾰ ῶ |
Punctuation and numbers
5The source of most invisible characters that survive a copy-paste.
| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+2000–U+206F | General Punctuation | Dashes, curly quotes, ellipsis — and most invisible spacesU+200B zero-width space and U+2060 word joiner live here. | – — “ ” … |
| U+2070–U+209F | Superscripts and Subscripts | Superscript and subscript digits and letters | ⁰ ⁴ ₂ |
| U+20A0–U+20CF | Currency Symbols | Currency signs Latin-1 has no room for — €, ₴, ₹, ₺ | € ₴ ₺ |
| U+2100–U+214F | Letterlike Symbols | Symbols built from letters — ™, №, ℃, ℓ | ™ ℃ № |
| U+2150–U+218F | Number Forms | Vulgar fractions and Roman numerals as single characters | ⅓ ⅕ Ⅷ |
Symbols and drawing
9| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+2190–U+21FF | Arrows | Arrows in every direction, single and double | ← ↔ ⇒ |
| U+2200–U+22FF | Mathematical Operators | Mathematical operators — ∀, ∑, ≠, ∞ | ∀ ∑ ≠ |
| U+2300–U+23FF | Miscellaneous Technical | Keyboard and technical symbols — ⌘, ⌥, ⌫, ⏎ | ⌘ ⌥ ⏎ |
| U+2500–U+257F | Box Drawing | Lines and corners for terminal boxes and tables | ─ ┌ ╬ |
| U+2580–U+259F | Block Elements | Solid and shaded blocks, used for terminal bar charts | █ ▓ ░ |
| U+25A0–U+25FF | Geometric Shapes | Squares, circles, triangles and diamonds | ■ ● ▲ |
| U+2600–U+26FF | Miscellaneous Symbols | Weather, astrology, chess, cards and the card suits | ☀ ★ ♥ |
| U+2700–U+27BF | Dingbats | Printer's ornaments — check marks, scissors, stars | ✂ ✓ ➜ |
| U+2800–U+28FF | Braille Patterns | All 256 Braille dot patterns, used for terminal graphics too | ⠁ ⣿ |
CJK and fullwidth
6Fullwidth forms are separate characters from their ASCII twins, not a font.
| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+3000–U+303F | CJK Symbols and Punctuation | Ideographic space and CJK punctuationU+3000 is a full-width space and is not trimmed by most whitespace rules. | 、 。 |
| U+3040–U+309F | Hiragana | The Japanese syllabary for native words and grammar | あ ん |
| U+30A0–U+30FF | Katakana | The Japanese syllabary for loanwords and emphasis | ア ン |
| U+4E00–U+9FFF | CJK Unified Ideographs | The main Han block — nearly 21,000 characters shared by Chinese, Japanese and Korean | 日 本 語 |
| U+AC00–U+D7AF | Hangul Syllables | All 11,172 precomposed Korean syllablesEach can also be written as separate Jamo, which normalisation reconciles. | 한 글 |
| U+FF00–U+FFEF | Halfwidth and Fullwidth Forms | Full-width copies of ASCII, and half-width KatakanaA is not A. Text pasted from CJK forms often needs folding to ASCII. | A ! ア |
Above U+FFFF, and the special cases
10Anything here takes two UTF-16 units, which is where string-length bugs come from.
| Range | Block | What is in it | Sample |
|---|---|---|---|
| U+D800–U+DB7F | High Surrogates | First half of a UTF-16 pair — never valid on its ownA lone surrogate is what a string cut through an emoji leaves behind. | |
| U+DC00–U+DFFF | Low Surrogates | Second half of a UTF-16 pair — never valid on its own | |
| U+E000–U+F8FF | Private Use Area | Unassigned by design, for private agreementsWhere icon fonts put their glyphs. Meaningless without the matching font. | |
| U+FFF0–U+FFFF | Specials | The replacement character and the byte order mark's relativesU+FFFD is what you see when a decoder gave up on a byte. | � |
| U+1D400–U+1D7FF | Mathematical Alphanumeric Symbols | Bold, italic, script and double-struck alphabets as charactersThe source of 𝐟𝐚𝐤𝐞 𝐛𝐨𝐥𝐝 in social media names. Unsearchable. | 𝐀 𝓪 𝟙 |
| U+1F300–U+1F5FF | Miscellaneous Symbols and Pictographs | Most non-face emoji — weather, objects, places, symbols | 🌍 🔥 📄 |
| U+1F600–U+1F64F | Emoticons | The face emoji | 😀 🙂 |
| U+1F680–U+1F6FF | Transport and Map Symbols | Vehicles, traffic and map symbols | 🚀 🛑 |
| U+1F900–U+1F9FF | Supplemental Symbols and Pictographs | Later emoji additions — later faces, gestures, objects | 🤖 🧠 |
| U+E0000–U+E007F | Tags | Invisible tag characters, used in flag sequencesAlso the vector for invisible text smuggled inside a normal-looking string. |