UTF-32 Encoding
UTF-32 (32-bit Unicode Transformation Format) is the fixed-width character encoding form that maps each Unicode scalar value to exactly one 32-bit code unit (4 bytes). In UTF-32, the numerical value of the 32-bit code unit is identically equal to the scalar value. Because it uses a 32-bit integer for every scalar value, UTF-32 requires no variable-width surrogates; however, surrogate code points (U+D800–U+DFFF) and values above U+10FFFF are strictly invalid.
Technical Quick Facts
U+0000–U+D7FF & U+E000–U+10FFFF
1,112,064 total valid scalar values
U+D800–U+DFFF (Prohibited)
Not scalar values; rejected in UTF-32
U+FEFF)
BE: 00 00 FE FF | LE: FF FE 00 00
Unicode Codespace & Valid Scalar Ranges
The Unicode codespace consists of $1,114,112$ code points ($U+0000$ to $U+10FFFF$). In UTF-32, every valid Unicode scalar value maps to exactly one 32-bit code unit (4 bytes). Surrogate code points ($U+D800$ to $U+DFFF$) and values exceeding $U+10FFFF$ are strictly prohibited by Unicode Standard 17.0 (§3.8 & §3.9).
| Codespace Section | Scalar Range | Code Units | Byte Length | UTF-32 Code Unit Range (Hex) | Conformance Status |
|---|---|---|---|---|---|
|
Basic Multilingual Plane (BMP, Lower)
Contains ASCII, Latin, Greek, Cyrillic, Hebrew, Arabic, and most world scripts.
|
U+0000–U+D7FF |
1 Code Unit |
4 Bytes
|
00000000–0000D7FF |
Valid Scalar Range |
|
High & Low Surrogate Range (Reserved)
Reserved exclusively for UTF-16 surrogate pairs. Surrogates are NOT Unicode scalar values.
|
U+D800–U+DFFF |
0 Units (Prohibited) | — | 0000D800–0000DFFF |
Prohibited (Surrogates) |
|
Basic Multilingual Plane (BMP, Upper)
Private Use Area, CJK compatibility, alphabetic presentation forms, and specials.
|
U+E000–U+FFFF |
1 Code Unit |
4 Bytes
|
0000E000–0000FFFF |
Valid Scalar Range |
|
Supplementary Planes (Planes 1–16)
Emojis, historic scripts (Linear B, Egyptian Hieroglyphs), musical symbols, rare CJK ideographs.
|
U+10000–U+10FFFF |
1 Code Unit |
4 Bytes
|
00010000–0010FFFF |
Valid Scalar Range |
|
Outside Unicode Codespace
Values exceeding U+10FFFF are outside the Unicode codespace and cannot be represented in UTF-32.
|
> U+10FFFF |
0 Units (Prohibited) | — | 00110000–FFFFFFFF |
Prohibited (Out of Codespace) |
UTF-32 32-Bit Code Unit Architecture
Unlike UTF-8 (1–4 bytes) or UTF-16 (2 or 4 bytes), UTF-32 is fixed-width at the code-unit level. The Unicode scalar value U+1F600 (😀) maps directly to 32-bit code unit 0001F600, which serializes into four 8-bit bytes.
00000000 00000001 11110110 00000000
00 01 F6 00
00 F6 01 00
Interactive UTF-32 Inspector & Validator
Inspect any text or character in real time to view its exact 32-bit code units, toggle between Big-Endian and Little-Endian byte serializations, or input raw hexadecimal code units to validate conformance.
0001F600
00 01 F6 00
00 F6 01 00
0001F600 maps directly to Unicode scalar value U+1F600 (Grinning Face) in Plane 1. Curated UTF-32 Character Examples Across Planes
Representative Unicode characters, mathematical symbols, historic scripts, combining marks, and complex emoji sequences encoded in UTF-32.
Latin Capital Letter A
Plane 0 (BMP) • Basic ASCIIU+0041
00000041
00 00 00 41
41 00 00 00
Standard ASCII character; numeric value 65 maps identically to code unit 00000041.
Latin Small Letter E with Acute
Plane 0 (BMP) • Latin-1 SupplementU+00E9
000000E9
00 00 00 E9
E9 00 00 00
Precomposed accented letter. In UTF-32, directly encoded as 000000E9.
Euro Sign
Plane 0 (BMP) • Currency SymbolsU+20AC
000020AC
00 00 20 AC
AC 20 00 00
Requires 3 bytes in UTF-8 and 2 bytes in UTF-16, but exactly 4 bytes in UTF-32.
Infinity
Plane 0 (BMP) • Mathematical OperatorsU+221E
0000221E
00 00 22 1E
1E 22 00 00
Mathematical operator represented as code unit 0000221E.
Grinning Face
Plane 1 (SMP) • Emoji & PictographsU+01F600
0001F600
00 01 F6 00
00 F6 01 00
Supplementary plane emoji. Unlike UTF-16 surrogate pairs, UTF-32 uses exactly 1 code unit (0001F600).
Linear B Ideogram B105F Mare
Plane 1 (SMP) • Historic ScriptsU+010083
00010083
00 01 00 83
83 00 01 00
Ancient Mediterranean script in Plane 1, encoded as 00010083.
CJK Unified Ideograph-2A6A5
Plane 2 (SIP) • CJK Extension BU+02A6A5
0002A6A5
00 02 A6 A5
A5 A6 02 00
Rare 64-stroke Chinese ideograph (dragon quadruple), represented as 0002A6A5.
e + Combining Acute Accent
Plane 0 (BMP) • Combining SequenceU+0065 + U+0301
00000065 00000301
00 00 00 65 00 00 03 01
65 00 00 00 01 03 00 00
Decomposed grapheme cluster: 2 scalars = 2 UTF-32 units (8 bytes), displaying as 1 visible letter.
Woman Technologist
Plane 1 & Plane 0 • Emoji ZWJ SequenceU+01F469 + U+200D + U+01F4BB
0001F469 0000200D 0001F4BB
00 01 F4 69 00 00 20 0D 00 01 F4 BB
69 F4 01 00 0D 20 00 00 BB F4 01 00
Comprises 3 scalars: Woman (U+1F469), ZWJ (U+200D), and Laptop (U+1F4BB) = 3 units (12 bytes).