Notes
What UTF-8 actually encodes
Character counts, code points, and bytes are three different rulers. This note walks one code point through UTF-8 and explains why URL encoding and JSON escapes look nothing like each other.
By Brook/Updated 2026-10-09/11 min read
Code points, code units, and bytes are different rulers
Unicode assigns each character a code point, written U+XXXX. A code point is a number, not a storage format. UTF-8, UTF-16, and UTF-32 are three rules for turning that number into bytes. The same text can share a code-point sequence and still have completely different bytes. A tool that only says "length" is incomplete until it says whether it counted code points, UTF-16 code units, or bytes.
JavaScript strings count UTF-16 code units. A typical non-ASCII letter in the BMP is one code unit and one code point. Characters on the supplementary planes, including many emoji, take two code units, a surrogate pair. String.length counts code units. Size in bytes needs another pass through UTF-8. "Limit this field to 32 characters" and "limit this field to 32 bytes" are different requirements, and they first disagree on non-ASCII text and emoji.
UTF-8 picks the byte count from the code point
The rule fits in a table. U+0000 to U+007F is one byte with the high bit clear, which is why ASCII survives unchanged. U+0080 to U+07FF is two bytes. U+0800 to U+FFFF is three bytes. U+10000 to U+10FFFF is four bytes. The lead byte uses high bits to say how many continuation bytes follow. Every continuation byte starts with 10. A reader can tell, mid-stream, whether it is standing on a character boundary.
- 01NumberCode pointFor example U+4E2D. Still just an integer.
- 02Split16 bits in three groups4 bits, then 6, then 6, for the three-byte template.
- 03Template1110xxxx 10xxxxxx 10xxxxxxThe lead byte says two continuation bytes follow.
- 04BytesThree bytesE4 B8 AD. This is what files and networks store.
U+4E2D in binary is 0100 1110 0010 1101. Split 4/6/6 into 0100, 111000, and 101101. Fill the template and you get 11100100 10111000 10101101, which is E4 B8 AD. The URL form %E4%B8%AD is those three bytes written as a percent sign plus two hex digits. It is not a second Unicode for the same character.
Why one transform can look like garbage
Mojibake is usually bytes read with the wrong encoding, not a broken font. The UTF-8 sequence E4 B8 AD displayed as Latin-1 or as a legacy double-byte encoding becomes unrelated characters. The reverse, legacy bytes read as UTF-8, fails where a continuation byte is illegal. A tool that shows bytes has to name the encoding it is using.
JSON \u4e2d is a different layer. It does not mean the UTF-8 bytes. It means the code point U+4E2D. The parser inserts that code point, and the language runtime decides whether the string is stored as UTF-16 or something else. One piece of data can show up as a source character, a JSON escape, and a percent-encoded URL. They convert into each other, but you cannot skip a step.
Which length an editor is counting
Code points
How many Unicode numbers.
- U+4E2D is 1 code point
- Good for "how many characters"
- Not the same as file size
UTF-16 units
What JavaScript length counts.
- U+4E2D is still 1
- Many emoji are 2
- A slice can cut a surrogate pair
UTF-8 bytes
File size and URL encoding.
- U+4E2D is 3 bytes
- ASCII is 1 byte
- Matches Content-Length
Slicing by code unit is the common bug. text.slice(0, n) on a title that may contain emoji can land inside a surrogate pair and leave a lone surrogate. Later encoding or storage then fails, or inserts a replacement character. Cutting by code point means walking code points first. Cutting by bytes also has to avoid splitting a UTF-8 character in half.
URL encoding encodes bytes, not "characters"
encodeURIComponent turns the string into UTF-8 bytes, then writes every byte outside the unreserved set as %HH. A space becomes %20, not +. The + for space belongs to application/x-www-form-urlencoded. Mixing the two functions produces query strings that decode with an extra plus, or plus signs that turn into spaces.
Encoding text that is already encoded turns % into %25. You need two decodes to get back. A double-encoded CJK sequence shows up as %25E4%25B8%25AD. Look for well-formed %HH pairs and decide how many layers you have. Repeated decoding because the text "looks wrong" will eventually eat real percent signs.
A few habits that avoid the usual bugs
- Exchange text as UTF-8, and name it in the HTTP header or file format. Do not rely on the other side guessing.
- Say the unit when you limit length. Code points or graphemes for people, bytes for protocols and storage.
- Normalize before you compare or search. The same visible letter can be different code points plus combining marks.
- Log both the original text and the encoded form. Percent-encoding alone makes the next person decode it in their head.
The URL tool on this site percent-encodes UTF-8 bytes. The Base64 tool handles bytes, not characters. Decide whether you are holding code points, raw bytes, or text that is already percent-encoded, then pick the tool.
Tools mentioned here
Keep reading
- How JSON parsing works, from characters to a valueFormat, minify, and tree view look like three buttons. Underneath they share one pipeline: split the text into tokens, then fold those tokens into a value. This note walks that path and shows where line numbers come from.
- The three parts of a JWT, and what verification checksDecoding the header and payload does not mean the token can be trusted. This note separates Base64URL, the signature, and claim checks, and explains why the algorithm and the key have to be chosen together.
- How an SSE stream reassembles a model reply from deltasModel APIs often split one reply across many SSE events. This note covers the frame format, how deltas merge by path, and why tool-call arguments are concatenated strings rather than JSON on every frame.