Notes

What UTF-8 actually encodes

Character counts, code points, and bytes are three different rulers. This note walks one code point through UTF-8 and explains why URL encoding and JSON escapes look nothing like each other.

By Brook/Updated 2026-10-09/11 min read

Code points, code units, and bytes are different rulers

Unicode assigns each character a code point, written U+XXXX. A code point is a number, not a storage format. UTF-8, UTF-16, and UTF-32 are three rules for turning that number into bytes. The same text can share a code-point sequence and still have completely different bytes. A tool that only says "length" is incomplete until it says whether it counted code points, UTF-16 code units, or bytes.

JavaScript strings count UTF-16 code units. A typical non-ASCII letter in the BMP is one code unit and one code point. Characters on the supplementary planes, including many emoji, take two code units, a surrogate pair. String.length counts code units. Size in bytes needs another pass through UTF-8. "Limit this field to 32 characters" and "limit this field to 32 bytes" are different requirements, and they first disagree on non-ASCII text and emoji.

UTF-8 picks the byte count from the code point

The rule fits in a table. U+0000 to U+007F is one byte with the high bit clear, which is why ASCII survives unchanged. U+0080 to U+07FF is two bytes. U+0800 to U+FFFF is three bytes. U+10000 to U+10FFFF is four bytes. The lead byte uses high bits to say how many continuation bytes follow. Every continuation byte starts with 10. A reader can tell, mid-stream, whether it is standing on a character boundary.

  1. 01NumberCode pointFor example U+4E2D. Still just an integer.
  2. 02Split16 bits in three groups4 bits, then 6, then 6, for the three-byte template.
  3. 03Template1110xxxx 10xxxxxx 10xxxxxxThe lead byte says two continuation bytes follow.
  4. 04BytesThree bytesE4 B8 AD. This is what files and networks store.
Figure 1. A code point in U+0800-U+FFFF, where most CJK characters live, becomes three bytes.

U+4E2D in binary is 0100 1110 0010 1101. Split 4/6/6 into 0100, 111000, and 101101. Fill the template and you get 11100100 10111000 10101101, which is E4 B8 AD. The URL form %E4%B8%AD is those three bytes written as a percent sign plus two hex digits. It is not a second Unicode for the same character.

Why one transform can look like garbage

Mojibake is usually bytes read with the wrong encoding, not a broken font. The UTF-8 sequence E4 B8 AD displayed as Latin-1 or as a legacy double-byte encoding becomes unrelated characters. The reverse, legacy bytes read as UTF-8, fails where a continuation byte is illegal. A tool that shows bytes has to name the encoding it is using.

JSON \u4e2d is a different layer. It does not mean the UTF-8 bytes. It means the code point U+4E2D. The parser inserts that code point, and the language runtime decides whether the string is stored as UTF-16 or something else. One piece of data can show up as a source character, a JSON escape, and a percent-encoded URL. They convert into each other, but you cannot skip a step.

Which length an editor is counting

Code points

How many Unicode numbers.

  • U+4E2D is 1 code point
  • Good for "how many characters"
  • Not the same as file size

UTF-16 units

What JavaScript length counts.

  • U+4E2D is still 1
  • Many emoji are 2
  • A slice can cut a surrogate pair

UTF-8 bytes

File size and URL encoding.

  • U+4E2D is 3 bytes
  • ASCII is 1 byte
  • Matches Content-Length
Figure 2. U+4E2D, three numbers, three questions. All three answers are right.

Slicing by code unit is the common bug. text.slice(0, n) on a title that may contain emoji can land inside a surrogate pair and leave a lone surrogate. Later encoding or storage then fails, or inserts a replacement character. Cutting by code point means walking code points first. Cutting by bytes also has to avoid splitting a UTF-8 character in half.

URL encoding encodes bytes, not "characters"

encodeURIComponent turns the string into UTF-8 bytes, then writes every byte outside the unreserved set as %HH. A space becomes %20, not +. The + for space belongs to application/x-www-form-urlencoded. Mixing the two functions produces query strings that decode with an extra plus, or plus signs that turn into spaces.

Encoding text that is already encoded turns % into %25. You need two decodes to get back. A double-encoded CJK sequence shows up as %25E4%25B8%25AD. Look for well-formed %HH pairs and decide how many layers you have. Repeated decoding because the text "looks wrong" will eventually eat real percent signs.

A few habits that avoid the usual bugs

  • Exchange text as UTF-8, and name it in the HTTP header or file format. Do not rely on the other side guessing.
  • Say the unit when you limit length. Code points or graphemes for people, bytes for protocols and storage.
  • Normalize before you compare or search. The same visible letter can be different code points plus combining marks.
  • Log both the original text and the encoded form. Percent-encoding alone makes the next person decode it in their head.

The URL tool on this site percent-encodes UTF-8 bytes. The Base64 tool handles bytes, not characters. Decide whether you are holding code points, raw bytes, or text that is already percent-encoded, then pick the tool.

Tools mentioned here

Keep reading