Notes

Why Base64 works in groups of four characters

Three bytes are 24 bits, which split cleanly into four 6-bit indexes. This note covers grouping, padding, and the URL-safe alphabet, then compares what JWT and data URLs strip off the outside.

By Brook/Updated 2026-10-09/10 min read

Why bytes have to become text first

Many channels only promise to carry a safe subset of characters: JSON strings, URLs, HTTP headers, HTML attributes. Raw bytes can contain NULs, newlines, and control characters that truncate a field or get read as syntax. Base64 maps arbitrary bytes onto 64 unproblematic characters. The cost is size: the text is about 4/3 of the original.

It is an encoding, not compression, and not encryption. Anyone can turn the text back into the original bytes. If the goal is secrecy, Base64 does not provide it. If the goal is a smaller payload, it makes the payload larger. It only answers "can these bytes travel through a text channel?".

24 bits become four indexes

  1. 01Bytes3 bytes = 24 bitsIf the tail is short, remember how many bytes are missing.
  2. 02SplitFour 6-bit groupsEach value lands between 0 and 63.
  3. 03AlphabetA-Z a-z 0-9 + /64 characters, one per index.
  4. 04Pad= fills the groupOne missing byte adds ==. Two missing bytes add a single =.
Figure 1. A group always consumes 3 bytes and emits 4 characters. A short tail is padded so the group is still full.

Six bits represent 0 through 63, so the alphabet has 64 symbols, not 256. Three bytes divide evenly. When the input length is not a multiple of three, the last group is short. One missing byte leaves two character slots, filled with =. Two missing bytes leave one slot. Padding tells the decoder how many trailing zero bits were not real data. Without that rule, those zeros become extra empty bytes.

The last two characters of the standard alphabet are + and /. A decoder should fail on anything outside the alphabet, not skip it. Whitespace policy varies. The safe approach is to strip known spaces and newlines first, and to say that you did. Quietly dropping odd characters turns corrupt input into different bytes.

The URL-safe variant changes two characters

+ in a query string can be read as a space, and / in a path is a separator. The URL-safe variant uses - instead of + and _ instead of /. The bit grouping is identical. Only the last two alphabet entries change names. = is also awkward in URLs, so some formats drop the trailing padding. Before decoding, restore it from the length: remainder 2 means add ==, remainder 3 means add =, remainder 1 is illegal.

JWT uses that variant: Base64URL, usually without padding. Pasting one segment into a decoder that only accepts +/ and required padding will fail. Swap the characters, restore =, and the bytes match standard Base64. The reverse is the same. Renaming the algorithm without changing the alphabet is not enough.

Data URLs and JWTs wrap another layer around Base64

Standard Base64

Alphabet includes + and /, and usually =.

  • Fits inside a JSON string
  • Awkward in a URL
  • Decoding returns raw bytes

Base64URL

Alphabet includes - and _, padding often omitted.

  • Each segment of a JWT
  • Restore padding first
  • Do not treat the signature as JSON

Data URL

data:type;base64, and then the payload.

  • Everything before the comma is a header
  • Decode only what follows
  • The type says how to read the bytes
Figure 2. Only the middle span is Base64. Strip the wrapper first, or those characters get decoded into the bytes.

A small image embedded in a page is often data:image/png;base64,....... The media type and the encoding sit before the comma. Only the rest is Base64. Decoding the whole string, including data:, adds header bytes and the file signature will not match. Some text data URLs skip Base64 and use percent-encoding. Check the header for the base64 token before you choose a decoder.

When decoding fails, check in this order

The error usually only says the input is not valid Base64. It does not say whether the wrapper, the alphabet, or the padding is wrong. Walking the checks below is faster than swapping decoders until one of them happens to accept the string.

  • Is a data URL, a pair of quotes, or a line break still attached? Peel back to alphabet characters.
  • Which alphabet is it? + / and - _ should not both appear in one well-formed value.
  • Is padding the only thing missing? After removing =, the length mod 4 must not be 1.
  • The result is bytes. If you expected text, decode those bytes as UTF-8. A failure there is an encoding problem, not a Base64 problem.

The Base64 tool on this site encodes and decodes standard Base64 over bytes. The JWT tool handles Base64URL and missing padding on its own. Do not paste the same string into both and expect one path: decide whether the text is going into a URL, into JSON, or back to file bytes.

Tools mentioned here

Keep reading