Notes
Why Base64 works in groups of four characters
Three bytes are 24 bits, which split cleanly into four 6-bit indexes. This note covers grouping, padding, and the URL-safe alphabet, then compares what JWT and data URLs strip off the outside.
By Brook/Updated 2026-10-09/10 min read
Why bytes have to become text first
Many channels only promise to carry a safe subset of characters: JSON strings, URLs, HTTP headers, HTML attributes. Raw bytes can contain NULs, newlines, and control characters that truncate a field or get read as syntax. Base64 maps arbitrary bytes onto 64 unproblematic characters. The cost is size: the text is about 4/3 of the original.
It is an encoding, not compression, and not encryption. Anyone can turn the text back into the original bytes. If the goal is secrecy, Base64 does not provide it. If the goal is a smaller payload, it makes the payload larger. It only answers "can these bytes travel through a text channel?".
24 bits become four indexes
- 01Bytes3 bytes = 24 bitsIf the tail is short, remember how many bytes are missing.
- 02SplitFour 6-bit groupsEach value lands between 0 and 63.
- 03AlphabetA-Z a-z 0-9 + /64 characters, one per index.
- 04Pad= fills the groupOne missing byte adds ==. Two missing bytes add a single =.
Six bits represent 0 through 63, so the alphabet has 64 symbols, not 256. Three bytes divide evenly. When the input length is not a multiple of three, the last group is short. One missing byte leaves two character slots, filled with =. Two missing bytes leave one slot. Padding tells the decoder how many trailing zero bits were not real data. Without that rule, those zeros become extra empty bytes.
The last two characters of the standard alphabet are + and /. A decoder should fail on anything outside the alphabet, not skip it. Whitespace policy varies. The safe approach is to strip known spaces and newlines first, and to say that you did. Quietly dropping odd characters turns corrupt input into different bytes.
The URL-safe variant changes two characters
+ in a query string can be read as a space, and / in a path is a separator. The URL-safe variant uses - instead of + and _ instead of /. The bit grouping is identical. Only the last two alphabet entries change names. = is also awkward in URLs, so some formats drop the trailing padding. Before decoding, restore it from the length: remainder 2 means add ==, remainder 3 means add =, remainder 1 is illegal.
JWT uses that variant: Base64URL, usually without padding. Pasting one segment into a decoder that only accepts +/ and required padding will fail. Swap the characters, restore =, and the bytes match standard Base64. The reverse is the same. Renaming the algorithm without changing the alphabet is not enough.
Data URLs and JWTs wrap another layer around Base64
Standard Base64
Alphabet includes + and /, and usually =.
- Fits inside a JSON string
- Awkward in a URL
- Decoding returns raw bytes
Base64URL
Alphabet includes - and _, padding often omitted.
- Each segment of a JWT
- Restore padding first
- Do not treat the signature as JSON
Data URL
data:type;base64, and then the payload.
- Everything before the comma is a header
- Decode only what follows
- The type says how to read the bytes
A small image embedded in a page is often data:image/png;base64,....... The media type and the encoding sit before the comma. Only the rest is Base64. Decoding the whole string, including data:, adds header bytes and the file signature will not match. Some text data URLs skip Base64 and use percent-encoding. Check the header for the base64 token before you choose a decoder.
When decoding fails, check in this order
The error usually only says the input is not valid Base64. It does not say whether the wrapper, the alphabet, or the padding is wrong. Walking the checks below is faster than swapping decoders until one of them happens to accept the string.
- Is a data URL, a pair of quotes, or a line break still attached? Peel back to alphabet characters.
- Which alphabet is it?
+/and-_should not both appear in one well-formed value. - Is padding the only thing missing? After removing
=, the length mod 4 must not be 1. - The result is bytes. If you expected text, decode those bytes as UTF-8. A failure there is an encoding problem, not a Base64 problem.
The Base64 tool on this site encodes and decodes standard Base64 over bytes. The JWT tool handles Base64URL and missing padding on its own. Do not paste the same string into both and expect one path: decide whether the text is going into a URL, into JSON, or back to file bytes.
Tools mentioned here
Keep reading
- How JSON parsing works, from characters to a valueFormat, minify, and tree view look like three buttons. Underneath they share one pipeline: split the text into tokens, then fold those tokens into a value. This note walks that path and shows where line numbers come from.
- What UTF-8 actually encodesCharacter counts, code points, and bytes are three different rulers. This note walks one code point through UTF-8 and explains why URL encoding and JSON escapes look nothing like each other.
- The three parts of a JWT, and what verification checksDecoding the header and payload does not mean the token can be trusted. This note separates Base64URL, the signature, and claim checks, and explains why the algorithm and the key have to be chosen together.