Notes
How JSON parsing works, from characters to a value
Format, minify, and tree view look like three buttons. Underneath they share one pipeline: split the text into tokens, then fold those tokens into a value. This note walks that path and shows where line numbers come from.
By Brook/Updated 2026-10-09/12 min read
A parser sees characters, not a tree
When you paste JSON into an editor, you already see objects, arrays, and fields. A parser does not. It receives a sequence of Unicode code points mixed with spaces, newlines, quotes, colons, and backslashes. Before it can say "this is an object", it has to decide what kind of token each stretch of characters belongs to.
The format button is not decorating the text you pasted. It throws the text away as structure, keeps the value in memory, and prints that value again with the indent you picked. Minify is a second print of the same value, with the spaces left out. If either step is wrong, the button looks like it "did something odd".
Split tokens before you talk about structure
The lexer answers one question: what is the next token? It does not care whether that token is legal in context. A " enters string mode until an unescaped closing quote. A - or a digit enters number mode until the next character cannot belong to a number. { } [ ] : and , are single-character tokens. true, false, and null must match the whole word. One extra letter is an illegal token.
- 01InputCharactersWhitespace, newlines, and escapes. No structure yet.
- 02LexTokensStrings, numbers, literals, and separators.
- 03ParseValue treeObjects, arrays, and scalars nested by the grammar.
- 04PrintText or viewIndented text, minified text, or an expanded tree.
Whitespace is dropped at this layer. JSON allows spaces, tabs, newlines, and carriage returns between tokens. It does not allow comments outside strings. // and /* */ are lexer errors, not a "lenient JSON". Samples copied from docs often include comments, and the failure lands on the first slash, which is the right place.
A stack folds tokens into a value
The parser walks a short grammar. A value is an object, an array, a string, a number, true, false, or null. An object is string keys, colons, and values, separated by commas. An array is values separated by commas. The usual implementation keeps a stack of "which container am I filling, and am I waiting for a key or a value?".
{ pushes a new object. } pops it, and the container must end on a complete member. A trailing comma is rejected because a comma must be followed by another member. That is stricter than a JavaScript object literal. JSON.parse is supposed to be stricter. A tool that quietly runs eval or Function will also accept comments, trailing commas, and code. That is not a JSON parser.
- Keys must be strings. In
{count: 1}the bare word is an illegal token. - A single top-level value is legal.
42and"ok"both parse. Formatting does not wrap them in an object. - Commas, colons, and brackets must pair. A missing one stops on that token, not at the end of the file.
Numbers, strings, and escapes fail in specific ways
A backslash in a string starts a short list of escapes: quote, backslash, slash, backspace, form feed, newline, carriage return, tab, and \u plus exactly four hex digits. \x61 and a backslash before a single quote are JavaScript, not JSON. A raw newline inside a string keeps string mode open, so the reported error often lands on a later quote. The opening mistake is earlier than the message looks.
Numbers cannot have leading zeros, and NaN and Infinity are not JSON numbers. Precision is the sharper edge. JSON does not fix an integer width. JavaScript numbers are IEEE 754 doubles, so integers past 2^53-1 change when parsed. Snowflake ids and long order numbers should stay strings, or go through a parser that keeps big integers. If a formatter runs ordinary JSON.parse first, pretty indent cannot restore digits that are already gone.
Three tools are three exits from one pipeline
Formatting wants a reprint. Indent changes only the print. Minify is the same reprint without whitespace between tokens. A tree view stops on the value: paths, types, and children, not the original spaces. All three can share one parse, and a failure should be the same failure for all three.
- 01TreeExpand the value by pathObjects and arrays stay nested. A JSON string inside a field is a second pass of the same pipeline.
- 02FormatPrint with indentTwo spaces, four spaces, or tabs change whitespace only. Key order depends on whether the implementation kept it.
- 03MinifyDrop whitespace between tokensSpaces inside strings stay. Minify is not encryption, and it does not shorten keys.
- 04ParseLex, then parseA failure carries a token offset. Line and column are converted from that offset.
JSON embedded in a string is another layer. Log fields and queue payloads are often "a string whose contents happen to be JSON". After the outer parse, that field is still a string. A tree view that drills in must parse the string again and append the new path to the old one. Outer failures and inner failures should be reported separately.
Boundaries worth keeping
A parser should reject syntax it does not know, rather than guess. Accepting one trailing comma means the next broken document gets "fixed" into a different meaning. Errors need a position: a 1-based line and column, and ideally a short slice of the source. "Parse failed" with no position is not actionable.
Test printing separately from parsing. Parse tests check that illegal input is rejected and legal input becomes the right value. Print tests check indent, key order, and Unicode escapes. Mixed together, two strings that "look close" can hide a value that has already changed. The JSON tools on this site follow that split: parse in the browser, point at the line on failure, then print or expand.
Tools mentioned here
Keep reading
- What UTF-8 actually encodesCharacter counts, code points, and bytes are three different rulers. This note walks one code point through UTF-8 and explains why URL encoding and JSON escapes look nothing like each other.
- The three parts of a JWT, and what verification checksDecoding the header and payload does not mean the token can be trusted. This note separates Base64URL, the signature, and claim checks, and explains why the algorithm and the key have to be chosen together.
- How an SSE stream reassembles a model reply from deltasModel APIs often split one reply across many SSE events. This note covers the frame format, how deltas merge by path, and why tool-call arguments are concatenated strings rather than JSON on every frame.