Notes
How an SSE stream reassembles a model reply from deltas
Model APIs often split one reply across many SSE events. This note covers the frame format, how deltas merge by path, and why tool-call arguments are concatenated strings rather than JSON on every frame.
By Brook/Updated 2026-10-09/13 min read
A stream is not one JSON document
Pasting a streamed response into a JSON formatter should fail. The text is many frames, separated by blank lines, each with its own data: line. One line may be JSON. The whole blob is not. Split frames first, then parse the data of each frame. The order matters.
Server-Sent Events are a one-way text push. The connection stays open and the server writes frames; the client handles each frame as it arrives. Unlike a WebSocket, the client does not send messages back on that connection. A typewriter effect only needs this: every time the model produces a bit more, the server pushes another frame.
What one frame looks like
- 01BytesRaw textIncludes data: prefixes, blank lines, and [DONE].
- 02SplitEventsBlank lines separate frames. Colon lines are comments.
- 03ParseJSON per frameA bad frame is kept aside. It does not sink the rest.
- 04MergeReply and tool callsAppend each delta onto the text you already have.
A frame is some field: value lines ended by a blank line. Model APIs usually send only data:. The space after data: is optional, and the value is usually JSON. Some servers also send event: to name the event. A line that starts with a colon is a comment and must be skipped. Two consecutive newlines end the frame. A single newline inside the frame only separates fields.
The end of a stream is often data: [DONE]. That is not JSON, so the parser has to recognize it before JSON.parse. If one frame contains broken JSON, record the raw frame and its position, then continue. One bad frame does not mean the rest of the reply is gone. A debugger that collapses all of that into a single failure hides which delta was corrupt.
Deltas append along a path. They do not replace
In the OpenAI-style stream, each frame is a chat.completion.chunk. The new text lives in choices[].delta, not in a full message. delta.content is the next piece of the reply, sometimes a single character. Merge by choice index: append the piece to the text already built for that choice. A later frame with no content means "nothing new this tick". It does not clear what you have.
Some APIs put reasoning in reasoning_content or a sibling field, separate from the user-visible reply. Keep those buffers separate. If you append both into one string, you cannot fold the reasoning away later. usage often arrives only on the last frame. Store it on its own. Do not require every frame to carry it.
- Append content for the same choice in arrival order.
- Do not merge different choice indexes into one reply.
- A missing field means "no delta", not "set the string to empty".
- A later finish_reason replaces an empty one. It does not rewrite the text you already appended.
Tool-call arguments are string fragments
When the model calls a tool, deltas contain tool_calls. Each item has an index. function.name often arrives whole in an early frame. function.arguments arrives as pieces of JSON text, and many frames are only { or "city":. Those pieces are not legal JSON until the stream ends. Merge by appending arguments as a string under that index, and parse once after the last frame. A parse error halfway through is normal. It is not an API failure.
- 01index 0Name is completefunction.name arrives on an early frame. Later frames may omit it.
- 02index 0Arguments are still fragmentsEach frame appends the next piece of arguments. Do not parse yet.
- 03AfterParse onceThe joined string is the JSON document. Failures before that can be ignored.
The index is the identity of the call. Some frames carry only the index and a short argument piece, without repeating the name. Keying by name collides when the same tool is called twice. Creating a new call on every frame shatters the arguments. The table you want is "choice index plus tool_calls index".
Other vendors change the path, not the merge
Anthropic streams use content_block_delta. Text arrives in delta.text, tool arguments in delta.partial_json, and the block identity is index. Gemini-style payloads often put the delta in candidates[].content.parts. The fields differ. The action is the same: find the identity of the block and the new text, then append. Do not replace.
Some APIs skip SSE and send one JSON object per line (NDJSON). The separator is a newline, there is no data: prefix, and there is often no [DONE]. A debugger can look at the first lines. If they start with data:, split SSE. If each line parses as JSON on its own, split NDJSON. Neither shape is "one JSON document for the whole body".
Compare frames first, then the sentence you built
A reply missing its ending is often a frame that was not recognized as data:, or a lost blank line that glued two frames together. A tool-argument parse error should be inspected on the concatenated string, not on the last frame alone. Reasoning mixed into the reply usually means both fields were written into one buffer. That is the client misreading the stream, not the model "answering badly".
The SSE debugger on this site follows that order: split frames, recognize [DONE] and comments, parse each frame, merge reply text, reasoning, and tool arguments by path, and keep frames that failed to parse. The input stays in the browser. Treat it as a reference: your client should build the same result under the same rules.
Tools mentioned here
Keep reading
- How JSON parsing works, from characters to a valueFormat, minify, and tree view look like three buttons. Underneath they share one pipeline: split the text into tokens, then fold those tokens into a value. This note walks that path and shows where line numbers come from.
- What UTF-8 actually encodesCharacter counts, code points, and bytes are three different rulers. This note walks one code point through UTF-8 and explains why URL encoding and JSON escapes look nothing like each other.
- The three parts of a JWT, and what verification checksDecoding the header and payload does not mean the token can be trusted. This note separates Base64URL, the signature, and claim checks, and explains why the algorithm and the key have to be chosen together.