Skip to content

Guides

Streaming

Get the answer while the model is still writing it: how to turn on streaming in each protocol, what the data looks like, and where usage appears.

With streaming, the model sends its answer back in small pieces while it is still generating, instead of returning everything at once when it is done. That is how chat apps make text appear word by word. For long answers, or models that think for a long time, streaming lets users see output sooner and keeps the connection from being cut by network equipment that drops idle connections.

Streaming and non-streaming requests cost exactly the same; only the way the answer is delivered differs.

All four protocols stream in SSE (Server-Sent Events) format: each message is a data: ... line, and messages are separated by a blank line.

Protocol How to turn it on How the stream ends
OpenAI Chat Completions Add "stream": true to the request body A final line data: [DONE]
OpenAI Responses Add "stream": true to the request body A response.completed event
Anthropic Messages Add "stream": true to the request body A message_stop event
Gemini Replace :generateContent in the URL with :streamGenerateContent?alt=sse The connection closes
终端窗口
# -N makes curl print each piece as it arrives instead of buffering
curl -N https://noviahub.com/v1/chat/completions \
-H "Authorization: Bearer $NOVIAHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"stream": true,
"messages": [{"role": "user", "content": "Write a four-line poem."}]
}'

Chat Completions: usage comes in the last chunk

Section titled “Chat Completions: usage comes in the last chunk”

When you stream through Chat Completions, NoviaHub by default sends one extra chunk with only usage just before the end: its choices is an empty array [] and its usage holds the token counts for the request. You get this chunk even if your request has no stream_options.

Here is a complete stream captured in a test environment (the content comes from a simulated upstream and only shows the format):

data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{"content":"Hello","role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{"content":" from the mock upstream."},"finish_reason":null}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":7,"total_tokens":19}}
data: [DONE]
  • If your program breaks on a chunk with empty choices, either check for empty choices in your code (as the Python sample does) or add "stream_options": {"include_usage": false} to the request; NoviaHub then doesn’t send the usage chunk.
  • Whether or not you ask for the usage chunk, NoviaHub charges for the actual usage.

Anthropic Messages: usage comes in message_delta

Section titled “Anthropic Messages: usage comes in message_delta”

An Anthropic-format stream sends message_start, content_block_start, a number of content_block_delta events, content_block_stop, message_delta and message_stop, in that order.

The usage in the message_delta event is the one to trust. When the model has to go through protocol conversion (for example, calling a GPT or DeepSeek model in Anthropic format), input_tokens in message_start is NoviaHub’s estimate at the start and may differ from the final count.

Each piece of a Gemini stream is a complete GenerateContentResponse object, and every piece has a usageMetadata field. The numbers in earlier pieces aren’t necessarily final (after protocol conversion, the input tokens in the first pieces are estimates and the output tokens are 0). Use the usageMetadata in the last piece.

  • Lines starting with a colon: in SSE, a line starting with : is a comment. The gateway can be configured to send such lines (: PING) during long waits to keep the connection alive. If you parse SSE yourself, skip lines that start with :; standard SSE client libraries do this for you.
  • Idle timeout: if the upstream model sends nothing for a long time, the gateway ends the streamed response.
  • Telling them apart in the logs: the Timing column in Usage logs marks each request Stream or Non-stream. Streaming requests also show First token, the time from sending the request to receiving the first piece of content.