Guides
Streaming
Get the answer while the model is still writing it: how to turn on streaming in each protocol, what the data looks like, and where usage appears.
With streaming, the model sends its answer back in small pieces while it is still generating, instead of returning everything at once when it is done. That is how chat apps make text appear word by word. For long answers, or models that think for a long time, streaming lets users see output sooner and keeps the connection from being cut by network equipment that drops idle connections.
Streaming and non-streaming requests cost exactly the same; only the way the answer is delivered differs.
How to turn it on
Section titled “How to turn it on”All four protocols stream in SSE (Server-Sent Events) format: each message is a data: ... line, and messages are separated by a blank line.
| Protocol | How to turn it on | How the stream ends |
|---|---|---|
| OpenAI Chat Completions | Add "stream": true to the request body |
A final line data: [DONE] |
| OpenAI Responses | Add "stream": true to the request body |
A response.completed event |
| Anthropic Messages | Add "stream": true to the request body |
A message_stop event |
| Gemini | Replace :generateContent in the URL with :streamGenerateContent?alt=sse |
The connection closes |
# -N makes curl print each piece as it arrives instead of bufferingcurl -N https://noviahub.com/v1/chat/completions \ -H "Authorization: Bearer $NOVIAHUB_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-v4-flash", "stream": true, "messages": [{"role": "user", "content": "Write a four-line poem."}] }'import os
from openai import OpenAI
client = OpenAI( base_url="https://noviahub.com/v1", api_key=os.environ["NOVIAHUB_API_KEY"],)
stream = client.chat.completions.create( model="deepseek-v4-flash", messages=[{"role": "user", "content": "Write a four-line poem."}], stream=True,)
for chunk in stream: # The last chunk carries only usage and has no choices, so check first if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True)Chat Completions: usage comes in the last chunk
Section titled “Chat Completions: usage comes in the last chunk”When you stream through Chat Completions, NoviaHub by default sends one extra chunk with only usage just before the end: its choices is an empty array [] and its usage holds the token counts for the request. You get this chunk even if your request has no stream_options.
Here is a complete stream captured in a test environment (the content comes from a simulated upstream and only shows the format):
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{"content":"Hello","role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{"content":" from the mock upstream."},"finish_reason":null}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-mock123","object":"chat.completion.chunk","created":1790609866,"model":"deepseek-v4-flash","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":7,"total_tokens":19}}
data: [DONE]- If your program breaks on a chunk with empty
choices, either check for emptychoicesin your code (as the Python sample does) or add"stream_options": {"include_usage": false}to the request; NoviaHub then doesn’t send the usage chunk. - Whether or not you ask for the usage chunk, NoviaHub charges for the actual usage.
Anthropic Messages: usage comes in message_delta
Section titled “Anthropic Messages: usage comes in message_delta”An Anthropic-format stream sends message_start, content_block_start, a number of content_block_delta events, content_block_stop, message_delta and message_stop, in that order.
The usage in the message_delta event is the one to trust. When the model has to go through protocol conversion (for example, calling a GPT or DeepSeek model in Anthropic format), input_tokens in message_start is NoviaHub’s estimate at the start and may differ from the final count.
Gemini: usage comes in usageMetadata
Section titled “Gemini: usage comes in usageMetadata”Each piece of a Gemini stream is a complete GenerateContentResponse object, and every piece has a usageMetadata field. The numbers in earlier pieces aren’t necessarily final (after protocol conversion, the input tokens in the first pieces are estimates and the output tokens are 0). Use the usageMetadata in the last piece.
Other things to know
Section titled “Other things to know”- Lines starting with a colon: in SSE, a line starting with
:is a comment. The gateway can be configured to send such lines (: PING) during long waits to keep the connection alive. If you parse SSE yourself, skip lines that start with:; standard SSE client libraries do this for you. - Idle timeout: if the upstream model sends nothing for a long time, the gateway ends the streamed response.
- Telling them apart in the logs: the Timing column in Usage logs marks each request Stream or Non-stream. Streaming requests also show First token, the time from sending the request to receiving the first piece of content.