Streaming responses
Flip stream to true on your chat completion request, you receive incremental deltas over HTTP using the familiar SSE framing (lines beginning with data:). Parse until the sentinel [DONE].
https://api.rodiumai.io/v1/chat/completionsSSE sentinel
Examples
…Wire format
…Usage in the last chunk
The gateway always asks the provider for token usage (stream_options.include_usage is forced on, whatever you send). The last data: chunk before [DONE] therefore carries a usage object and an empty choices array.
…- Guard against the empty
choicesarray: indexchoices[0]only when it exists, as in the Python example above. usageholds token counts only, never a RODI amount. See Pricing and billing.- Responses are sent as
text/event-streamwithCache-Control: no-cacheand proxy buffering disabled. With smart routing, the selected model is also announced in theX-RodiumAI-Selected-Modelheader.
Anthropic Messages events
POST /v1/messages with stream: true emits native Anthropic events (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop), so the Anthropic SDK stream helpers work unchanged. Usage arrives in message_start (input) and message_delta (output). There is no [DONE] sentinel.
…Responses API events
POST /v1/responses with stream: true relays OpenAI Responses events: response.created, response.output_text.delta for text, then response.completed, whose response.usage carries the token counts.
…Errors and interruptions
- An error after the stream started arrives in-band: on chat completions a
data:event with anerrorobject followed bydata: [DONE]. Formats for every endpoint: Errors during a stream. - If you close the connection early, generation stops and you are billed for the input plus the output delivered up to then (details).
- Long generations can stay silent for a while before the first token (reasoning models): use a read timeout of several minutes.
Full SSE notes plus JSON chunk shapes live in API Streaming.