Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Guides

Streaming responses

Flip stream to true on your chat completion request, you receive incremental deltas over HTTP using the familiar SSE framing (lines beginning with data:). Parse until the sentinel [DONE].

POSThttps://api.rodiumai.io/v1/chat/completions

Examples

…

Wire format

…

Usage in the last chunk

The gateway always asks the provider for token usage (stream_options.include_usage is forced on, whatever you send). The last data: chunk before [DONE] therefore carries a usage object and an empty choices array.

…
  • Guard against the empty choices array: index choices[0] only when it exists, as in the Python example above.
  • usage holds token counts only, never a RODI amount. See Pricing and billing.
  • Responses are sent as text/event-stream with Cache-Control: no-cache and proxy buffering disabled. With smart routing, the selected model is also announced in the X-RodiumAI-Selected-Model header.

Anthropic Messages events

POST /v1/messages with stream: true emits native Anthropic events (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop), so the Anthropic SDK stream helpers work unchanged. Usage arrives in message_start (input) and message_delta (output). There is no [DONE] sentinel.

…

Responses API events

POST /v1/responses with stream: true relays OpenAI Responses events: response.created, response.output_text.delta for text, then response.completed, whose response.usage carries the token counts.

…

Errors and interruptions

  • An error after the stream started arrives in-band: on chat completions a data: event with an error object followed by data: [DONE]. Formats for every endpoint: Errors during a stream.
  • If you close the connection early, generation stops and you are billed for the input plus the output delivered up to then (details).
  • Long generations can stay silent for a while before the first token (reasoning models): use a read timeout of several minutes.

Full SSE notes plus JSON chunk shapes live in API Streaming.