Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
API

Streaming

Set stream:true on chat completions to receive SSE events. Each meaningful payload arrives as its own data: line terminating with \n.

POSThttps://api.rodiumai.io/v1/chat/completions

Terminate consumption when encountering the literal sentinel data: [DONE].

Examples

…

Wire format

…

Usage in the last chunk

The gateway always asks the provider for token usage (stream_options.include_usage is forced on, whatever you send). The last data: chunk before [DONE] therefore carries a usage object and an empty choices array.

…
  • Guard against the empty choices array: index choices[0] only when it exists, as in the Python example above.
  • usage holds token counts only, never a RODI amount. See Pricing and billing.
  • Responses are sent as text/event-stream with Cache-Control: no-cache and proxy buffering disabled. With smart routing, the selected model is also announced in the X-RodiumAI-Selected-Model header.

Anthropic Messages events

POST /v1/messages with stream: true emits native Anthropic events (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop), so the Anthropic SDK stream helpers work unchanged. Usage arrives in message_start (input) and message_delta (output). There is no [DONE] sentinel.

…

Responses API events

POST /v1/responses with stream: true relays OpenAI Responses events: response.created, response.output_text.delta for text, then response.completed, whose response.usage carries the token counts.

…

Errors and interruptions

  • An error after the stream started arrives in-band: on chat completions a data: event with an error object followed by data: [DONE]. Formats for every endpoint: Errors during a stream.
  • If you close the connection early, generation stops and you are billed for the input plus the output delivered up to then (details).
  • Long generations can stay silent for a while before the first token (reasoning models): use a read timeout of several minutes.

Conceptual primer: Streaming guide