POST
https://api.rodiumai.io/v1/responsesAPI
Create response
Opaque OpenAI Responses passthrough. Set stream:true for response.output_text.delta SSE events. Billed in RODI.
Auth
Bearer rd_sk_… only (Authorization header). model is required and must resolve to an OpenAI upstream.
Request body parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | Catalogue id served by an OpenAI upstream (e.g. openai/gpt-4o). |
| input | string | array | Required | Prompt text or Responses input items. |
| stream | boolean | Optional | When true, returns SSE events (response.output_text.delta, …). |
| instructions | string | Optional | Optional system-level instructions passed through to the model. |
| max_output_tokens | integer | Optional | Upper bound on generated tokens, reasoning included. Bounds the cost; set it on every call. |
| temperature | number | Optional | Sampling temperature when supported. |
Examples
…RodiumAi SDK
…Streaming
…Scope of /v1/responses
- Served for OpenAI models only (for example
openai/gpt-4o,openai/gpt-5.1,openai/gpt-5.1-codex). Other models are not available on this endpoint: use/v1/chat/completions. - The body is passed to the provider as-is, apart from
model; fields such asinstructions,toolsorreasoningwork as documented by OpenAI.service_tiervaluespriority,fast,ultrafastandscaleare not forwarded (the request runs at the standard tier);auto,defaultandflexare. background: trueandconversationare rejected with400 unsupported_parameter: usestream: truefor long runs.previous_response_idworks only with the id of a response your account created through RodiumAI; any other id returns404 previous_response_not_found(param: previous_response_id).- Provider web searches (
web_search_calloutput items) are billed per call on top of tokens, see Pricing. stream: truereturns the Responses event stream (response.output_text.delta, thenresponse.completedwithusage).- Authenticate with
Authorization: Beareronly. Smart aliases and custom models are not supported here. - Set
max_output_tokensto bound the cost; reasoning tokens are billed as output.
Streamed deltas → Streaming reference