POST
https://api.rodiumai.io/v1/chat/completionsCreate chat completion
Generates a model response from a conversational message list, fields and streaming semantics are defined for RodiumAi's API.
Request body parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | Catalogue id (e.g. openai/gpt-4o, anthropic/claude-sonnet-4-6), a smart alias (rodiumai/smart, rodium/fast, …) or one of your custom models. List ids with GET /v1/models. |
| messages | array | Required | Array of message objects (system, user, assistant, tool) with role and content. On vision models user content can mix text and image_url parts. |
| max_tokens | integer | Optional | Maximum tokens to generate. Also sizes the pre-flight hold (2,048 when omitted). When omitted, OpenAI models use their own maximum and other providers receive 4,096. |
| max_completion_tokens | integer | Optional | Alias of max_tokens used by newer OpenAI clients; read when max_tokens is absent. On OpenAI reasoning models the gateway sends max_completion_tokens for you. |
| temperature | number | Optional | Sampling temperature between 0 and 2. GPT-5 and o-series models only accept the default (1); other values are dropped. |
| top_p | number | Optional | Nucleus sampling probability mass, between 0 and 1. Dropped on GPT-5 and o-series models. |
| stream | boolean | Optional | If true, stream deltas as Server-Sent Events. The last chunk carries usage. |
| stop | string | array | Optional | Sequences where generation stops. |
| tools | array | Optional | OpenAI function tools. Translated for Claude and Gemini models; send results back as role tool messages. |
| tool_choice | string | object | Optional | auto, none, required or a specific function, as in OpenAI. |
| response_format | object | Optional | json_object or json_schema output when the model supports JSON mode (supports_json_mode in GET /v1/models). |
| n | integer | Optional | Number of choices. Each choice is billed as output; only OpenAI-compatible providers honour values above 1. |
| session_id | string | Optional | RodiumAI extension for custom models: conversation id for short-term memory. Ignored by catalogue models. |
OpenAI SDK mode (Python, JavaScript)
…Example request JSON
…Example response JSON
…Output limit and defaults
- Set
max_tokens(or its aliasmax_completion_tokens) on every request: it bounds the answer, the cost and the pre-flight hold. - When you omit both, models served by OpenAI use their own maximum, while models served by other providers (Claude, Gemini, Llama, …) receive 4,096 output tokens. The pre-flight hold assumes 2,048 in that case.
- On OpenAI reasoning models (GPT-5, o-series) the gateway forwards your limit as
max_completion_tokensand drops sampling parameters they reject (temperatureother than 1,top_p, penalties,logit_bias,logprobs). Reasoning tokens count as output.
Differences between providers
One request shape works for every catalogue model, but many models are served through providers with their own API (Anthropic, Google Gemini, Amazon Bedrock). For those the gateway translates the request and the response back to the OpenAI format.
| Field | OpenAI-compatible providers | Claude, Gemini, Bedrock-hosted models |
|---|---|---|
| messages (system, user, assistant, tool) | Forwarded | Translated |
| max_tokens / max_completion_tokens | Forwarded | Translated (4,096 when omitted) |
| temperature, top_p, stop | Forwarded | Translated |
| tools, tool_choice | Forwarded | Translated |
| response_format | Forwarded | Gemini only |
| stream | Forwarded (usage forced on) | Translated to OpenAI chunks |
| n, seed, logit_bias, logprobs, frequency_penalty, presence_penalty, user, parallel_tool_calls, store, prediction, audio, modalities, reasoning_effort | Forwarded when the model supports them | Not forwarded |
service_tier is forwarded only as auto, default or flex; priority, fast, ultrafast and scale are dropped and the request runs at the standard tier.
Tools on Claude through chat
Tool calling works through
/v1/chat/completions on Claude models. For long agentic loops with streaming, the native Anthropic Messages endpoint keeps every Anthropic feature.Response
modelin a non-streaming response is the catalogue id that answered (the selected model when you used smart routing).usageholdsprompt_tokens,completion_tokensandtotal_tokens, plus cached and reasoning details when the provider reports them. There is no RODI amount: see how to get the cost.rodiumai_routingis added only for smart aliases (Smart routing).- Images in
image_urlparts can behttps://URLs ordata:image/…;base64,…URLs; a malformed data URL returns400 invalid_valuewithparampointing at the part.
Streamed deltas → Streaming reference