Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Guides

Chat completions

Our primary endpoint, assemble messages[], set model to a provider-scoped id (e.g. openai/gpt-4o), tune decoding, and optionally stream SSE.

POSThttps://api.rodiumai.io/v1/chat/completions

When to use

  • Assistants, copilots, and chat UIs that need multi-turn text.
  • Coding agents that generate or refactor full files.
  • Vision: describe or reason over images in the same request.
  • Let rodiumai/smart pick a strong model when you do not want to hardcode one.

Recipes

Coding agent

Ask for a complete Express + JWT server (or any scaffold). Prefer models tagged coding, or model rodiumai/smart.

Note: Keep max_tokens high enough for full files; stream for UX.

Vision (image_url)

Pass a user message with content parts: text + image_url (data URL or HTTPS). Works on vision-capable catalogue models.

Note: Huge images increase prompt tokens, resize when possible.

Tools / function calling

Send tools + tool_choice like OpenAI. Round-trip tool results as role tool messages.

Note: On Anthropic upstream via chat, prefer stream:false for tool rounds.

OpenAI SDK mode + cURL

…

Parameters

ParameterTypeRequiredDescription
modelstringRequiredCatalogue id (e.g. openai/gpt-4o, anthropic/claude-sonnet-4-6), a smart alias (rodiumai/smart, rodium/fast, …) or one of your custom models. List ids with GET /v1/models.
messagesarrayRequiredArray of message objects (system, user, assistant, tool) with role and content. On vision models user content can mix text and image_url parts.
max_tokensintegerOptionalMaximum tokens to generate. Also sizes the pre-flight hold (2,048 when omitted). When omitted, OpenAI models use their own maximum and other providers receive 4,096.
max_completion_tokensintegerOptionalAlias of max_tokens used by newer OpenAI clients; read when max_tokens is absent. On OpenAI reasoning models the gateway sends max_completion_tokens for you.
temperaturenumberOptionalSampling temperature between 0 and 2. GPT-5 and o-series models only accept the default (1); other values are dropped.
top_pnumberOptionalNucleus sampling probability mass, between 0 and 1. Dropped on GPT-5 and o-series models.
streambooleanOptionalIf true, stream deltas as Server-Sent Events. The last chunk carries usage.
stopstring | arrayOptionalSequences where generation stops.
toolsarrayOptionalOpenAI function tools. Translated for Claude and Gemini models; send results back as role tool messages.
tool_choicestring | objectOptionalauto, none, required or a specific function, as in OpenAI.
response_formatobjectOptionaljson_object or json_schema output when the model supports JSON mode (supports_json_mode in GET /v1/models).
nintegerOptionalNumber of choices. Each choice is billed as output; only OpenAI-compatible providers honour values above 1.
session_idstringOptionalRodiumAI extension for custom models: conversation id for short-term memory. Ignored by catalogue models.

Sample JSON

…
…

Output limit and defaults

  • Set max_tokens (or its alias max_completion_tokens) on every request: it bounds the answer, the cost and the pre-flight hold.
  • When you omit both, models served by OpenAI use their own maximum, while models served by other providers (Claude, Gemini, Llama, …) receive 4,096 output tokens. The pre-flight hold assumes 2,048 in that case.
  • On OpenAI reasoning models (GPT-5, o-series) the gateway forwards your limit as max_completion_tokens and drops sampling parameters they reject (temperature other than 1, top_p, penalties, logit_bias, logprobs). Reasoning tokens count as output.

Differences between providers

One request shape works for every catalogue model, but many models are served through providers with their own API (Anthropic, Google Gemini, Amazon Bedrock). For those the gateway translates the request and the response back to the OpenAI format.

FieldOpenAI-compatible providersClaude, Gemini, Bedrock-hosted models
messages (system, user, assistant, tool)ForwardedTranslated
max_tokens / max_completion_tokensForwardedTranslated (4,096 when omitted)
temperature, top_p, stopForwardedTranslated
tools, tool_choiceForwardedTranslated
response_formatForwardedGemini only
streamForwarded (usage forced on)Translated to OpenAI chunks
n, seed, logit_bias, logprobs, frequency_penalty, presence_penalty, user, parallel_tool_calls, store, prediction, audio, modalities, reasoning_effortForwarded when the model supports themNot forwarded

service_tier is forwarded only as auto, default or flex; priority, fast, ultrafast and scale are dropped and the request runs at the standard tier.

Response

  • model in a non-streaming response is the catalogue id that answered (the selected model when you used smart routing).
  • usage holds prompt_tokens, completion_tokens and total_tokens, plus cached and reasoning details when the provider reports them. There is no RODI amount: see how to get the cost.
  • rodiumai_routing is added only for smart aliases (Smart routing).
  • Images in image_url parts can be https:// URLs or data:image/…;base64,… URLs; a malformed data URL returns 400 invalid_value with param pointing at the part.

Endpoint reference: POST /v1/chat/completions · Anthropic Messages