Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Guides

Smart routing

Send a virtual model id to POST /v1/chat/completions and the gateway picks a concrete catalogue model for each request.

Virtual model ids

modelHow the model is chosenExtra hop
rodiumai/smart (also smart, rodium/smart)An LLM router reads your request and picks the best candidateYes, one router call
rodium/basic or basicCheapest member of the profile (by RODI price)No
rodium/fast or fastFastest member (speed tier)No
rodium/pro or proFirst member by profile priority (balanced)No
rodium/max or maxHighest quality tierNo
rodium/auto or autoQuick heuristic: a vision model when the request has images, a coding model for code or tools, the highest tier for long or hard prompts (over 2,500 characters), otherwise the cheapestNo
  • Aliases are case-insensitive and only resolved by POST /v1/chat/completions. Other endpoints (/v1/messages, /v1/responses, …) answer 404 model_not_found for them.
  • GET /v1/models lists them first, with rodiumai_kind set to smart_router or smart_profile and null prices (you pay the selected model).
  • Every candidate respects the key's allowed_models (and provided-credit scope). If none is left, the gateway answers 403 smart_pool_empty (smart) or 404 profile_empty (profiles).

How rodiumai/smart decides

  1. The gateway builds a candidate pool (the auto profile, up to 25 models) filtered by your key.
  2. A router model (openai/gpt-4o-mini by default) receives an excerpt of your prompt plus signals (images, tools, response_format, length) and returns the selected model, alternates, intent and confidence. It has 12 seconds.
  3. Your request runs unchanged on the selected model, with streaming, tools and every other parameter you sent.

Fallback: if the router times out, fails, returns something unparseable or picks a model outside the pool, the gateway uses an alternate it proposed or the first active member of the pro profile, and sets fallback_used: true. The request is not failed because of the router.

Examples

…

Request body

…

rodiumai_routing

Non-streaming responses carry a rodiumai_routing object, and model is the model that answered.

…
FieldMeaning
requestedCanonical alias you sent (rodiumai/smart, rodium/fast, …)
router_modelModel used for the router hop; null for profiles
selected_modelCatalogue model that produced the answer
intentLabels detected by the router (coding, chat, vision, reasoning, cheap, long_form, tools)
reasonShort explanation, or a fallback reason (router_timeout, router_upstream_error, router_parse_failed, …); for profiles the strategy result (fastest_by_speed_tier, …)
confidenceRouter confidence between 0 and 1; 1.0 for profiles
alternatesOther suitable candidates proposed by the router
fallback_usedtrue when the router's choice could not be used
latency_router_msDuration of the router hop (smart only)
parent_request_idRequest id of the router hop (smart only)
strategy, profileProfile strategy (cheapest, fastest, balanced, highest_quality, dynamic) and profile name (profiles only)

Rule profile

…

Streaming

With stream: true the chunks are those of the selected model; the routing decision is sent in response headers instead of the body: X-RodiumAI-Selected-Model and X-RodiumAI-Routing (the same object as JSON).

…

Billing and limits

  • rodiumai/smart bills the router hop at the router model's price plus your request at the selected model's price. If the router call times out or fails, it is not billed.
  • Rule profiles bill only the selected model.
  • The router hop also counts against the router model's rate limits for your key.
  • Pin a concrete model id when you need predictable cost, latency or behaviour; use smart routing when prompts vary widely.

Same endpoint as regular chat: POST /v1/chat/completions. Pricing details: Pricing and billing.