Smart routing
Send a virtual model id to POST /v1/chat/completions and the gateway picks a concrete catalogue model for each request.
Virtual model ids
| model | How the model is chosen | Extra hop |
|---|---|---|
rodiumai/smart (also smart, rodium/smart) | An LLM router reads your request and picks the best candidate | Yes, one router call |
rodium/basic or basic | Cheapest member of the profile (by RODI price) | No |
rodium/fast or fast | Fastest member (speed tier) | No |
rodium/pro or pro | First member by profile priority (balanced) | No |
rodium/max or max | Highest quality tier | No |
rodium/auto or auto | Quick heuristic: a vision model when the request has images, a coding model for code or tools, the highest tier for long or hard prompts (over 2,500 characters), otherwise the cheapest | No |
- Aliases are case-insensitive and only resolved by
POST /v1/chat/completions. Other endpoints (/v1/messages,/v1/responses, …) answer404 model_not_foundfor them. GET /v1/modelslists them first, withrodiumai_kindset tosmart_routerorsmart_profileandnullprices (you pay the selected model).- Every candidate respects the key's
allowed_models(and provided-credit scope). If none is left, the gateway answers403 smart_pool_empty(smart) or404 profile_empty(profiles).
How rodiumai/smart decides
- The gateway builds a candidate pool (the
autoprofile, up to 25 models) filtered by your key. - A router model (
openai/gpt-4o-miniby default) receives an excerpt of your prompt plus signals (images, tools,response_format, length) and returns the selected model, alternates, intent and confidence. It has 12 seconds. - Your request runs unchanged on the selected model, with streaming, tools and every other parameter you sent.
Fallback: if the router times out, fails, returns something unparseable or picks a model outside the pool, the gateway uses an alternate it proposed or the first active member of the pro profile, and sets fallback_used: true. The request is not failed because of the router.
Examples
…Request body
…rodiumai_routing
Non-streaming responses carry a rodiumai_routing object, and model is the model that answered.
…| Field | Meaning |
|---|---|
| requested | Canonical alias you sent (rodiumai/smart, rodium/fast, …) |
| router_model | Model used for the router hop; null for profiles |
| selected_model | Catalogue model that produced the answer |
| intent | Labels detected by the router (coding, chat, vision, reasoning, cheap, long_form, tools) |
| reason | Short explanation, or a fallback reason (router_timeout, router_upstream_error, router_parse_failed, …); for profiles the strategy result (fastest_by_speed_tier, …) |
| confidence | Router confidence between 0 and 1; 1.0 for profiles |
| alternates | Other suitable candidates proposed by the router |
| fallback_used | true when the router's choice could not be used |
| latency_router_ms | Duration of the router hop (smart only) |
| parent_request_id | Request id of the router hop (smart only) |
| strategy, profile | Profile strategy (cheapest, fastest, balanced, highest_quality, dynamic) and profile name (profiles only) |
Rule profile
…Streaming
With stream: true the chunks are those of the selected model; the routing decision is sent in response headers instead of the body: X-RodiumAI-Selected-Model and X-RodiumAI-Routing (the same object as JSON).
…Billing and limits
rodiumai/smartbills the router hop at the router model's price plus your request at the selected model's price. If the router call times out or fails, it is not billed.- Rule profiles bill only the selected model.
- The router hop also counts against the router model's rate limits for your key.
- Pin a concrete model id when you need predictable cost, latency or behaviour; use smart routing when prompts vary widely.
Same endpoint as regular chat: POST /v1/chat/completions. Pricing details: Pricing and billing.