What is rodiumai/smart?
rodiumai/smart is a virtual model on the RodiumAI gateway. You do not call a fixed upstream (Claude, GPT, Gemini…). You call rodiumai/smart, and RodiumAI selects the ideal catalogue model for that request.
The goal is simple:
Save cost when a short or easy prompt does not need a frontier model.
Keep quality when the task needs coding, reasoning, vision, or tools.
Stay fast by preferring flash / mini / lite models when latency matters more than peak intelligence.
It is the same OpenAI-shaped API you already use (chat.completions), with one change: the model field is rodiumai/smart (aliases: smart, rodium/smart).
How it works
When you send a request with model: "rodiumai/smart", the gateway:
Reads signals from your payload: message text, images, tools, approximate length.
Calls a router model (default
openai/gpt-4o-mini, configurable) with a short candidate list from the catalogue pool (built around theautosmart profile plus complements).Picks one catalogue slug (for example
google/gemini-2.5-flashoranthropic/claude-sonnet-4-6).Forwards your request to that model and returns the normal chat completion.
The router is instructed to:
Prefer coding / reasoning models for code, algorithms, debugging.
Prefer vision-capable models when images are present.
Prefer cheaper / faster models (flash, mini, nano, lite) for short Q&A.
Prefer higher quality models for hard reasoning, long-form design, or agentic tool use.
Never invent a slug: the choice must be one of the candidates.
Both hops are billed in RODI: the small router hop, then the selected model hop. On non-stream responses you get routing metadata (rodiumai_routing). On stream, look at headers such as X-RodiumAI-Selected-Model.
Your app → rodiumai/smart → router picks slug → real model (Claude / GPT / Gemini / …)
Why use smart instead of a fixed model?
1. One key, many vendors
You keep a single rd_sk_… key and a single base_url. Smart decides whether this turn needs Haiku-class speed, Sonnet-class coding, or a Flash-class bargain.
2. Cost control without micromanagement
Pinning every call to Opus or GPT-5 burns wallet balance on greetings and tiny edits. Smart routes those to affordable models and reserves frontier capacity for harder work.
3. Speed when it matters
For short prompts, the pool includes low-latency options (Gemini Flash, GPT mini/nano, DeepSeek Flash, Claude Haiku). You get answers sooner without rewriting your client.
4. OpenAI-compatible drop-in
If your stack already uses the OpenAI SDK, Cursor, or any OpenAI-compatible agent, point base_url to RodiumAI and set model="rodiumai/smart". No new SDK.
5. Local payment rails
Same RodiumAI wallet: Mobile Money, bank transfer, RODI billing, usage in the dashboard. Smart does not change how you pay, only which model runs.
Quick start
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("RODIUMAI_API_KEY"),
base_url="https://api.rodiumai.io/v1",
)
response = client.chat.completions.create(
model="rodiumai/smart",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=256,
)
print(response.choices[0].message.content)
print(response.model) # actual catalogue slug selected for this turn
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RODIUMAI_API_KEY,
baseURL: "https://api.rodiumai.io/v1",
});
const response = await client.chat.completions.create({
model: "rodiumai/smart",
messages: [{ role: "user", content: "Hello!" }],
max_tokens: 256,
});
console.log(response.choices[0].message.content);
console.log(response.model);
curl https://api.rodiumai.io/v1/chat/completions \
-H "Authorization: Bearer rd_sk_…" \
-H "Content-Type: application/json" \
-d '{
"model": "rodiumai/smart",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256
}'
Tip: Log
response.model(or the routing headers) in production. It is the best way to see which slug smart chose and to tune prompts or fall back to a pinned model when needed.
Smart vs rule-based profiles
RodiumAI also ships rule-based smart profiles (no LLM hop). They resolve deterministically from a SmartModelProfile:
ProfileStrategyBest forbasiccheapestHigh volume, low stakesfastfastestLowest latencyprobalancedDefault quality / pricemaxhighest qualityFrontier reasoningautodynamic heuristicsPrompt-shaped routing without a router LLM
rodiumai/smart is different: it uses an LLM router over a candidate pool for finer intent matching (coding vs chat vs vision vs cheap). Profiles like fast or pro are cheaper on the routing hop itself (no classifier call) but less adaptive turn by turn.
NeedSuggested callAdaptive pick every requestrodiumai/smartAlways prioritize latencyrodiumai/fast (or fast)Always prioritize costrodiumai/basicAlways prioritize qualityrodiumai/maxKnow the exact vendor modelPin anthropic/claude-sonnet-4-6, openai/gpt-5.4-mini, etc.
Example scenarios
User requestWhat smart tends to prefer"Hello!" / short FAQCheap, fast models (Flash / mini / Haiku-class)Full Express + JWT server codeStronger coding / reasoning modelsScreenshot or chart analysisVision-capable modelsLong design / multi-step planHigher quality / reasoning members of the poolTool-heavy agent turnModels with solid tool support in the pool
You still pay for what runs. Smart reduces waste: frontier rates only when the prompt warrants them.
Billing notes
Router hop + selected model hop are both charged in RODI.
On simple prompts the router cost is small compared to avoiding a frontier completion.
Check live RODI rates with
GET /v1/modelsor the models catalogue.Recharge and usage stay in the RodiumAI dashboard.
When not to use smart
You need a guaranteed vendor slug for compliance or evals: pin the model.
You need zero routing hop latency and a fixed strategy: use
fast,basic,pro, ormax.The workload is image or video generation: use
POST /v1/images/generationsorPOST /v1/videos/generationswith the right model, not chat smart routing.
Get started
Create an API key in the dashboard.
Set
base_urltohttps://api.rodiumai.io/v1.Call
model="rodiumai/smart"onchat.completions.Inspect
response.modelto see the selected slug.Promote smart as the default for product chat; pin frontier models only where you must.
One API, one wallet, many models. Let RodiumAI choose the ideal model for each request so you gain on cost and speed without rewriting your stack.


