Forge is available: build a website with AI from your RodiumAi account.

Try Forge

Use cases

rodiumai/smart on RodiumAI: let the gateway pick the right model for cost and speed

rodiumai/smart is RodiumAI's LLM router: one OpenAI-compatible call, and the gateway picks the catalogue model that best matches your request. Less overpaying on simple prompts, more power when the task is hard. Here is how it works and how to use it.

  • rodiumai
  • smart
  • routing
  • auto
  • cost
  • latency
  • models
  • openai-compatible
  • api
5 min read69 views

What is rodiumai/smart?

rodiumai/smart is a virtual model on the RodiumAI gateway. You do not call a fixed upstream (Claude, GPT, Gemini…). You call rodiumai/smart, and RodiumAI selects the ideal catalogue model for that request.

The goal is simple:

  • Save cost when a short or easy prompt does not need a frontier model.

  • Keep quality when the task needs coding, reasoning, vision, or tools.

  • Stay fast by preferring flash / mini / lite models when latency matters more than peak intelligence.

It is the same OpenAI-shaped API you already use (chat.completions), with one change: the model field is rodiumai/smart (aliases: smart, rodium/smart).


How it works

When you send a request with model: "rodiumai/smart", the gateway:

  1. Reads signals from your payload: message text, images, tools, approximate length.

  2. Calls a router model (default openai/gpt-4o-mini, configurable) with a short candidate list from the catalogue pool (built around the auto smart profile plus complements).

  3. Picks one catalogue slug (for example google/gemini-2.5-flash or anthropic/claude-sonnet-4-6).

  4. Forwards your request to that model and returns the normal chat completion.

The router is instructed to:

  • Prefer coding / reasoning models for code, algorithms, debugging.

  • Prefer vision-capable models when images are present.

  • Prefer cheaper / faster models (flash, mini, nano, lite) for short Q&A.

  • Prefer higher quality models for hard reasoning, long-form design, or agentic tool use.

  • Never invent a slug: the choice must be one of the candidates.

Both hops are billed in RODI: the small router hop, then the selected model hop. On non-stream responses you get routing metadata (rodiumai_routing). On stream, look at headers such as X-RodiumAI-Selected-Model.

Your app  →  rodiumai/smart  →  router picks slug  →  real model (Claude / GPT / Gemini / …)

Why use smart instead of a fixed model?

1. One key, many vendors

You keep a single rd_sk_… key and a single base_url. Smart decides whether this turn needs Haiku-class speed, Sonnet-class coding, or a Flash-class bargain.

2. Cost control without micromanagement

Pinning every call to Opus or GPT-5 burns wallet balance on greetings and tiny edits. Smart routes those to affordable models and reserves frontier capacity for harder work.

3. Speed when it matters

For short prompts, the pool includes low-latency options (Gemini Flash, GPT mini/nano, DeepSeek Flash, Claude Haiku). You get answers sooner without rewriting your client.

4. OpenAI-compatible drop-in

If your stack already uses the OpenAI SDK, Cursor, or any OpenAI-compatible agent, point base_url to RodiumAI and set model="rodiumai/smart". No new SDK.

5. Local payment rails

Same RodiumAI wallet: Mobile Money, bank transfer, RODI billing, usage in the dashboard. Smart does not change how you pay, only which model runs.


Quick start

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("RODIUMAI_API_KEY"),
    base_url="https://api.rodiumai.io/v1",
)

response = client.chat.completions.create(
    model="rodiumai/smart",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=256,
)

print(response.choices[0].message.content)
print(response.model)  # actual catalogue slug selected for this turn
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RODIUMAI_API_KEY,
  baseURL: "https://api.rodiumai.io/v1",
});

const response = await client.chat.completions.create({
  model: "rodiumai/smart",
  messages: [{ role: "user", content: "Hello!" }],
  max_tokens: 256,
});

console.log(response.choices[0].message.content);
console.log(response.model);
curl https://api.rodiumai.io/v1/chat/completions \
  -H "Authorization: Bearer rd_sk_…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "rodiumai/smart",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 256
  }'

Tip: Log response.model (or the routing headers) in production. It is the best way to see which slug smart chose and to tune prompts or fall back to a pinned model when needed.


Smart vs rule-based profiles

RodiumAI also ships rule-based smart profiles (no LLM hop). They resolve deterministically from a SmartModelProfile:

ProfileStrategyBest forbasiccheapestHigh volume, low stakesfastfastestLowest latencyprobalancedDefault quality / pricemaxhighest qualityFrontier reasoningautodynamic heuristicsPrompt-shaped routing without a router LLM

rodiumai/smart is different: it uses an LLM router over a candidate pool for finer intent matching (coding vs chat vs vision vs cheap). Profiles like fast or pro are cheaper on the routing hop itself (no classifier call) but less adaptive turn by turn.

NeedSuggested callAdaptive pick every requestrodiumai/smartAlways prioritize latencyrodiumai/fast (or fast)Always prioritize costrodiumai/basicAlways prioritize qualityrodiumai/maxKnow the exact vendor modelPin anthropic/claude-sonnet-4-6, openai/gpt-5.4-mini, etc.


Example scenarios

User requestWhat smart tends to prefer"Hello!" / short FAQCheap, fast models (Flash / mini / Haiku-class)Full Express + JWT server codeStronger coding / reasoning modelsScreenshot or chart analysisVision-capable modelsLong design / multi-step planHigher quality / reasoning members of the poolTool-heavy agent turnModels with solid tool support in the pool

You still pay for what runs. Smart reduces waste: frontier rates only when the prompt warrants them.


Billing notes

  1. Router hop + selected model hop are both charged in RODI.

  2. On simple prompts the router cost is small compared to avoiding a frontier completion.

  3. Check live RODI rates with GET /v1/models or the models catalogue.

  4. Recharge and usage stay in the RodiumAI dashboard.


When not to use smart

  • You need a guaranteed vendor slug for compliance or evals: pin the model.

  • You need zero routing hop latency and a fixed strategy: use fast, basic, pro, or max.

  • The workload is image or video generation: use POST /v1/images/generations or POST /v1/videos/generations with the right model, not chat smart routing.


Get started

  1. Create an API key in the dashboard.

  2. Set base_url to https://api.rodiumai.io/v1.

  3. Call model="rodiumai/smart" on chat.completions.

  4. Inspect response.model to see the selected slug.

  5. Promote smart as the default for product chat; pin frontier models only where you must.

One API, one wallet, many models. Let RodiumAI choose the ideal model for each request so you gain on cost and speed without rewriting your stack.


About the author

R

RodiumAI

You might also like