Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Guides

Pricing and billing

Every request is priced in RODI from the model's published rates, held before the call and settled on actual usage.

RODI, the billing unit

  • 1 RODI = 1 XOF (FCFA) at top-up. Published model prices already include a 3% FX buffer and the platform margin; there is no extra per-request fee.
  • Your wallet is prepaid: top up with Mobile Money in Dashboard → Billing. Requests that do not fit in the spendable balance fail with 402 insufficient_balance before anything is generated.
  • Some models include a daily free allowance, shown on their model page. Usage inside it is not charged.

What is metered

Token prices are quoted per 1 million tokens, separately for each kind of token. A call is the sum of the buckets it uses.

BucketWhat it countsPrice
InputPrompt tokens: messages, system prompt, tool definitions, images sent to a chat modelinput_per_1m
OutputGenerated tokens, including reasoning tokens of reasoning modelsoutput_per_1m
Cached inputPrompt tokens the provider served from its cachecached_input_per_1m (input price when absent)
Cache writeAnthropic prompt-caching writes5-minute writes 1.25 × input, 1-hour writes 2 × input, unless the model sets its own price
AudioAudio tokens in (transcription, audio models) and out (text-to-speech)Per 1M audio tokens, on the model page
Image tokensGPT Image models: text, input-image and output-image tokensPer 1M tokens, on the model page
Per imageFlat-priced image models, by quality and sizeper_image / model page
VideoGenerated video lengthPer second, on the model page
Web searchSearches run by the provider's hosted web search tool: server_tool_use.web_search_requests on /v1/messages, web_search_call items on /v1/responsesPer call, on top of tokens: the provider's $10 per 1,000 searches, converted to RODI like token prices

Requests always run at the standard price tier: service_tier values priority, fast, ultrafast and scale are not forwarded (auto, default and flex are), and on Claude models a speed other than standard is not forwarded. You are never billed a premium tier.

GET /v1/models and GET /v1/pricing expose the input, output, cached-input and per-image prices. Audio, image-token and per-second video prices are shown on each model's page in the catalogue; for video models the token fields in /v1/models read 0.0.

Rounding

Each bucket a call uses is converted to RODI and rounded up to the next 0.1 RODI, then the buckets are added. A short chat call that uses input and output therefore costs at least 0.2 RODI.

Worked example

At the time of writing, openai/gpt-4o is listed at 1882.9 RODI per 1M input tokens and 7531.3 RODI per 1M output tokens. A call with 1,200 input tokens and 400 output tokens:

BucketTokensComputationCharged
Input1,2001,200 / 1,000,000 × 1882.9 = 2.262.3 RODI
Output400400 / 1,000,000 × 7531.3 = 3.013.1 RODI
Total5.4 RODI

The gateway computes each bucket from the underlying rate, so a figure you derive from the published per-1M price can differ by 0.1 RODI per bucket. Prices change: read them from GET /v1/models rather than hard-coding them.

ModelInput / 1MOutput / 1MCached input / 1M
openai/gpt-4o1882.97531.3941.5
openai/gpt-4o-mini113.0451.956.5
anthropic/claude-sonnet-4-62259.411296.9226.0
google/gemini-2.5-flash226.01882.9226.0
deepseek/deepseek-v3.2467.01393.3—
openai/text-embedding-3-small15.1——

Images, video and audio

  • Images: flat-priced models charge per image, by quality and size; GPT Image models (openai/gpt-image-*) charge per token for the prompt, input images and generated image. The hold uses a conservative estimate per image (up to 4,160 output image tokens for high quality).
  • Video: priced per second. The hold uses the requested duration_seconds (8 by default); you are billed for the duration the provider reports, or the requested duration when it reports none.
  • Text to speech: input text tokens plus output audio tokens. Transcription: audio input tokens plus text output tokens, as reported by the provider.

Hold, then settle

  1. Estimate. Before calling the provider, the gateway estimates the cost: input tokens from your messages and tools (about 3 characters per token), plus max_tokens (or max_completion_tokens) output tokens per choice. Without either, the hold assumes 2,048 output tokens.
  2. Hold. That estimate is reserved from your spendable balance. If it does not fit, the request fails with 402 and nothing is called. balance_rodi from GET /v1/wallet excludes in-flight holds.
  3. Settle. When the call completes, the actual usage reported by the provider is priced and you pay the actual amount, not the estimate: the bill is never capped by the hold. If it is lower, the rest of the hold returns to your balance. If it is higher, the excess is taken from your available balance, never from other requests' holds. The hold is only a pre-flight estimate.
  4. Failover. If the gateway has to switch to another provider for the same model and that provider is more expensive, you never pay more than the primary provider's price for those tokens.

Failed and interrupted requests

  • A request that fails before generation starts (validation, authentication, balance, rate limit, provider refusal or outage) is not billed: its hold is released.
  • A cancelled or interrupted stream (you disconnect, or the provider stops mid-way) is billed on the input plus the output actually delivered to you up to that point.
  • Errors are listed in the error reference.

Keep cost bounded

The output limit is the main lever. Set max_tokens (or max_completion_tokens) on chat, max_tokens on /v1/messages and max_output_tokens on /v1/responses. When you omit it, providers other than OpenAI receive a default of 4,096 output tokens, and a smaller hold (2,048) is reserved up front.

…
  • Give each API key a monthly RODI quota (403 api_key_quota_exceeded once reached) and an allowed_models list.
  • Prefer smaller models for classification and extraction; pin a concrete model instead of rodiumai/smart when you need predictable cost.

See what a call cost

API responses report token counts in usage (prompt_tokens, completion_tokens, cached and reasoning details when the provider sends them). They never include a RODI amount or your balance. To know the cost:

  • Dashboard → Usage lists every request with its model, tokens and RODI cost.
  • `GET /v1/wallet` before and after a call: the difference is the cost, provided nothing else spends from the wallet meanwhile. Settlement is asynchronous, so wait a moment before the second read.
  • Compute it from usage and the prices in GET /v1/models or GET /v1/pricing, as in the example above.
…