Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Core concepts

Text to speech

Send JSON with model and input; receive audio bytes. Billed in RODI (incl. output_audio when priced).

POSThttps://api.rodiumai.io/v1/audio/speech

Call POST /v1/audio/speech with a TTS model id. OpenAI voices (alloy, …) or Gemini prebuilt voices (Kore, Puck, …).

When to use

  • IVR prompts and product narration in French or English.
  • Choose Gemini expressive voices or OpenAI gpt-4o-mini-tts.
  • Save the binary response directly to a file.

Recipes

French narration

input in French, voice Kore (Gemini) or alloy (OpenAI), response_format wav or mp3.

Note: Write the response body as binary, do not JSON.parse it.

Gemini vs OpenAI

google/gemini-2.5-flash-tts or google/gemini-2.5-pro-tts for Gemini voices; openai/gpt-4o-mini-tts for OpenAI voices.

Note: Confirm the slug in GET /v1/models before shipping.

WAV note (Gemini)

Gemini upstream is PCM; Rodium wraps WAV for non-pcm formats so browsers can play the file.

Note: Use response_format=pcm only if you decode L16 yourself.

Examples

Gemini TTS

…

OpenAI TTS

…

Request parameters

ParameterTypeRequiredDescription
modelstringRequiredTTS model id (e.g. openai/gpt-4o-mini-tts, google/gemini-2.5-flash-tts).
inputstringRequiredText to synthesize into speech.
voicestringOptionalVoice id when supported (e.g. "alloy", "nova").
response_formatstringOptionalAudio format when supported (e.g. "mp3", "opus", "wav").
speednumberOptionalPlayback speed multiplier when supported (typically 0.25–4.0).
instructionsstringOptionalOptional speaking style / delivery instructions when the model supports them.

API reference: speech