Forge is available: build a website with AI from your RodiumAi account.

Try Forge
RodiumAi docs
Core concepts

Audio transcription

Upload an audio file (multipart) and get text back. Same base_url and rd_sk_* key as chat.

POSThttps://api.rodiumai.io/v1/audio/transcriptions

RodiumAi exposes OpenAI-compatible speech-to-text. Send a multipart form with file and model, use the official OpenAI package or plain HTTP.

When to use

  • Meeting notes and call summaries from a recording.
  • Force language + glossary via language and prompt.
  • Need timestamps? Prefer verbose_json when supported.

Recipes

Meeting notes

Upload the full WAV/MP3, model openai/gpt-4o-transcribe or a Gemini id, then summarize the text with chat.

Note: Large files increase latency; chunk long meetings if needed.

Language + prompt

Set language=fr and a prompt listing product names to stabilize spelling.

Note: Prompt is a soft hint, not a hard constraint.

verbose_json

response_format=verbose_json returns richer metadata when the upstream supports it.

Note: Billing still uses token estimates from the transcription path.

Examples

…

Request parameters

ParameterTypeRequiredDescription
filefileRequiredAudio file to transcribe (multipart form field). The whole request must stay under 10 MiB.
modelstringRequiredTranscription model id (e.g. openai/gpt-4o-transcribe, openai/gpt-4o-mini-transcribe, google/gemini-2.5-flash).
languagestringOptionalOptional ISO-639-1 language hint (e.g. en, fr).
promptstringOptionalOptional text to guide style or spelling of the transcript.
response_formatstringOptionalOutput format when supported (e.g. "json", "text", "verbose_json").
chunking_strategystringOptionalOpenAI transcription models: chunking for long audio (e.g. "auto"). Forwarded as-is.
temperaturenumberOptionalSampling temperature between 0 and 1 when supported.

API reference: transcriptions

Limits and behaviour

  • Send multipart/form-data with file and model. Missing fields return 422 with a detail array instead of the usual error envelope.
  • The whole request, file included, must stay under 10 MiB (413 payload_too_large above). Split or compress longer recordings.
  • Both Authorization: Bearer and x-api-key are accepted.
  • language, prompt, response_format, temperature and chunking_strategy are forwarded to the model; the response is the provider's JSON (text, or richer fields with verbose_json).
  • Billed from the audio and text tokens the provider reports, at the model's price.