Core concepts
Audio transcription
Upload an audio file (multipart) and get text back. Same base_url and rd_sk_* key as chat.
POST
https://api.rodiumai.io/v1/audio/transcriptionsRodiumAi exposes OpenAI-compatible speech-to-text. Send a multipart form with file and model, use the official OpenAI package or plain HTTP.
Model ids
Use provider-scoped ids such as openai/gpt-4o-transcribe, openai/gpt-4o-mini-transcribe or google/gemini-2.5-flash. List them with GET /v1/models.
When to use
- Meeting notes and call summaries from a recording.
- Force language + glossary via language and prompt.
- Need timestamps? Prefer verbose_json when supported.
Recipes
Meeting notes
Upload the full WAV/MP3, model openai/gpt-4o-transcribe or a Gemini id, then summarize the text with chat.
Note: Large files increase latency; chunk long meetings if needed.
Language + prompt
Set language=fr and a prompt listing product names to stabilize spelling.
Note: Prompt is a soft hint, not a hard constraint.
verbose_json
response_format=verbose_json returns richer metadata when the upstream supports it.
Note: Billing still uses token estimates from the transcription path.
Examples
…Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| file | file | Required | Audio file to transcribe (multipart form field). The whole request must stay under 10 MiB. |
| model | string | Required | Transcription model id (e.g. openai/gpt-4o-transcribe, openai/gpt-4o-mini-transcribe, google/gemini-2.5-flash). |
| language | string | Optional | Optional ISO-639-1 language hint (e.g. en, fr). |
| prompt | string | Optional | Optional text to guide style or spelling of the transcript. |
| response_format | string | Optional | Output format when supported (e.g. "json", "text", "verbose_json"). |
| chunking_strategy | string | Optional | OpenAI transcription models: chunking for long audio (e.g. "auto"). Forwarded as-is. |
| temperature | number | Optional | Sampling temperature between 0 and 1 when supported. |
Limits and behaviour
- Send
multipart/form-datawithfileandmodel. Missing fields return422with adetailarray instead of the usualerrorenvelope. - The whole request, file included, must stay under 10 MiB (
413 payload_too_largeabove). Split or compress longer recordings. - Both
Authorization: Bearerandx-api-keyare accepted. language,prompt,response_format,temperatureandchunking_strategyare forwarded to the model; the response is the provider's JSON (text, or richer fields withverbose_json).- Billed from the audio and text tokens the provider reports, at the model's price.