Providers & pricing
| Provider | Input | Output | Free tier |
|---|---|---|---|
Vertex Default | 376.8 RODI/M~0.500 USD/M | — RODI/M | — |
Google AI Studio | 376.8 RODI/M~0.500 USD/M | — RODI/M | — |
Pricing
| Rate | RODI | USD (ref.) | Unit |
|---|---|---|---|
| In | 376.8 | ~ 0.500 | USD/M · RODI/M |
| Audio out | — | ~ 10.00 | USD/s · RODI/s |
RODI prices include RodiumAi markup and upstream fees. USD figures are wholesale reference rates.
Capabilities
Streaming
Tool calling
Vision
JSON mode
Reasoning
About this model
Gemini 2.5 Flash TTS is Google's fast text-to-speech model for low-latency, high-fidelity narration. It converts text into spoken audio via generateContent (response modalities AUDIO) or POST /v1/audio/speech, with controllable prebuilt voices.
API usage
Synthesize speech with client.speech() — binary POST /v1/audio/speech.
RodiumAi SDK
speech() returns raw audio bytes (save to disk or stream to clients).
…Compatible SDKs (OpenAI / Anthropic packages)
…