Image generation
From zero: install openai, set base_url, call images.generate. Pass optional reference images to edit. Billed per image or per token in RODI.
https://api.rodiumai.io/v1/images/generationsRodiumAi exposes OpenAI-compatible image generation. Text-to-image works with OpenAI GPT Image and Google Gemini Image models. Pass image or images on the same endpoint to edit: GPT Image requests are routed to the provider's edit API, Gemini Image models edit natively.
Model ids & input images
When to use
- Product heroes, ads, and social creatives from a text brief.
- Edit or restyle a still with GPT Image or Gemini Image (image / images).
- Batch variants (n, size, quality) or feed a still into Veo.
Recipes
Product hero
Describe lighting, angle, and brand mood in the prompt. Start with 1024x1024 medium quality.
Note: Billed per image in RODI, check GET /v1/pricing before large n.
Image edit
Use a GPT Image or Gemini Image model and pass image (or images[]) as { b64_json } or a data URL. HTTP(S) URLs are rejected.
Note: gs:// references are accepted by Gemini Image models only; send b64_json or a data URL to GPT Image.
Image → video
Generate or edit a still here, then POST /v1/videos/generations with image.b64_json (not an HTTPS URL).
Note: See the video guide image-to-video recipe.
Text → image examples
…Image → image (edit)
Send prompt plus image or images (max 14). Same encodings as video: { b64_json }, data URL, or gs://. HTTP(S) image URLs are rejected.
…Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | Image model id (e.g. openai/gpt-image-1.5, google/gemini-3.1-flash-image). List them with GET /v1/models. |
| prompt | string | Required | Text description of the image to generate or edit instruction. |
| image | object | string | Optional | Optional reference image to edit. Accepts { b64_json }, a data URL or raw base64; gs:// URIs on Gemini Image models only. HTTP(S) URLs are rejected. On GPT Image models the request is sent to the provider's edit API. Aliases: input_image, image_url. |
| images | array | Optional | Several reference images (same encodings as image), max 14. Alias: input_images. Use for multi-image edits and merges. |
| n | integer | Optional | Number of images (1–10). Default 1. |
| size | string | Optional | Output size, e.g. "1024x1024", "1536x1024", "1024x1536" (mapped to an aspect ratio on Gemini Image models). |
| quality | string | Optional | Quality when supported (GPT Image: low, medium, high). Also selects the price tier of flat-priced models. |
| mask | object | string | Optional | GPT Image edits only: PNG mask (same encodings as image) marking the area to change. |
| background / output_format | string | Optional | Forwarded to GPT Image models (e.g. background: transparent, output_format: webp). |
Models and input images
| Family | Text to image | Input images | Encodings |
|---|---|---|---|
GPT Image (openai/gpt-image-1.5, openai/gpt-image-1-mini, …) | Yes | Yes: the gateway calls the provider's edit API; mask supported | { b64_json }, data URL, raw base64 |
Gemini Image (google/gemini-3.1-flash-image, google/gemini-3-pro-image, …) | Yes | Yes, up to 14 (image or images) | { b64_json }, data URL, raw base64, gs:// |
Remote http(s):// image URLs are never fetched by the gateway: encode the file as base64. Keep the whole request under 10 MiB, or it is rejected with 413.
Response
- Images come back as base64 in
data[].b64_json(OpenAI routes also addmime_type). GPT Image models always answer in base64, soresponse_formatis not needed. - Decode
b64_jsonand write the bytes to a file, as in the examples above. Some models may returndata[].urlinstead; handle both. - Billing: per image (flat-priced models, by
qualityandsize) or per token (GPT Image), see Pricing.