Errors & retries
One table to decide whether to fix the request, top up, wait or retry, plus a retry loop you can copy.
Decision table
| Status | Meaning | Retry? | What to do |
|---|---|---|---|
| 400 · 404 · 413 · 422 | The request itself is wrong | No | Fix the body, the model id or the payload size. |
| 401 | Credential rejected | No | Check the key, or refresh the OIDC token. |
| 402 | Balance too low for the pre-flight hold | No | Top up, lower max_tokens or pick a cheaper model. |
| 403 | Key restriction (scope, allowed models, monthly quota) | No | Change the key's settings or use another key. |
| 429 | Rate limited (yours or the provider's) | Yes, after Retry-After | Back off and spread traffic over time. |
| 500 | Unexpected gateway error | Once or twice | Keep the request_id for support. |
| 502 · 503 · 504 | Provider or platform temporarily unavailable | Yes, exponential backoff | Consider a fallback model. |
Every status, type and code is listed in the API errors reference.
Retry with backoff
Retry only 429 and 5xx. Honour Retry-After (seconds) when present; otherwise back off exponentially with jitter and cap the number of attempts. The official OpenAI SDKs already retry 429 and 5xx twice by default: set max_retries=0 / maxRetries: 0 when you run your own loop so attempts do not multiply.
…What gets billed
- A request that fails before generation starts (validation, authentication, balance, limits, provider refusal) costs nothing: its pre-flight hold is released.
- A cancelled or interrupted stream is billed on the input plus the output actually delivered.
- Details: Pricing and billing.
Errors inside a stream
A stream that fails after it started still has HTTP status 200. On chat completions the gateway sends a data: event with an error object (the same sanitised error as outside a stream), then data: [DONE]. Treat that event as a failure of the whole response and decide from its code whether to retry. Responses and Messages formats: Errors during a stream.
…Provider failures
RodiumAI never forwards raw provider errors. Account, billing or permission problems on the provider side (provider 401, 402, 403) become 502 provider_unavailable; provider throttling becomes 429 rate_limit_exceeded with Retry-After; timeouts become 504 timeout and other outages 502 provider_unavailable; provider 400, 404, 413 and 422 keep their status with a stable code and a sanitised message. None of the 5xx cases means your key or your wallet is wrong.
Contacting support
- The
X-Request-Idof the failing response (orerror.request_idon a500). - The UTC time, endpoint, model id, HTTP status and
error.code. - Never send your API key or the full content of sensitive prompts.