Error format#
Errors follow the OpenAI format — an HTTP status plus an error object:
{
"error": {
"message": "Invalid API key",
"type": "invalid_request_error",
"code": "invalid_api_key"
}
}Status codes#
| Code | Meaning | What to do |
|---|---|---|
400 | Bad request | Check the request body and parameters. |
401 | Missing or invalid key | Check the Authorization header and the key itself. |
402 | Insufficient balance | Top up your balance in the dashboard. |
403 | Key restricted | The key hit its spending limit or was revoked. |
404 | Model not found | Check the model id against the catalog. |
429 | Too many requests | Retry later with exponential backoff. |
500 | Internal error | Retry; if it persists, contact support. |
502, 503 | Provider unavailable | Retry later or pick another model. |
504 | Timed out | Lower max_tokens or switch to streaming. |
Retries#
Retry only 429 and 5xx errors, with a growing pause: 1, 2, 4, 8 seconds. If the response has a Retry-After header, wait as long as it says. Other 4xx errors will not succeed on retry — the request needs fixing.
The official OpenAI and Anthropic SDKs already retry these for you (twice by default); tune it with max_retries.
Timeouts#
Large models and long answers take time — sometimes minutes. Give your client a timeout of at least 120 seconds, or use streaming: the first tokens arrive sooner and the connection never sits idle.
Limits#
Rate limits depend on the model and current load. When you exceed them the API answers 429 — pause and retry. The spending limit on each key is yours to set in the dashboard.