Requests are rate-limited at two levels: per API Key and per user (all keys of a user share the user-level budget). Actual RPM / QPS quotas depend on your plan — see the console.
Over-limit response
HTTP 429 with an OpenAI-compatible error body:
{ "error": { "message": "rate limit exceeded", "type": "rate_limit_error", "code": "rate_limited" }}Related error codes:
| HTTP | code | Meaning |
|---|---|---|
| 429 | rate_limited |
Platform key/user rate limit hit |
| 429 | key_quota_exhausted |
Key / user quota exhausted |
| 402 | user_balance_exhausted |
Insufficient account balance |
| 429 | upstream_rate_limited |
All upstream providers busy during the retry window |
| 502 | upstream_exhausted |
Upstream unavailable after retries |
Retry guidance
- On 429 / 502, retry with exponential backoff plus jitter (e.g. 1s → 2s → 4s)
- Poll async task status every 5–10 seconds — faster polling does not speed up generation
- For long synchronous requests (image generation), don't immediately resend on client timeout; confirm the previous request finished to avoid duplicate charges
