Rate Limits

Two-level limiting per API Key and per user; 429 on excess — retry with exponential backoff

Requests are rate-limited at two levels: per API Key and per user (all keys of a user share the user-level budget). Actual RPM / QPS quotas depend on your plan — see the console.

Over-limit response

HTTP 429 with an OpenAI-compatible error body:

{  "error": {    "message": "rate limit exceeded",    "type": "rate_limit_error",    "code": "rate_limited"  }}

Related error codes:

HTTP code Meaning
429 rate_limited Platform key/user rate limit hit
429 key_quota_exhausted Key / user quota exhausted
402 user_balance_exhausted Insufficient account balance
429 upstream_rate_limited All upstream providers busy during the retry window
502 upstream_exhausted Upstream unavailable after retries

Retry guidance

  • On 429 / 502, retry with exponential backoff plus jitter (e.g. 1s → 2s → 4s)
  • Poll async task status every 5–10 seconds — faster polling does not speed up generation
  • For long synchronous requests (image generation), don't immediately resend on client timeout; confirm the previous request finished to avoid duplicate charges