Rate Limiting

Q: What are the default concurrency limits? How can I increase them?

A:

  • Default quotas: QPS=[X], QPM=[X], daily Tokens=[X]

  • Ways to increase them:

    1. Console → “Help & Tickets” → “Submit Quota Increase Request” → Enter your requirements → Review (usually within 24 hours)

    2. Enterprise customers can contact their account manager to customize quotas

    3. For urgent requirements, contact technical support to request a temporary capacity increase

Q: How should I design a high-availability architecture?

A: Recommendations:

  1. Primary/standby failover: Configure primary and standby API endpoints and automatically switch if the primary endpoint becomes unavailable

  2. Exponential backoff: When encountering a 429/500 error, retry at intervals of 1s→2s→4s→8s, up to 3 times

  3. Request queue: In high-concurrency scenarios, use a message queue (such as RabbitMQ/Kafka) to handle traffic spikes

  4. Multi-Key rotation: Create multiple API Keys and distribute requests through load balancing

  5. Timeout control: Set a reasonable timeout (60~120s recommended) to prevent threads from being blocked

  6. Monitoring and alerts: Monitor success rates, latency, and Token consumption, and issue timely alerts when anomalies occur

Q: Is batch processing (Batch API) supported?

A: Yes. It is suitable for large-scale offline, non-real-time tasks:

  • Submit task files in JSONL format

  • The platform processes tasks asynchronously, usually completing them within 24 hours

  • Pricing is approximately 50% of real-time API calls

  • Console → “Batch Tasks” → Upload File → View progress and results

Suitable for data labeling, bulk content generation, testing, evaluation, and similar scenarios.