Q: What are the default concurrency limits? How can I increase them?
A:
Default quotas: QPS=[X], QPM=[X], daily Tokens=[X]
Ways to increase them:
Console → “Help & Tickets” → “Submit Quota Increase Request” → Enter your requirements → Review (usually within 24 hours)
Enterprise customers can contact their account manager to customize quotas
For urgent requirements, contact technical support to request a temporary capacity increase
Q: How should I design a high-availability architecture?
A: Recommendations:
Primary/standby failover: Configure primary and standby API endpoints and automatically switch if the primary endpoint becomes unavailable
Exponential backoff: When encountering a 429/500 error, retry at intervals of 1s→2s→4s→8s, up to 3 times
Request queue: In high-concurrency scenarios, use a message queue (such as RabbitMQ/Kafka) to handle traffic spikes
Multi-Key rotation: Create multiple API Keys and distribute requests through load balancing
Timeout control: Set a reasonable timeout (60~120s recommended) to prevent threads from being blocked
Monitoring and alerts: Monitor success rates, latency, and Token consumption, and issue timely alerts when anomalies occur
Q: Is batch processing (Batch API) supported?
A: Yes. It is suitable for large-scale offline, non-real-time tasks:
Submit task files in JSONL format
The platform processes tasks asynchronously, usually completing them within 24 hours
Pricing is approximately 50% of real-time API calls
Console → “Batch Tasks” → Upload File → View progress and results
Suitable for data labeling, bulk content generation, testing, evaluation, and similar scenarios.
