| title | Rate Limits |
|---|---|
| description | Understanding ModelsLab API rate limits, the RateLimit response headers, request queuing, and concurrency limits across subscription plans. |
Our API uses request queue limits to manage server load and ensure optimal performance. The limits vary based on your subscription plan:
**5 queued API requests**Perfect for individual developers and small projects getting started with our APIs.
Ideal for growing businesses and applications with moderate usage requirements.
Designed for enterprise applications and high-volume usage scenarios.
Every ModelsLab API response carries its own rate-limit state, so you can throttle against the real remaining budget instead of discovering a limit by exceeding it.
| Header | Meaning |
|---|---|
RateLimit-Limit |
Requests permitted in the current window. |
RateLimit-Remaining |
Requests still available. Throttle against this. |
RateLimit-Reset |
Seconds until the window resets. |
RateLimit-Policy |
Structured field: "<policy>";q=<quota>;w=<window seconds>. |
RateLimit |
Structured field: "<policy>";r=<remaining>;t=<reset seconds>. |
Retry-After |
Seconds to wait. Sent when the limit is exhausted. |
The legacy X-RateLimit-Limit and X-RateLimit-Remaining headers are still sent
alongside these and are not going away.
curl -sD - -o /dev/null https://modelslab.com/api/agents/v1/changelog | grep -i ratelimitRateLimit-Limit: 120
RateLimit-Remaining: 119
RateLimit-Reset: 60
RateLimit-Policy: "default";q=120;w=60
RateLimit: "default";r=119;t=60
| Surface | Limit |
|---|---|
| Authenticated control plane | 120 requests/minute |
| Billing and wallet mutations | 15 requests/minute |
| Auth endpoints | 20 requests/minute per IP and email |
| Signup | 3/minute, 10/hour, 20/day per client IP |
Buckets are keyed per credential, so two tokens on one account do not consume each other's budget.
Generation endpoints (`/api/v6`, `/api/v7`, `/api/v8`) signal a rate-limit refusal as **HTTP 200** with `"status": "error"` and `"code": "rate_limited"` in the body — not as a `429`. Branch on the body, not the HTTP status. See [Error Codes](/error-codes).Request queuing ensures that API calls are processed sequentially in a controlled manner. Here's what you need to know:
- Sequential Processing: Requests are processed one after another in queue order
- Queue Management: New requests are added to the queue and processed when previous ones complete
- Per Account: Limits apply to your entire account, not per API endpoint
- Real-time: The limit is enforced in real-time as requests come in
When you reach your queue limit:
- Queue Full: Additional requests are rejected with a rate limit error
- Sequential Processing: Requests are processed one after another in queue order
- FIFO Order: Requests are processed in First-In-First-Out order
- Automatic Processing: Queued requests are automatically processed as previous ones complete
When you hit rate limits, you'll receive an HTTP 429 status code with details about the limit:
{
"status": "error",
"message": "Rate limit exceeded. Maximum 5 queued requests allowed.",
"retry_after": 30
}If you need higher queue limits:
- Log in to your ModelsLab account
- Navigate to the billing section
- Select a higher tier plan
- New limits take effect immediately
Need help with rate limits or want to discuss custom solutions?
- Documentation: Check our API Reference for detailed endpoint information
- Support: Contact us at support@modelslab.com
- Discord: Join our Discord community for real-time help