Rate Limits
Rate Limiting Guide
Rate limits protect the platform and ensure fair usage across all tenants. Limits are applied per API key, per minute, and vary by tier and face.
Limits by Tier
| Tier | Global | ee.ai | ee.storage | ee.billing |
|---|---|---|---|---|
| Starter | 100/min | 60/min | 200/min | 50/min |
| Business | 1,000/min | 500/min | 2,000/min | 500/min |
| Enterprise | Custom | Custom | Custom | Custom |
Some faces have lower per-face limits than the global limit. The AI face is the most constrained due to compute costs. Enterprise customers can negotiate custom limits.
Rate Limit Headers
Every API response includes rate limit headers so you can monitor your usage in real time.
| Header | Description |
|---|---|
| X-RateLimit-Limit | The maximum number of requests allowed in the current window. |
| X-RateLimit-Remaining | The number of requests remaining in the current window. |
| X-RateLimit-Reset | Unix timestamp (seconds) when the rate limit window resets. |
| Retry-After | Seconds to wait before making another request (only on 429 responses). |
429 Response Handling
When you exceed your rate limit, the API returns a 429 status code with a Retry-After header indicating how long to wait.
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1697400000
Retry-After: 42
{
"error": {
"code": "RATE_LIMITED",
"message": "Rate limit exceeded. Retry after 42 seconds.",
"requestId": "req_abc123"
}
}Best Practices
- Implement exponential backoff — the SDK has built-in retry support
- Use batch operations — combine multiple calls into one request to reduce rate limit consumption
- Cache responses— avoid redundant calls for data that rarely changes
- Monitor X-RateLimit-Remaining — throttle proactively before hitting the limit
- Use scoped tokens — different tokens have independent rate limits
SDK Automatic Retry
The SDK can automatically handle rate limiting with exponential backoff. Enable it in the client configuration.
import { createClient } from '@evileye/sdk'
const ee = createClient({
token: process.env.EE_TOKEN!,
entity: 'my-company',
// Built-in retry with exponential backoff
retry: {
maxRetries: 3,
backoff: 'exponential', // 1s, 2s, 4s
retryOn: [429, 500, 503],
},
})
// The SDK automatically handles 429 responses
const result = await ee.ai.generate({ prompt: 'Hello' })Related
- Error Handling— all error codes including 429
- Batch Operations— reduce rate limit consumption with batch calls
- Pricing— compare tier limits and upgrade options