Key Takeaways
- Rate limiting protects your API from abuse and ensures fair usage across all clients
- Token Bucket algorithm is the most popular choice - allows burst while enforcing limits
- Always include
X-RateLimit-* headers to inform clients of their status
- Plan for 2-3x peak multiplier over average traffic
- Implement exponential backoff with jitter for retry logic
Understanding API Rate Limiting
| Limit Type |
Description |
Unit |
Example |
| RPM | Requests per minute | req/min | 60 RPM = 1 req/sec |
| RPS | Requests per second | req/sec | 10 RPS = 600 RPM |
| RPH | Requests per hour | req/hr | 3,600 RPH = 1 RPS |
| RPD | Requests per day | req/day | 100,000 RPD |
| Burst limit | Short-term spike allowed | Varies | 100 in 5 seconds |
| Concurrent | Simultaneous active requests | req | 5 concurrent |
| Quota | Monthly/periodic allowance | Total | 1M calls/month |
| Algorithm |
How It Works |
Pros |
Cons |
| Fixed window | Count resets at window start | Simple | Spike at window boundary |
| Sliding window | Rolling window from each request | Smooth | More complex |
| Token bucket | Tokens fill at fixed rate; requests spend tokens | Allows burst | Burstable past average |
| Leaky bucket | Queue drains at fixed rate | Consistent output | No burst absorption |
| Sliding window log | Track exact timestamps | Most accurate | High memory use |
| Fixed window + burst | Fixed window with short-term burst capacity | Practical | Two limits to manage |
| Strategy |
Description |
When to Use |
| Exponential backoff | Retry after 2^n seconds (1, 2, 4, 8...) | 429 or 503 responses |
| Jitter | Add random delay to backoff | Multiple clients, avoid thundering herd |
| Request queuing | Batch requests in a queue, throttle to limit | High-volume pipelines |
| Cache responses | Store and reuse recent API responses | Repeated identical queries |
| Respect retry-after | Read the Retry-After header | All 429 responses |
| Pre-emptive throttle | Self-throttle below the limit (e.g. 80%) | Production integrations |
| Circuit breaker | Stop retrying after N failures; reopen after timeout | Resilient systems |
API rate limiting is a critical strategy for managing server resources and ensuring fair usage across all clients. This guide will help you understand rate limiting concepts and how to plan effective rate limit policies for your APIs.
What is API Rate Limiting?
Rate limiting is a technique used to control the number of requests a client can make to an API within a specified time period. It protects your servers from being overwhelmed, prevents abuse, and ensures consistent service quality for all users.
Key Metrics in Rate Limiting
- Requests Per Second (RPS): The most granular measure of API traffic, essential for capacity planning.
- Requests Per Minute (RPM): Common rate limit window that balances granularity with flexibility.
- Requests Per Hour (RPH): Useful for broader usage quotas and billing purposes.
- Burst Capacity: Short-term allowance for traffic spikes above the sustained rate.
Common Rate Limiting Strategies
1. Token Bucket Algorithm
The token bucket algorithm is one of the most popular rate limiting approaches. It works by:
- Maintaining a bucket that holds tokens
- Adding tokens at a fixed rate (e.g., 10 tokens per second)
- Each request consumes one token
- Requests are rejected when the bucket is empty
- The bucket has a maximum capacity (burst limit)
This algorithm naturally allows for burst traffic while enforcing long-term rate limits.
2. Leaky Bucket Algorithm
Similar to token bucket but processes requests at a constant rate:
- Requests enter a queue (the bucket)
- Requests are processed at a fixed rate
- Excess requests overflow and are rejected
- Provides smooth, consistent output rate
3. Fixed Window Counter
The simplest approach that counts requests within fixed time windows:
- Divide time into fixed windows (e.g., per minute)
- Count requests in each window
- Reset counter at window boundary
- Simple but can allow burst at window edges
4. Sliding Window Log
A more precise method that tracks individual request timestamps:
- Store timestamp of each request
- Count requests in rolling time window
- More accurate but higher memory usage
- Eliminates boundary burst issues
5. Sliding Window Counter
A hybrid approach combining fixed windows with weighted averages:
- Combines current and previous window counts
- Weights based on position in current window
- Good balance of accuracy and efficiency
- Popular in production systems
Rate Limit Tier Best Practices
Free Tier
Designed for evaluation and small-scale usage:
- Lower limits (100-1,000 requests/day)
- Stricter burst limits
- May have feature restrictions
- Good for testing and development
Basic Tier
For production applications with moderate traffic:
- Moderate limits (10,000-100,000 requests/day)
- Reasonable burst capacity
- SLA guarantees
- Priority support
Pro Tier
For high-traffic applications:
- Higher limits (1M+ requests/day)
- Generous burst allowance
- Advanced features
- Dedicated support
Enterprise Tier
Custom solutions for large-scale deployments:
- Custom rate limits
- Dedicated infrastructure options
- Custom SLAs
- Dedicated account management
Implementing Rate Limits
HTTP Headers
Standard headers for communicating rate limit status:
X-RateLimit-Limit: Maximum requests allowed
X-RateLimit-Remaining: Requests remaining in window
X-RateLimit-Reset: Time when limit resets (Unix timestamp)
Retry-After: Seconds to wait before retrying (on 429 response)
Response Codes
200 OK: Request successful
429 Too Many Requests: Rate limit exceeded
503 Service Unavailable: Server overloaded
Tips for API Consumers
1. Implement Exponential Backoff
When rate limited, wait progressively longer between retries:
- First retry: 1 second
- Second retry: 2 seconds
- Third retry: 4 seconds
- Add random jitter to prevent thundering herd
2. Cache Responses
Reduce API calls by caching responses when appropriate:
- Respect Cache-Control headers
- Implement local caching
- Use ETags for conditional requests
3. Batch Requests
Combine multiple operations into single requests when possible:
- Use bulk endpoints
- Aggregate data fetching
- Reduce round trips
4. Monitor Usage
Track your API usage to avoid unexpected rate limiting:
- Log rate limit headers
- Set up usage alerts
- Plan for capacity increases
Conclusion
Effective API rate limiting is essential for building scalable and reliable services. By understanding the various rate limiting strategies and planning appropriate tiers for your user base, you can ensure fair resource allocation while protecting your infrastructure from abuse. Use this calculator to plan your rate limits based on expected traffic patterns and scale appropriately as your user base grows.
Common API Rate Limit Examples
| API Provider |
Free Tier |
Paid Tier |
Strategy |
| Twitter API |
500 req/15min |
Custom |
Fixed Window |
| GitHub API |
60 req/hour |
5,000 req/hour |
Fixed Window |
| Stripe API |
100 req/sec |
Custom |
Token Bucket |
| OpenAI API |
20 req/min |
3,500 req/min |
Token Bucket |
Frequently Asked Questions
How accurate are the results?
The API Rate Limit Planner applies a standard formula to your inputs — accuracy depends on how precisely you measure those inputs. For planning and estimation, results are reliable. For high-stakes or professional decisions, cross-check the output with a domain expert or primary source.
Can I use this on mobile?
Yes — the calculator is designed to work on any device. For complex multi-input calculations on small screens, landscape orientation gives more room to see all fields and results simultaneously.
What is an API rate limit?
An API rate limit is a constraint placed by an API provider on how many requests a client can make within a specified time period. Rate limits protect the provider's infrastructure from abuse, ensure fair usage among multiple clients, and manage server costs. Types of rate limits: Requests per second (RPS): the most common burst-level limit. 10 RPS = 10 requests in any given second. Requests per minute (RPM): common for many REST APIs. 60 RPM = 1 request per second average. Requests per day (RPD) or per month: quota-style limits that cap total usage. Concurrent requests: limit on how many requests can be in-flight simultaneously. Token or credit limits: some APIs (like LLM APIs) rate-limit by token count rather than request count. How rate limits are communicated: Most modern APIs include rate limit information in HTTP response headers: X-RateLimit-Limit: total limit for the window. X-RateLimit-Remaining: requests left in current window. X-RateLimit-Reset: Unix timestamp when the window resets. Retry-After: seconds to wait after a 429 Too Many Requests error. When you exceed a rate limit: The server returns HTTP 429 Too Many Requests. Well-designed clients handle 429 with exponential backoff and retry. Ignoring 429 and continuing to send requests can result in temporary IP bans or account suspension. Why rate limits matter for developers: Rate limit design affects architecture decisions — do you need request queuing, caching, or a distributed rate limiter if your app scales?
How do I calculate how many API requests I can make?
Converting between rate limit formats is a common calculation. The core conversions: If limit = X requests per minute (RPM): Per second: X / 60. Per hour: X × 60. Per day: X × 60 × 24. If limit = X requests per second (RPS): Per minute: X × 60. Per hour: X × 3,600. If limit = X requests per day (RPD): Per second: X / 86,400. Per minute: X / 1,440. Example calculations: 100 RPM: 100/60 = 1.67 req/sec average. 1,000 requests/day: 1,000/1,440 = 0.69 req/min average = one request every ~87 seconds. 5 RPS: 5 × 60 = 300 RPM; 5 × 3,600 = 18,000 RPH. Estimating whether your use case fits within limits: Monthly API calls = daily calls × 30. If you process 500 records/hour and each needs 3 API calls: 500 × 3 = 1,500 calls/hour = 25 calls/minute = 0.42 calls/second. If the API limit is 60 RPM, you have ~2.4× headroom. Pipeline throughput: Max throughput (records/hour) = (Rate limit per minute × 60) / calls per record. With 60 RPM and 3 calls/record: 60 × 60 / 3 = 1,200 records/hour maximum. When to add safety margin: Never plan at exactly the limit — bursts in traffic, retries, and monitoring calls consume headroom. Target 70–80% of the stated limit in production design.
How do I handle API rate limit errors (HTTP 429)?
A 429 Too Many Requests response means you've exceeded the API's rate limit. Handling it well prevents data loss and account suspension. Step 1 — Read the response headers: Retry-After: if present, wait exactly this many seconds before retrying. X-RateLimit-Reset: Unix timestamp of when the limit resets. Respecting these headers is the most important first step. Step 2 — Implement exponential backoff: wait = initial_delay × 2^attempt + random_jitter. Attempt 1: wait ~1 second. Attempt 2: wait ~2 seconds. Attempt 3: wait ~4 seconds. Attempt 4: wait ~8 seconds. Etc. (cap at a maximum, e.g. 60 seconds). Add jitter (random 0–1 second): prevents multiple clients from retrying simultaneously ("thundering herd"). Step 3 — Queue your requests: rather than firing requests as fast as possible and hitting 429s, maintain a request queue. Dequeue at the target rate (slightly below the limit). This is more efficient than retry-based approaches. Example in Python (conceptual): time.sleep(1.0 / rate_limit_per_second) between requests. Use asyncio with semaphores or ratelimit library for async code. Step 4 — Implement a circuit breaker: after N consecutive 429s, stop sending requests for a defined cooldown period. Open the circuit again after the cooldown. Prevents battering the API while your backoff catches up. Step 5 — Monitor and alert: track 429 rate in your metrics. Consistently high 429 rates signal you need: a higher tier plan, more efficient request batching, or better caching. Best practices summary: Always respect Retry-After. Use exponential backoff with jitter. Pre-throttle to 70–80% of limit. Cache responses where possible. Never ignore 429s — repeated violations can lead to account suspension.
What is exponential backoff and when should I use it?
Exponential backoff is a retry strategy where each successive retry waits exponentially longer than the previous one, giving the server time to recover. The formula: wait_time = base_delay × 2^attempt + random_jitter. Example progression (base = 1 second): Attempt 1 fails → wait ~1 second. Attempt 2 fails → wait ~2 seconds. Attempt 3 fails → wait ~4 seconds. Attempt 4 fails → wait ~8 seconds. Attempt 5 fails → wait ~16 seconds. Maximum cap: typically 32–64 seconds (prevents infinite growth). Jitter: adding random(0, 1) seconds to each wait time prevents multiple clients from retrying in perfect synchrony — this avoids the "thundering herd" problem where all clients hammer the API simultaneously after a shared rate limit window resets. When to use exponential backoff: HTTP 429 (Too Many Requests). HTTP 503 (Service Unavailable). Timeout errors. Network connectivity errors. When NOT to use it: HTTP 400 (Bad Request) — the request itself is malformed; retrying without changing the request will always fail. HTTP 401/403 (Unauthorized) — authentication/authorization issues; retrying won't help. HTTP 404 (Not Found) — resource doesn't exist; no point retrying. Implementation: Max retries: set a limit (3–5 attempts) — infinite retries can mask bugs and waste resources. Log each retry: for debugging and alerting. Circuit breaker integration: after max retries exhausted, open the circuit for a longer cooldown. Alternatives for rate limiting: For rate limits specifically, a pre-emptive request queue (throttling to ~80% of limit) is often more efficient than retry-based backoff because you avoid hitting the limit at all. Use backoff as a fallback for unexpected rate limits, not as the primary strategy.