How to Retry API Requests Safely
Your code calls an API and the call times out. The server may never have received the request, or it may have…
Your requests to an API start failing, and the error says "rate limit exceeded" or "quota exhausted". Both messages mean the provider is restricting your usage, but they describe different restrictions.
An API rate limit controls how fast you can send requests. An API quota controls how much you can use in total, over a longer period or within an assigned allowance. You can hit a rate limit while most of your quota is still left, and you can use up your quota without ever going over the rate limit.
A rate limit counts requests over a short interval, such as 100 requests per second, 1,000 calls per minute or 20 requests every 90 seconds. Some APIs count tokens, bandwidth or another unit instead of requests. Azure API Management's rate-limit policy, for example, allows a set number of calls per renewal period and rejects the extra calls from a client that goes faster than that.
Rate limits protect a service from sudden bursts that would overload its databases, queues or other downstream systems. They also share access fairly between users. The counter can apply per API key, user, account, IP address or route, so "100 requests per minute" means little until you know what it counts.
When a client goes over a rate limit, the API usually returns HTTP 429 Too Many Requests. RFC 6585 defines this status for a client that has sent too many requests in a given amount of time. The response can include a Retry-After header that says when to try again.
A quota sets the total amount you can use over a longer period, or the total allowance of an account, project, subscription or resource. Typical quotas look like this:
Quotas govern consumption, cost, capacity and what a plan includes. Google Cloud describes quotas as limits on how much of a resource a project can use, and that resource can be API requests, hardware, network components or project capacity.
Suppose an API plan allows 1,000,000 requests per month and 100 requests per minute. A client that sends 150 requests in one minute gets rate-limited, although it has used only a small part of its monthly quota. A client that sends a steady 50 requests per minute never hits the rate limit, but it uses up the whole monthly quota in under two weeks.
Once a quota is used up, the service rejects requests until the quota resets. With some providers you need an approved quota increase or a different plan instead. Quota errors do not share one status code. Some providers return 429, others return their own error.
| Feature | API rate limit | API quota |
|---|---|---|
| Main question | How fast can requests be sent? | How much total usage is allowed? |
| Typical period | Seconds, minutes or short rolling windows | Hours, days, months, billing cycles or a fixed allowance |
| Main purpose | Prevent bursts and protect service performance | Govern consumption, cost, capacity or plan entitlements |
| Typical response | Temporary throttling | Blocking until reset, approval or plan change |
| Common example | 100 requests per minute | 1,000,000 requests per month |
| Usual client action | Wait, slow down and retry | Reduce total use or request more capacity |
Many APIs use both. Azure's throttling examples combine a short-term call limit with a longer-term call or bandwidth quota. An API might allow bursts of up to 100 requests, a sustained rate of 20 requests per second and 5 million requests per month. The burst limit stops sudden spikes, the sustained rate controls ongoing traffic and the monthly quota caps total use.
Short-term limits are often built as a token bucket. The bucket holds up to a maximum number of tokens, each request uses tokens from it, and the bucket refills at a fixed rate. The bucket size decides how big a burst can be, and the refill rate decides the sustained speed. Amazon EC2 documents this model, and Amazon API Gateway describes its throttling as a steady-state rate plus a burst capacity. A token bucket lets a client burst for a moment without letting it keep that speed.
Other systems count requests in fixed or sliding windows. A fixed window counts requests inside set intervals. A sliding window looks at a moving period, which is what Azure's rate-limit policy uses. Two APIs that both allow 100 requests per minute can therefore behave differently near the edge of a window.
A rate-limit error is a traffic problem. Check for 429 Too Many Requests, read Retry-After and any headers that report the remaining limit or reset time, and retry with exponential backoff and random jitter. Then send fewer parallel requests, spread the traffic out and find out whether the limit counts per route, key, user, region or account.
A quota error is a consumption problem. Check the quota's scope and reset period, and stop retrying if the allowance will not reset soon. Cut unnecessary calls with caching, batching, pagination or request deduplication, watch your usage before you reach the next limit, and ask for a quota increase or a different plan if the provider offers one.
The error text alone does not always tell you which limit you hit. Check the response body, the headers, the provider documentation and your account's usage dashboard.
"Rate limit", "quota", "limit" and "throttling" do not mean the same thing on every platform. Microsoft usually separates short-term rate limits from longer-term quotas. AWS calls the bucket size and the refill rate of an EC2 throttle limit two quotas. Google Cloud uses "rate quota" for limits expressed as requests per minute.
So judge a limit by what it counts and over which period, whatever the provider calls it. A short allowance on request frequency works as a rate limit, and a long-term or total allowance works as a quota. Burst, concurrency and capacity ceilings need their own reading of the documentation.
A 429 alone does not always tell you the cause either. Microsoft Fabric's throttling documentation describes request-rate limits and capacity limits that both lead to 429 responses.
Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.
Take a look at vroni.com