Currently Available: Need a skilled Software Developer for your next project?
Categories
Tech Stack

How to Retry API Requests Safely

Your code calls an API and the call times out. The server may never have received the request, or it may have finished the work and lost the response on the way back. In the first case, sending the request again is the right fix. In the second case, a careless retry charges a customer twice, overwrites newer data or adds load to a service that is already failing.

Before you retry, check whether the failure is temporary, whether the operation is safe to repeat and whether there is still time left for another attempt.

Decide whether retrying is safe

A write is safe to retry when repeating it cannot cause a second effect, or when the API has a way to prevent duplicates. The failure also has to be plausibly temporary.

Separate reads from writes

Reads are the easiest to retry. A failed GET or list request normally does not change anything on the server, so the client can retry a temporary network error or a suitable HTTP error, with a deadline and backoff.

Some writes are naturally idempotent. An operation is idempotent when repeating it leaves the same intended final state as doing it once. A PUT that replaces a resource is idempotent, because every request sets the resource to the supplied representation. A DELETE usually is too, because deleting an already deleted resource changes nothing.

The HTTP method does not settle the question. The HTTP semantics specification defines PUT and DELETE as idempotent, but an application can still send a notification or create a billing event on every call. Only the API documentation tells you what the endpoint does.

Treat these operations as unsafe to repeat without protection:

  • POST requests that create orders, accounts or payments
  • increments such as "add one"
  • appends
  • requests that trigger an external action
  • read-modify-write updates without a version check

An API can make a POST safe to retry by supporting idempotency keys. A PUT is still unsafe under concurrent updates if the retry overwrites a change that another client made in the meantime.

Use idempotency keys

An idempotency key identifies one logical operation across all of its network attempts. If a user submits one order and the client tries three times, all three requests carry the same key. A second order gets a new key. Generate the key before the first request and reuse it for every retry. Do not generate a new key after a timeout.

On the server, supporting idempotency keys means four things:

  • store the key together with the operation and its parameters;
  • record the result atomically with the side effect;
  • return the original status and response when the key is used again;
  • reject the key if a later request sends materially different parameters.

Stripe's idempotency model saves the status code and body of the first request for a key, including failed results, and returns that saved result for every later request with the same key. It also compares the parameters, so a client cannot use one operation's key for a different operation by accident.

The server has to keep idempotency records long enough to cover realistic retries, including delays from client libraries, queues, gateways, scheduled jobs and manual reconciliation. Each provider picks its own retention window.

Handle unknown outcomes

A timeout does not tell you that the server failed. For a protected write, retry with the same idempotency key and get the original result. For an unprotected write, do not send the request again blindly. Mark the outcome as unknown and recover through one of these:

  • a status endpoint that accepts the operation ID
  • a query by a reference that the client generated
  • a reconciliation job
  • a conditional request with If-Match
  • a resource version or generation check

Conditional requests stop a retry from overwriting newer data. The client reads the resource together with its entity tag (ETag) and sends the update with an If-Match header. If another client changed the resource in the meantime, the server rejects the stale update. Google Cloud Storage's retry guidance uses generation and metageneration preconditions in the same way, to make conditionally idempotent operations safe to retry.

Reporting the outcome as unknown is safer than guessing failure or success. The application can reconcile it later without creating a duplicate.

Classify retryable failures

Retry specific errors, not every exception. The usual temporary failures are:

  • connection resets and temporary DNS or network failures;
  • socket or response timeouts;
  • 408 Request Timeout and 429 Too Many Requests;
  • 500 Internal Server Error, 502 Bad Gateway, 503 Service Unavailable and 504 Gateway Timeout.

The Cloud Storage retry strategy treats these status codes, socket timeouts and dropped connections as worth retrying, as long as the operation also meets its idempotency rules. A 503 means the server is temporarily unable to handle the request, and the response can include a Retry-After header.

Do not retry malformed requests, invalid parameters, failed authentication or authorization, unsupported methods or business-rule errors. Sending them again returns the same error. The request, the credentials or the permissions have to change first.

404 Not Found and 409 Conflict depend on the API. A 404 can be temporary when a newly created resource has not reached every replica yet. A 409 can be worth retrying when the API documents a repeatable read-modify-write sequence. Google IAM asks clients to handle a 409 with the status ABORTED by repeating the whole read-modify-write sequence with backoff, because retrying only the final write keeps failing.

Control timing and retry load

A retry policy decides how long one operation can take and how much extra traffic your service sends during an outage. Immediate retries turn a brief slowdown into a bigger failure, because they send more work to a dependency that is already overloaded.

Respect Retry-After

The Retry-After header holds either a delay in seconds or an HTTP date. APIs usually send it with 429 and 503 responses. Parse both forms, handle invalid values safely and wait at least as long as the server asks. The HTTP specification defines the header, and Microsoft's transient fault guidance tells clients to follow it.

Your own limits still apply. If the server asks for a longer wait than the time left for the request, stop and return a controlled failure, or move the work to an asynchronous process.

Add backoff and jitter

Wait longer after each failed attempt. A simple schedule grows from one second to two, four and eight seconds and then stays at a configured cap. The right values depend on the dependency, the request type and the traffic volume. Google IAM documents this truncated exponential backoff with a small random delay added to every wait.

Jitter adds randomness so that many clients do not retry at the same moment. A common version is full jitter:

cap = min(max_delay, initial_delay × multiplier^retry_number)
delay = random(0, cap)

When the server sends Retry-After, never wait less than it asks:

delay = max(parsed_retry_after, random(0, cap))

If the calculated cap is four seconds, each client picks a delay between zero and four seconds, but none retries before the server's stated minimum. Full jitter, equal jitter and decorrelated jitter all work. No formula is best for every workload, so test how retries bunch up and how the system recovers under failure.

Set attempts and deadlines

Each attempt needs its own connection and response timeout. The whole logical operation also needs one end-to-end deadline that covers request time, response time and the waits between attempts.

Without a shared deadline, every retry starts a fresh timeout, and a user request, thread, connection or queue slot stays busy for an unbounded time. Microsoft's transient fault recommendations count the total time of an operation as the failed attempts plus the retry delays.

Use both a maximum number of attempts and a total deadline, and stop when either one is reached:

deadline = current_time + total_budget
attempts = 0

while attempts < max_attempts and current_time < deadline:
    remaining = deadline - current_time
    timeout = min(per_attempt_timeout, remaining)

    response = send_request(timeout)

    if success:
        return response

    if not retryable(response) or not operation_is_protected:
        return failure

    delay = calculate_server_aware_backoff(response, attempts)

    if current_time + delay >= deadline:
        return failure

    wait(delay)
    attempts += 1

return failure

Send the remaining deadline downstream when you can. A service should not spend almost all of a user's time retrying a dependency and then have no time left to answer.

Limit total retry traffic

Per-request limits do not protect a service when hundreds or thousands of requests retry at the same time. If 1,000 original requests each make two retries, the dependency can receive up to 3,000 attempts. Retry loops at several layers multiply that further.

Set a retry budget for each dependency at the process or service level. Microsoft's example is at most 60 retries per minute against one dependency. A budget can limit retries per time interval, separate retry traffic by endpoint and reserve capacity for high-priority work. When the budget runs out, fail fast, drop low-priority work, use a fallback or put background work on a queue.

Report retries separately from original requests. Retry storms are one of the documented causes of metastable failures, in which retried work keeps a partly failing system stuck in overload even after the original trigger is gone.

Choose one retry owner

A typical request travels through several layers:

client → gateway → service A → service B → database

If every layer retries on its own, one user request can turn into a large number of database attempts. Prefer one retry owner for each dependency call. If several layers must retry, split the total attempt count, deadline and retry budget between them.

Send useful context downstream, such as the remaining deadline, the operation ID and the retry state. A downstream service should know whether it is handling an original attempt or work that an upstream retry has already delayed.

When retries are not enough

Retries fit short faults. Longer outages need mechanisms that take pressure off the dependency and keep the work without holding synchronous resources open.

Use circuit breakers

A circuit breaker lets calls through normally in the closed state. After failures cross a configured threshold, it opens, and calls fail fast or use a fallback. In the half-open state, a limited number of trial calls check whether the dependency has recovered. Microsoft describes this three-state model in its circuit breaker pattern.

The breaker protects the dependency from further retry traffic and protects the caller's threads, connections and memory. Tune the thresholds and the trial rate to the dependency. A threshold copied from another service can open too early for a low-volume dependency or too late for a high-volume one. Combine the breaker with bulkheads, meaning separate worker and connection pools, so that one failing dependency cannot use up all local capacity.

Queue work for later

Background operations should not keep a user request open while they wait through many retries. Put the work on a durable queue, set a maximum retry age and move messages that run out of attempts to a dead-letter queue, where you can inspect them or replay them in a controlled way.

For long-running operations, return an operation ID and let the client poll a status endpoint. The server can use Retry-After to control how often the client polls. Microsoft's asynchronous request-reply pattern describes this approach.

Log every attempt

Record enough to tell recovery apart from hidden failure:

  • the dependency, endpoint and operation;
  • the HTTP method and status code, or the transport error type;
  • the attempt number, the calculated delay and the parsed Retry-After;
  • the elapsed time and the remaining deadline;
  • whether the call used an idempotency key or a precondition;
  • the final outcome and the reason for stopping.

Track the retry rate, retries per original request, the success-after-retry rate, retry-induced latency, circuit openings and budget exhaustion. A high success-after-retry rate means the retries recover from faults. A rising retry rate together with a falling success-after-retry rate points to overload or a failing dependency.

Log a first transient failure at a low level instead of treating it as an incident. Raise the final failed operation and sustained increases in retry traffic to warnings or errors.

Starting points by operation type

Operation Default posture Required safeguard
Read-only GET or list Retry transient failures Deadline, backoff and jitter
PUT replacement Usually retryable Confirm semantics; use version conditions for concurrency
DELETE Often retryable Define behavior for an already deleted resource
POST create Do not retry blindly Idempotency key or server-side deduplication
Payment or order creation Retry only with strong protection Idempotency key and reconciliation endpoint
Increment or append Treat as unsafe Unique operation ID or transactional deduplication
Read-modify-write update Retry the whole sequence ETag, version or compare-and-set condition
Background job Retry asynchronously Durable queue, maximum age and dead-letter handling
What I'm building

Delegate tasks. Get software.

Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.

Take a look at vroni.com

Email updates

Usually a new article and a few links I found interesting.

No spam. Unsubscribe with one click.

Leave a Reply

Your email address will not be published. Required fields are marked *