All stories

Scrapebento journal

Avoiding HTTP 429 Errors and Rate Limiting

A delay in each worker does not control the total traffic your scraper sends. Servers may enforce limits by IP address, API token, account, endpoint, time window, or the number of requests running at once. This guide explains how to diagnose 429 responses, interpret Retry-After, coordinate workers around a shared request budget, and retry without creating another burst. The goal is to match the limit the server actually enforces, rather than keep increasing a guessed delay.

You add a two-second pause, restart the scraper, and watch the first few pages load. Then the 429s return.

The obvious reaction is to increase the delay. Before doing that, check what the server actually receives.

Suppose ten workers each send one request every two seconds. Together, they send five requests per second, or 300 per minute. If the server allows 60 requests per minute for the token they share, every worker can obey its own delay while the scraper exceeds the allowance fivefold.

That is the difference between a local pause and a shared request budget.

The server might count traffic over a rolling window rather than a clock minute. It might group requests by IP address, token, account, or endpoint. It might also restrict how many requests can remain active at the same time. The HTTP status code does not tell you which counting method is in use.

Start with the response and the service’s documentation. Find out what is being counted, which workers contribute to that count, and when another request is allowed.

This guide focuses on request volume and concurrency. Browser fingerprinting and challenge-based access controls are separate problems, although the distinction matters when deciding whether another retry makes sense.

What a 429 Actually Tells You

A 429 Too Many Requests means the server is reporting excessive request activity. It does not, by itself, identify the quota, explain how the server groups your requests, or promise that a short pause will resolve the problem.

Nearby status codes offer different clues. A 403 Forbidden generally means access was refused, while a 503 Service Unavailable indicates temporary unavailability. But implementations vary: some APIs also use 403 responses for rate limiting. Read the response body and documented behavior rather than assigning every status code a rigid retry rule.

Consider this request:

GET /api/products?page=12&limit=50 HTTP/1.1
Host: catalog.example
Accept: application/json
Authorization: Bearer <redacted>

The server responds:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 30
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 0

{
  "error": "rate_limit_exceeded",
  "message": "Limit of 60 requests per minute per API token.",
  "scope": "token"
}

All requests and responses in this article are illustrative. The endpoints, quotas, and application-specific fields are fictional, and unrelated headers are omitted.

The most useful detail here is per API token. It explains why slowing one worker will not help while the other nine continue using the same allowance.

The response gives you three actionable facts: the shared allowance is exhausted, the server asks for a 30-second wait, and the affected scope is the token. Put that token’s outgoing work on hold, not just the URL that happened to receive the error.

Headers such as X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset can provide additional context. Check the API’s definitions, especially the reset format. A reset value may be a timestamp rather than a number of seconds to sleep.

Capture enough evidence to reconstruct the failure: timestamp, method, endpoint, status, request ID, relevant headers, and a safe body excerpt. Redact authorization headers, cookies, tokens, and personal data before writing them to logs. Debugging throttling should not create a credentials leak.

Honor Retry-After Before You Guess

Read Retry-After on every relevant response, not just the first failure.

A value of Retry-After: 30 asks you to wait at least 30 seconds after receiving the response. It is a waiting instruction, not a reservation for your next request. Another request can still be rejected after that time.

For the earlier example, hold the affected token’s requests, then let the failed request re-enter the shared queue after the cooldown. If capacity is available, repeating the same request might produce:

HTTP/1.1 200 OK
Content-Type: application/json
X-RateLimit-Limit: 60
X-RateLimit-Remaining: 59

{
  "page": 12,
  "items": [
    {
      "id": "p_501",
      "name": "Desk lamp"
    }
  ],
  "next_page": null
}

That successful response is not a reason to release every waiting worker simultaneously. Continue with the shared pacing policy that keeps total traffic within the allowance.

Retry-After can also contain an HTTP date:

HTTP/1.1 429 Too Many Requests
Date: Mon, 05 Oct 2026 10:00:00 GMT
Retry-After: Mon, 05 Oct 2026 10:01:30 GMT

Here, the retry time is 90 seconds after the response’s stated server time. Parse HTTP dates with a timezone-aware parser. A valid server Date can help account for clock differences; when it is unavailable, calculate against a correctly synchronized UTC clock.

Handle missing or malformed values without crashing the job. Use a bounded backoff policy when there is no usable waiting instruction.

Most importantly, a local backoff cap must not shorten a longer server-provided wait. A 60-second cap in your configuration does not turn Retry-After: 120 into permission to retry after one minute.

Exponential Backoff With Jitter, Not sleep()

The problem is not the sleep() function. The problem is giving every worker the same fixed delay regardless of what happened.

Exponential backoff increases the retry interval after repeated failures. Jitter randomizes the timing so workers do not all return together. With full jitter, the first retry might wait between zero and one second, the next between zero and two seconds, then zero and four seconds, with a configurable ceiling.

When a valid Retry-After is present, keep that wait as a floor. One possible policy is to add the randomized backoff after the required wait.

For example:

Attempt 1: 429, Retry-After: 30
Wait: 30.7 seconds

Attempt 2: 429, Retry-After: 45
Wait: 46.4 seconds

Attempt 3: 200 OK

The extra fractions are illustrative jitter. Notice that the second response’s waiting instruction is honored too. The scraper does not fall back to a shorter delay simply because it has already waited once.

Set an explicit retry budget. For example, allow one initial attempt and up to three retries, for four total attempts. Also set a total time budget. If the required wait would exceed the job’s deadline, defer the work rather than shorten the wait.

Choose retryable failures deliberately. Read-only GET requests can often be retried after temporary failures, but do not blindly retry every error. Authentication failures, permission problems, and invalid URLs need investigation or a documented recovery action. Requests that create or change data need an idempotency guarantee before automatic retries, otherwise a lost response can lead to the same action being performed twice.

Finally, stop a repeatedly failing target from consuming the entire job. A circuit breaker temporarily pauses the affected scope after persistent failures, then allows a small test request after the cooldown. If it succeeds, resume gradually. If it fails, keep the remaining work deferred.

Pace, Concurrency and Distribution: Not Reaching the Limit at All

Retries help a scraper recover. Shared pacing helps it avoid the rejection in the first place.

Start with a budget per host, then match the additional scopes the service enforces. An account-wide allowance must be shared across that account’s workers. An endpoint-specific limit needs its own control alongside the broader allowance. Independent local counters do not enforce a shared budget.

For the fictional 60-request-per-minute token limit, an initial configuration could look like this:

Quota scope: one API token across all workers
Documented allowance: 60 requests per minute
Chosen starting rate: 48 requests per minute
Minimum gap between request starts: 1.25 seconds
Maximum requests in flight: 2

The 48-request target is an example, not a universal safety margin. It leaves some room below the stated allowance, but it does not override a tighter endpoint limit or another documented restriction.

A token bucket can enforce the shared rate: requests spend tokens, and tokens replenish at a configured rate. Pay attention to bucket capacity as well as refill speed, because accumulated tokens permit bursts. A leaky-bucket-style queue can instead release work at a steady pace. Choose the behavior that fits the server’s policy.

Concurrency needs a separate control. A fixed worker pool limits simultaneous work, but it does not establish a fixed request rate. Two workers completing requests every tenth of a second could still produce roughly 20 requests per second.

Now consider a slower endpoint:

GET /api/reports/daily?date=2026-10-04 HTTP/1.1
Host: catalog.example
Accept: application/json
Authorization: Bearer <redacted>

Suppose two reports are already being generated. A third request receives:

HTTP/1.1 429 Too Many Requests
Content-Type: application/json

{
  "error": "concurrency_limit_exceeded",
  "message": "Only two report requests may run at once.",
  "scope": "account",
  "max_in_flight": 2
}

The immediate problem is not requests per minute. It is that two requests are still running. Queue the third until a slot becomes available, while continuing to enforce any request-rate limit. Long-running requests can trigger concurrency limits even when request starts are relatively infrequent.

Distribution does not remove the need for coordination. In the earlier example, deploying three machines with independent 48-request-per-minute allowances would produce 144 requests per minute against the same 60-request token quota. Use a shared coordinator or allocate portions of the total budget explicitly.

Changing IP addresses also does not create more capacity under an account-wide allowance. The account remains the thing being counted.

Before adding more infrastructure, remove unnecessary requests. Cache reusable responses where permitted, de-duplicate equivalent requests, and skip pagination you do not need. Keep request parameters and authentication context in mind when deciding whether two requests are genuinely interchangeable. Fetching fewer fields may reduce payload size, but it does not necessarily reduce a request-count quota.

When Rate Limiting Is Enforcement, Not Capacity

A scraper that still receives 429s after slowing down has not necessarily encountered bot detection.

First check the less dramatic explanations. Another job may be using the same account. A search endpoint may have a tighter allowance. A longer time window may still be exhausted. Or a few slow requests may be filling the concurrency limit. Some services enforce several restrictions at once.

Treat the response body as evidence. A message that identifies a depleted quota points toward pacing and scheduling. A message that explicitly says automated access is not permitted points toward an access-policy problem.

A missing Retry-After header does not settle the question, because that header is optional on a 429 response. Its presence is not proof that request volume is the only relevant restriction either.

When the response clearly requires authorization or prohibits the requested access, stop spending the retry budget on an unchanged request. Check the permitted interface, request appropriate access, or discuss the required volume with the operator.

Browser fingerprinting and challenge handling sit outside this article’s scope. The useful distinction here is whether your next action should change the traffic schedule or resolve an access requirement.

Making Rate Limits a Monitored, Non-Breaking Condition

A recovered 429 and a stalled crawl should not look identical in your dashboard.

Track throttling by host and, where useful, by the affected quota scope. Record the proportion of attempts receiving 429s, retries per completed item, successful throughput, and time spent waiting. Also watch queue age and data freshness, because a scraper can reduce its error count simply by doing almost no useful work.

Consider a reporting interval with 1,000 request attempts and 200 throttled responses. That is a 20% rejection rate. The next question is whether the retries recovered the work and the queue caught up, or whether unfinished pages continued to accumulate.

Likewise, reducing the rejection rate from 20% to 2% is not automatically a success if the oldest queued work keeps getting older.

Alert on sustained failure or meaningful impact, rather than treating every recoverable retry as an incident. Keep the early failures visible in logs, and escalate when retries are exhausted or the job can no longer meet its requirements.

Make progress resumable. Save completed pages or cursors so a throttled job does not restart from the beginning. Keep result writes idempotent so a retried page does not create duplicate records.

Where the application can tolerate it, retain the last successfully fetched data and show when it was checked. Do not replace a valid dataset with an empty result just because a refresh was throttled. Where stale data would be misleading, mark it unavailable instead. The aim is to make the limitation visible without unnecessarily stopping unrelated work.

A Checklist for Scraping Within the Limits

Before increasing the delay again, check the whole request path:

  • Capture useful error context. Keep relevant headers and a safe body excerpt, with credentials redacted.
  • Honor every valid Retry-After. Apply the cooldown to the workers sharing the affected quota.
  • Bound retries. Use jitter, a retry count, and a total time budget.
  • Share the request budget. Include all workers, machines, and retry attempts that contribute to the same limit.
  • Control concurrency separately. Limit active requests as well as request starts.
  • Remove unnecessary requests. Cache appropriately, de-duplicate work, and avoid needless pagination.
  • Check what the rejection means. Distinguish quota exhaustion from an authentication or access-policy problem.
  • Monitor useful progress. Watch completed work, queue age, and freshness alongside the 429 rate.

The question is not whether every worker slept. It is whether all the requests, including retries, stayed within the budget the server actually enforces.