A rate limit error is not one situation, it is two, and they look identical in a log until you check a single response header. One is temporary throttling: wait as instructed and the request goes through. The other is a quota or spend cap: retrying, including your SDK's automatic retries, will keep failing until access is restored, so effort spent backing off is pure waste. This is the part of the LLM API rate limit story that a status code alone cannot tell you, and both major vendors document the distinction on their rate limit pages, with specific guidance on how to react. This guide walks through classifying a 429, using Retry-After correctly, building a retry budget that cannot spiral, and understanding why limits behave the way they do at the architecture level.
The LLM API rate limit decision: retryable or not
The first source to read is the provider's own documentation: OpenAI's rate limits guide and Anthropic's rate limits page both spell out which conditions deserve a retry.
For temporary throttling, OpenAI documents 429 responses with error type rate_limit_error and code slow_down, meaning your request rate increased too quickly; the guidance is to follow Retry-After when it is present, reduce your request rate, and then increase it gradually. Anthropic documents rate limits measured in requests and tokens per period; exceeding them returns a 429 describing which limit was exceeded, along with a retry-after header indicating how long to wait. Both of these are wait-and-retry situations.
Quota and spend caps are different. Anthropic states that when usage reaches an organization's spend cap, the response is a 429 whose error type is rate_limit_error, the same type as a rate limit, but with no retry-after header, and that retrying, including the SDK's automatic retries, fails until access resumes; on the Messages API you can separate the two cases by checking the error details code for an enforced spend limit. Its documentation also notes that raising the tier restores access, which is a human decision, not a backoff strategy. OpenAI's documentation is equally explicit from the other direction: Retry-After may be present on 429 responses caused by a temporary rate limit, and it "does not mean that quota, billing, or other errors that require user action can be resolved by retrying." Its hard spend limits make affected requests return 429 as well, and the fix is adjusting the limit, not waiting longer.
That yields the classification rule this article is built around: a 429 carrying Retry-After is a throttling signal: obey it. A 429 without Retry-After should be treated as "check quota, billing, and tier first," not as an invitation to retry harder. Providers can also use the same status code for other errors, so inspect the error body before choosing a recovery action.
Reading 429 Retry-After properly
When Retry-After is present, read its semantics precisely: it is a minimum wait, not a target. OpenAI's guidance is to wait at least that long, and to add a small random delay on top so that multiple clients do not all retry at the same moment. That second half is easy to skip and expensive to skip: a fleet of workers that all obey the header to the second will stampede the moment the window opens.
When the header is missing or invalid, or when you are handling conditions outside its scope, fall back to exponential backoff with jitter: wait briefly, then increase the delay after each unsuccessful attempt, with randomization to spread out retries. The same logic covers temporary overload responses; OpenAI documents 503 responses with service_unavailable_error and code server_is_overloaded, where the guidance is to follow Retry-After when present, then retry, increasing the delay between attempts if the error continues.
One more nuance from the same documentation: a slow_down can occur even when your traffic sits within its requests-per-minute and tokens-per-minute limits, because it reflects how quickly traffic increased rather than whether you exhausted the quota. The documented remedy is a ramp that rises gradually instead of a step change.
Exponential backoff with jitter is not optional
Teams that treat jitter as a nice-to-have tend to rediscover why it is standard practice. Deterministic backoff synchronizes clients: every worker that failed at the same moment waits the same amount and fires at the same moment, converting a transient throttle into a self-inflicted traffic spike that keeps the limit tripped. Jitter decorrelates the retries. Combined with Retry-After, the pattern is simple: wait the instructed minimum plus a random fraction, and apply exponential growth only when no instruction is available. Neither part of this is exotic or vendor-specific; it is the minimum viable retry policy for any synchronous API integration.
Retry budgets: cap attempts and total time
The guidance that prevents most retry storms is boring and specific: limit both the number of attempts and the total time spent retrying. A count without a deadline lets a long sequence of waits stack up behind one user request; a deadline without a count lets a burst of quick failures hammer the endpoint. You want both, and you want them enforced by one piece of code that all callers share.
The most common way teams violate their own budget is nesting. If your application manages retries and the SDK underneath also retries automatically, the effective attempt count is multiplicative, not additive, and the total time can stretch unpredictably. OpenAI's documentation addresses this directly: if you manage retries in your application, disable SDK retries or account for them in your limits so nested retry loops do not multiply requests. It also warns that each official SDK retries eligible 429 and 503 responses subject to its own retry settings, and that handling of Retry-After, especially long delays, varies by SDK version and configuration, so verify the behavior of your installed version instead of assuming a server delay is fully respected.
Finally, decide what happens when the budget is exhausted, and make it loud: fail the request fast, log the classification (Retry-After present or not), and alert when a non-retryable 429 appears, because that class of error means a human has something to do. Quietly retrying a quota error until timeout is how a five-minute incident becomes a fifty-minute one.
Why token bucket rate limiting rejects average-rate traffic
A frequent surprise: your measured average rate looks comfortably under the limit, yet requests are still rejected. The mechanism explains it. Anthropic documents that its API uses the token bucket algorithm, where capacity is continuously replenished up to the maximum rather than being reset at fixed intervals, so a burst can exhaust the bucket even if the day's average is modest. Smoothing your traffic is therefore worth more than lowering your daily totals. A client-side queue that paces requests, and an API concurrency limit per caller, do more for stability than any retry policy can.
Bursts are not the only shape of the problem. Anthropic also documents acceleration limits: a sharp increase in usage can itself produce 429 errors, with the guidance to ramp up traffic gradually and maintain consistent usage patterns. The practical reading is that capacity planning should include how fast you grow, not just how much you send: warm the ramp, avoid step changes at deploy time, and keep traffic patterns stable enough that a dashboard spike does not become a rejection spike.
Layered limits: organization, workspace, project
Limits in both ecosystems are enforced at levels, and knowing the level changes where you fix things. Anthropic enforces service-configured limits at the organization level, and additionally supports user-configurable limits per workspace, which lets a workspace be capped lower than the organization so one workload cannot starve the others. OpenAI exposes project-scoped token limits through response headers, which may be present when a project-scoped token limit applies. Anthropic's Message Batches API adds its own layer: a separate rate limit shared across models, plus a cap on batch requests sitting in the processing queue, where a request counts until it has been successfully processed.
The architectural implication is to mirror the provider's structure in your own: separate workspaces or projects per workload, per team, or per product surface, so that a runaway batch job degrades its own slice instead of the shared pool. This is also where attribution matters: when a limit trips, the first question is who consumed the capacity, and the metrics needed to answer it are the same ones described in the usage attribution guide.
Rate limit headers as control signals
Headers are not just diagnostics; they are live control signals. OpenAI exposes a family of headers covering the limit, the remaining amount, and the reset timing, separately for requests and tokens, plus the project-scoped variants above; the reset values are duration-style strings. A client that reads the remaining count and the reset window can throttle itself before it earns a 429: when remaining capacity is low or the reset is imminent, slow the queue down preemptively rather than sprinting into the wall. Two cautions from the documentation: header presence and meaning should be confirmed against the providers' current pages, and values should never be hardcoded into your client. Treat them as inputs to an adaptive limiter, not as constants to bake in.
Where this fits, and the checklist
This article is the deep end of a pool the error code guide introduces: that piece maps the codes, this one handles the 429 properly. To reduce how often you meet one at all, the caching guide covers cutting repeat work at the source, and the aggregation discussion covers when a multi-provider layer is the right architectural answer to capacity questions. The endpoint and protocol details for whatever service you use live in its documentation.
The checklist, in one place:
- Retry only retryable errors: temporary throttling and overload, not quota, billing, or anything demanding a human.
- If
Retry-Afteris present, wait at least that long, plus a small random delay. - If it is absent or invalid, use exponential backoff with jitter.
- Budget both attempts and total time, in shared code.
- Disable or account for SDK retries so loops do not nest.
- Pace traffic client-side: queue, smooth, cap the API concurrency limit per caller.
- Mirror limits per workspace or project, and attribute usage so throttling decisions have an owner.
- Read rate limit headers adaptively, and alert on non-retryable
429s instead of retrying them.
None of this is glamorous. It is the difference between an integration that degrades gracefully under pressure and one that turns every capacity bump into an incident review.