LLM API error codes are a small vocabulary with a large gap between reading them and acting on them. Two things make the gap: the same code can mean different things at different providers, and some errors are not worth retrying under any policy. The way through is to triage by layer, authentication, permission and account, endpoint and path, capacity and rate, and then decide whether the error is even retryable. If you have not yet made one successful minimal call to your endpoint, do that first, the pattern in our quickstart applies, because every triage table assumes a working baseline.
Triage by layer: API error 401 vs 403, 404, and the rest
Sorting by layer instead of by number keeps you honest, because the fix differs by layer even when the number repeats. The rows below cover the codes you will meet first; each is the official reading, per OpenAI's error code guide and Anthropic's errors reference.
| Code | Layer | First thing to check |
|---|---|---|
| 400 | Request | Malformed request or unsupported parameter value; OpenAI also returns it for a disallowed service tier, Anthropic as invalid_request_error |
| 401 | Authentication | The key itself: OpenAI distinguishes invalid authentication from an incorrect key, and Anthropic's authentication_error covers malformed, revoked, or expired keys |
| 402 | Account | Anthropic only: billing_error means the billing or payment information needs attention |
| 403 | Permission and region | OpenAI's known 403 is a country, region, or territory not supported; Anthropic's permission_error means the key lacks access to the resource, so check organization and workspace settings |
| 404 | Path | The endpoint path or resource identifier in the URL does not exist, per Anthropic's not_found_error |
| 409 | State | Anthropic's conflict_error: the request conflicts with the resource's current state, such as a concurrent modification; resolve, then retry |
One pattern worth noticing early: on an OpenAI-compatible setup, 401 and 404 are frequently configuration problems rather than code problems, a truncated key or a base URL that lost its path. That failure mode has its own walkthrough in the provider switching guide, including the environment and settings surfaces where a stale value hides.
The capacity layer: OpenAI 429 rate limit subtypes, 503, and Anthropic 529 overloaded
This is where the error vocabulary gets genuinely vendor-specific, because the same 429 status carries several different situations. OpenAI's documentation splits it into named cases: rate limit reached for requests, which tells you to pace requests and follow the Retry-After header when present; slow_down, which can appear even when your traffic is inside its nominal limits and asks you to follow the Retry-After header and ramp more gradually; and credit_balance_exhausted, which is not a pacing problem at all, since the organization has no prepaid credits and must add them. A related case, organization_spend_limit_exceeded, fires when an enforced spend limit has been reached.
Anthropic's 429 splits differently, and one variant deserves a callout: a tier spend-cap 429 arrives with no retry-after header and keeps failing until access resumes, so a naive retry loop will spin against it indefinitely. The provider also notes that Claude Code workspace spend limits can surface as a 429 where related conditions surface as a 400, and that sharp usage increases can trigger acceleration-limit 429s, which is an argument for ramping traffic deliberately. For long-running requests, timeout conditions have their own home as well, 504 timeout_error.
Server-side capacity errors mostly look the same but are read carefully: OpenAI's 503 arrives as service_unavailable_error with the code server_is_overloaded, meaning the requested model is temporarily overloaded, and the guidance is to follow Retry-After when present and increase the delay when it is missing. A 500 is the generic server error. Anthropic's 500 is api_error, retried with exponential backoff, and if it persists the provider asks for the request ID when you contact support. And then, 529 overloaded_error, a status code outside the standard set, signifying the API is temporarily overloaded across users. Non-standard codes are exactly the ones that slip past error mappers written from memory; a handler must name it explicitly or it falls into a catch-all.
When capacity is the layer that fails most often, the strategic question is where retries should land: same endpoint, or a fallback. The architecture trade-offs of routing across endpoints and providers are covered in our aggregation versus direct comparison.
Retry or stop: the decision table
The single most useful distinction in API error handling is between errors that might succeed later and errors that will not succeed until something changes outside the request. Here is the split, following the two providers' own guidance.
| Situation | Retry? |
|---|---|
| 429 with a Retry-After header | Yes, honoring the header |
| 503 or 529 (overload) | Yes, with backoff, honoring Retry-After when present |
| 500-class server errors | Yes, bounded, with backoff |
| Connection errors before a response | Yes, standard practice |
| Any 401 or 403 | No. Fix the key or the permission first |
| 404 and 409 | No. Fix the path or resolve the state conflict first |
| Billing and quota conditions, including credit_balance_exhausted and Anthropic's 402 and tier spend-cap 429 | No. The official line is blunt: retrying billing, spend, or quota errors will not restore access, and the spend-cap 429 keeps failing by design until access resumes |
Three principles keep retry logic honest regardless of vendor. First, exponential backoff with jitter, so bursts do not synchronize. Second, always honor the Retry-After header API responses carry when it is present; treat a missing header as a reason for more patience, not less. Third, retry only operations that are safe to repeat; for anything with side effects, make the operation idempotent before it ever reaches the retry layer. And note what your SDK already does: both major SDKs ship automatic retries for a small set of errors with exponential backoff and configurable limits, which is helpful until it silently doubles a retry storm you did not know you were running. Know the setting.
Same problem, different codes across vendors
The most practical reason to branch by layer rather than by number is that business-level problems map to different codes at different providers. A depleted account is a 402 billing_error at Anthropic and a credit_balance_exhausted 429 at OpenAI. A permissions issue is Anthropic's 403 permission_error, while OpenAI's distinct 403 is geographic. Rate pressure can be 429 at both, but the sub-types differ, and one of Anthropic's 429 variants (the spend cap) behaves unlike anything OpenAI's rate limiter produces, since it is not a pacing signal at all.
That has a concrete consequence for your code: keep a mapping table of your own, one row per layer, with the provider-specific codes that land in it, and route handling by layer. When you add a provider or a compatible endpoint, you extend the table instead of rewriting the branches. Treat the codes as the dialect and the layers as the grammar.
A minimal error-handling skeleton
You do not need an elaborate framework to handle this well. You need one handler per layer, with a clear disposition:
- Authentication and permission (401, 403): no retry; fail fast, surface "check key or permission," and log the endpoint and credential source so the fix is obvious.
- Account and billing (402, credit balance and spend cap conditions): no retry; alert an operator, because this class waits on a human action.
- Path and state (404, 409): no retry; log the exact URL and resource IDs, since this layer is almost always a configuration or concurrency bug that should never have been retried into noise.
- Capacity (429 with Retry-After, 503, 529, 500-class): retry with the three principles above; if the retry budget is exhausted, degrade to a fallback model or endpoint rather than dropping the request on the floor.
- Unknown: log and alert, never silently swallow. Generic handlers should treat an unrecognized status, including the non-standard ones, as its own class.
Two logging details pay for themselves in every incident review. Log the error's structured type and code, not just the status line, because both providers layer meaning onto the number: OpenAI asks you to inspect error.code for billing and rate cases, and Anthropic notes that error.type values may grow over time, which is another reason your mapping table, not your memory, holds the truth. And log the request ID the provider returns along with the endpoint and model, because it is the one handle that turns "something failed" into a specific timeline. For the related class where the request succeeds but the response shape is off, streaming events and tool calls being the familiar suspects, that surface has its own checklist at the end of this article; the endpoint context itself, base URL, bearer auth, and the model listing, is documented in the platform's documentation.
Where to go next
- A worked example of the configuration layer. Most 401s and 404s in practice come from credential and endpoint surfaces, and the Cursor custom endpoint setup walks one real case end to end, including its own triage table.
- The shape layer next door. When status codes are fine but responses misbehave, the compatibility checklist breaks the surface into testable pieces.
Read the code, name the layer, decide whether retrying could possibly help, and only then write the branch.