A 403 is the most misread status code in LLM integrations. When a runbook starts with an LLM request failed: 403 line, the instinct is to treat it like a rate limit that went to the wrong number, or an authentication hiccup that a retry will clear. Neither is right. This status code covers four completely different problems: regional policy blocks, credential permission scope, policy and content denials, and blocks that happen inside your own network path before the request ever reaches the provider. The fixes are different in each case, and for all four, retrying is the wrong first move. This is the single-code deep dive that follows the full error code triage; if you arrived here from a 429, the sibling piece on retryable limits is the rate limits guide.
LLM API 403 forbidden is not one error
Before diagnosing, internalize what the four classes are, because the entire rest of this article is telling them apart.
- Region and territory blocks. The request comes from a location the provider does not serve, and the denial is policy, not malfunction.
- Credential permission scope. The key is valid and recognized, but it is not allowed to do this particular thing, here.
- Policy and content denials. The request itself is not permitted, independent of who is asking.
- Gateway, proxy, and edge blocks. Something in your own path (a corporate egress control, a self-managed gateway, a security product) answered instead of the provider. The request never arrived.
Notice what is absent: nothing on that list is a transient condition. That absence drives the first rule.
Rule one: 401 asks who you are; 403 asks what you may do
The 401 and 403 pair is often treated as interchangeable, and the error tables make the division clearer than the folklore does. OpenAI's error reference lists 401 cases about identity and organization: invalid authentication, an incorrect key, an account that is not part of an organization, and a request IP that does not match the configured allowlist. Anthropic's error table draws the same shape for authentication issues on its side. The 403 family, meanwhile, is about authorization: geography, permission scope, and policy. The distinction matters because the fixes live in different places: identity problems are solved in your credential configuration (the Claude Code precedence piece covers how several credentials interact on that side), while authorization problems are solved in permissions, policy, or location. A useful shorthand: 401 means the door does not know you; 403 means the door knows exactly who you are and is still closed.
Which leads to the operational rule this article exists to state: do not apply backoff-retry logic to a 403. The retry machinery from the rate limits guide exists for conditions that can change on their own; a 403 is a decision, and decisions do not expire on a timer. Retrying a region block or a permission denial just spends budget re-asking a question whose answer is on file.
The four kinds of 403, one by one
Region and territory policy
OpenAI labels this entry in its error table as 403 - Country, region, or territory not supported, with the cause stated plainly: you are accessing the API from an unsupported country, region, or territory, and the solution points to the supported countries page. That is worth reading twice: the provider treats this as a policy boundary with a documentation trail, not as a bug. The handling path follows from that. Confirm the location policy on the provider's page (for the platform you are reading this on, the equivalent surface is our own supported regions page), and if your situation needs an exception or a formal answer, the path runs through the provider's official channels: account representatives or support. This is a compliance and account conversation, and it is the one 403 class where "fix your code" is the wrong frame entirely.
Credential permission scope
The second class looks like authentication but is authorization: the key is fine, the permission is not. OpenAI's documentation handles this case in its authentication troubleshooting, listing an API key that does not have the required permissions for the endpoint among the causes, and pointing the resolution at the key, the organization ID, and project-level keys. At the SDK layer, the same situation surfaces as a PermissionDeniedError, whose documented cause is having no access to the requested resource and whose documented solution is to check the API key, organization ID, and resource ID. Anthropic names the case directly in its error table: 403 permission_error, meaning the API key does not have permission to use the specified resource, with the fix path being your organization's access settings and workspace settings in the console. So the API key permission denied at the endpoint is a configuration problem with a specific checklist: right key, right organization, right workspace, right resource ID. If those check out and the denial persists, the next step is an administrator who can see the access settings on the provider side.
Policy and content denials
Not every 403 is about access at all; some are the platform declining the request itself. This class intersects with the other 4xx families and deserves a clear boundary: a 400-class validation error means the request was malformed, a model-side refusal is the model declining inside a successful response, and a policy denial is the platform refusing the operation. The practical first move for this class is always the same: read the error body. Both ecosystems return structured errors with a type and a message, so the type field tells you which family you are in before you attach meaning to the status code. Then go to the policy documentation that governs the action you attempted, rather than iterating on your prompt and hoping. If the request triggers a policy boundary, prompt engineering is not the tool; reading the rule is.
Gateway, proxy, and edge blocks
The fourth class never reaches the provider, and that is its signature. A self-managed gateway, a corporate egress path, a CDN, or a security appliance can all answer with a 403-shaped denial of its own. The diagnostic is shape: API error responses are JSON with an error object, a type, a message, and a request ID; an edge block usually returns something else, an HTML page, a plain-text line, a vendor-specific structure, and no request ID because no request exists. Anthropic's documentation even shows the concept from the provider side: it notes that for one of its error conditions, the response is returned by its edge infrastructure before the request reaches the API servers, which is exactly the pattern to look for. The fix is network archaeology, not account work: compare a direct call and a call through the full path, keep both raw responses as evidence, and work with whoever owns the middle.
A decision path you can follow in order
- Read the body, not just the code. Get
error.typeanderror.messageinto the incident notes first. - Check for a request ID. If neither the body nor the headers carry one, suspect class four before class one.
- Classify into the four. Region, permission, policy, or edge.
- For region: policy page first, official channels second. Nothing else.
- For permissions: re-verify key, organization, workspace, resource ID, then route to an administrator.
- Record the conclusion. Which class, which evidence, what changed. The field list for that record, request ID included, is the logging guide's territory, and it is what turns the next occurrence from a mystery into a lookup.
What to attach to a support report
Both vendors design around one artifact: the request ID. Anthropic's guidance is explicit that every response includes a unique request-id header, that the same identifier appears as the request_id field in error bodies, and that support conversations should include it. OpenAI's error reference, for its part, asks you to provide the request data and headers you sent when you escalate a persistent error. When you escalate, attach: the request ID; the UTC timestamp; the endpoint and model; the full error type and message; whether the call went through any gateway; reproduction steps; and what you have already ruled out. That list is deliberately boring, and it is the difference between a ticket that moves and a ticket that bounces.
The compliance boundary, stated plainly
Regional restrictions on AI services are policy and compliance decisions made by vendors under real legal constraints, and the only durable way to work with them is through the same official channels that made them: the supported regions documentation, and the provider's support or account team when an exception or clarification is needed. This article deliberately does not discuss, endorse, or hint at any method for working around a regional restriction, because none of those approaches are reliable, and all of them put the account, the data, and the business relationship at risk. If your team operates in a constrained location, the productive conversation is about the officially supported path: what the provider offers there, what alternatives exist for the workloads that must run, and what the documentation says today. That conversation, with your account team and your legal counsel, is also the one that produces an answer you can defend in an audit.
Where this fits
This article sits inside a diagnostics cluster. The full error code triage maps every status family; this article is the deep dive on one of them. The retry guide explains when retrying is right, which this article's 403 rule deliberately contradicts for this code. The logging guide supplies the evidence fields used above. And when whatever endpoint you call sits behind a compatible interface, the endpoint-level facts live in the platform's own documentation.
The 403 checklist
- Classify first: region, permission, policy, or edge.
- Confirm the source: an API error body with a request ID means the provider; its absence points at the path.
- Region class: policy page and official channels only. No exceptions, no cleverness.
- Permission class: verify key, organization, workspace, and resource ID in that order.
- Policy class: read the error type and the governing policy document before touching a prompt.
- Edge class: compare direct and full-path responses, and preserve both as evidence.
- Escalate with the artifact list, request ID first.
- Never backoff-retry a 403; and baseline the same key across local, pipeline, and production environments, because the fourth class hides in exactly those differences.
Four problems, one code, and one discipline: read the body, find the class, fix the right layer. Boring diagnosis beats clever retries every time it matters.