Most LLM observability problems are not tooling problems; they are schema problems. The teams that debug incidents quickly, attribute cost confidently, and answer audit questions without drama are not logging everything. They are logging the right handful of fields, consistently, and they draw one hard line: secrets and customer content never enter the log. This article covers the field list, the things that must stay out, and a part of LLM API data retention that surprises people in both directions: what the vendor keeps is not your audit trail, and what you keep needs to be deliberate, because the controls exist but they are not defaults. The operations articles on this site have been citing logs for weeks; this is the piece about the logs themselves.
LLM API logging: four purposes, one schema
A useful log line for an LLM API call serves four purposes, and the fields you need map directly onto them.
Debugging. The HTTP status code; the error type and message from the response body; the error code where one exists; the request ID; the retry count for this logical request; and the elapsed time. With these, a failure stops being a story and becomes a record: what was attempted, what came back, and where in the retry sequence it died.
Attribution. The model ID actually used, the endpoint, and the caller identity: which feature, team, or customer initiated the call. Where the provider exposes organization or workspace identifiers in response headers, log those too, because they let you reconcile your records against the provider's view. API usage attribution logs are the connective tissue between a line item on an invoice and the code path that generated it. The dimensions worth tagging, and why grouping by them pays off, are covered in the usage attribution guide.
Cost. The token counts from the usage object on each response, and the cache-related fields that show whether cache reads actually happened. Logging those fields turns cost review from archaeology into a query; the mechanics and the fields that prove a cache hit are in the caching guide.
Stability. Counters that answer "is this getting worse": rate-limit responses over time, server-error responses, retry budget exhaustion, and backoff durations actually applied. Those metrics are the observability half of the retry discipline described in the rate limits guide.
One schema, four consumers. The practical trick is a single structured log entry per logical request (not per underlying attempt), enriched with everything above, so that the same record serves the on-call engineer, the finance spreadsheet, and the auditor.
The request ID: your API debugging anchor
If you log only one field beyond the token counts, log the request ID. Anthropic's errors documentation is precise about it: every API response includes a unique request-id header, the same identifier appears as the request_id field in error response bodies, and when you contact support about a specific request, you include that ID. That last part is the point. Support conversations about intermittent failures go nowhere without an identifier both sides can query; with one, they become a lookup.
Reading it programmatically is straightforward but differs by SDK. The Python and TypeScript SDKs expose the request ID as a _request_id property on top-level response objects; the C#, Go, Java, and PHP SDKs expose it through their raw-response accessors, and Ruby through middleware. Reading other response headers, such as anthropic-organization-id and anthropic-workspace-id, uses the same raw-response accessors in every SDK except Ruby, and the same middleware in Ruby. On Claude Platform on AWS, responses carry two IDs: the AWS request ID (x-amzn-requestid), which is indexed in CloudTrail and is the one to use for CloudTrail lookups, and the Anthropic request ID, which is the one for support tickets. Log both if you run there; they answer different questions. The general principle predates any vendor: the request ID plus the error code is the pair that pins one failure to one call, and it is exactly what the error code guide assumes you have when triaging.
What never goes in the logs
Two categories, no exceptions worth making. First, credentials: API keys, bearer tokens, and authentication headers never belong in log output, at any level, including debug. It is worth stating plainly because the leak paths are boring: an over-eager request logger that prints headers, an error handler that serializes the whole request object, a support ticket where someone pastes a config. The same discipline applies identically in CI pipelines and in application logs.
Second, raw customer content: full prompts and completions, especially when they contain personal or regulated data. This is the harder line to hold, because raw content is genuinely useful for debugging. The replacement toolkit: log lengths and hashes instead of bodies; log classification labels or redacted samples when you need a feel for the data; keep a separate, access-controlled store if regulatory requirements force content retention, with its own retention timer. A working pattern is a three-part filter: a whitelist of field names that may be logged, a scrub pass before write that drops anything matching a credential shape and raises an alert when it fires, and access control plus retention limits on the log store itself. The alert matters; if a key shape ever appears in a log stream, you want to know in minutes, not at the next audit.
LLM API data retention: what the vendor keeps
Here is the part many teams misunderstand in one direction. Vendor-side logs exist, and you do not control them by default. OpenAI's data controls documentation states both halves clearly.
The reassuring half: as of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models unless the customer explicitly opts in to share data.
The operational half: abuse monitoring logs are generated for API usage to enforce usage policies, and they may contain customer content, such as prompts and responses, plus derived metadata like classifier outputs. By default these logs are retained for up to 30 days, unless longer retention is required by law or reasonably necessary to protect services or third parties. Eligible customers can get customer content excluded from those logs by being approved for Zero Data Retention or Modified Abuse Monitoring, both subject to prior approval and additional requirements. The conclusion for your architecture is blunt: vendor retention is not your audit trail. If you need records of what happened, when, at what cost, and with which caller, you keep those metadata yourself.
Do not overcorrect in the other direction either. "Not used for training" does not mean "nothing is retained," and the 30-day figure is a documented retention window for one specific log category, not a promise that every trace of every request is unrecoverable after 30 days. Write your internal policy against the actual controls, which the same document details per endpoint.
Two retention controls people misread
Two details from that documentation are worth engraving, because they break assumptions.
First, Zero Data Retention changes behavior, not just policy: with it enabled, the store parameter for the Responses and Chat Completions endpoints is always treated as false, even if a request attempts to set it to true. So an application that assumes "I set store=true, therefore this conversation is retrievable" is simply wrong under ZDR, and the failure mode is silent, because the request succeeds. If your product depends on stored state, verify the endpoint behavior under your organization's actual data policy instead of assuming your parameter wins.
Second, the Responses API has a 30-day application state retention period by default, or when store is set to true; in those cases response data is stored for at least 30 days. Teams that treat the endpoint as stateless can be surprised in both directions: data they assumed was gone is retained, and data they assumed was kept is subject to policy-dependent behavior. The same page lists retention per endpoint, including which endpoints are eligible for the stricter controls at all, and it notes that data residency is a project-level configuration option subject to eligibility, with pricing that may differ from standard endpoints. Confirm both against the current page and your account, not against a blog post, including this one.
Logging and retention as a procurement dimension
Once you see logging as a capability surface, it belongs in vendor selection, not in the fine print. Four questions worth asking during evaluation: can you retrieve a request ID for any individual call, including failures; can you read organization and workspace identifiers from responses so attribution is provable; is the retention posture configurable to the level your compliance team needs, and what does approval require; and where is data processed and stored, if residency matters to you. These slot directly into the architecture trade-offs in the aggregation vs direct access comparison, because an intermediary layer changes who holds the logs and which identifiers you can see. The endpoint-level details for whatever platform you run against live in its own documentation.
The field map
The operations cluster on this site has grown into a set of articles each relying on the same observability substrate. Here is how they connect to the fields in this schema:
| Article | What it needs from your logs |
| :- | :- |
| Usage and cost attribution | Token counts, model, caller tags, org and workspace identifiers |
| Error code triage | Status codes, error type and code, request ID |
| Prompt caching | Cache read and write token fields, to prove hits |
| CI integration | Secrets hygiene: credentials never in logs at all |
| Rate limits and retries | Rate-limit and error counters, backoff and budget metrics |
The implementation checklist
- Define a whitelist schema: field name, type, and nullability for every logged field, and review it when features change.
- Capture the request ID and error code on every call, and store them as first-class fields, not inside a message string.
- Log usage token fields and cache fields as structured numbers, per logical request.
- Tag every call with feature, team, and customer identifiers, and reconcile those tags against the provider's headers.
- Filter before write: drop credential-shaped strings and known PII patterns, and alert when the filter fires.
- Separate debug logs from audit records, with different retention windows for each.
- Restrict who can read the log store, and log access to it.
- Re-audit the schema quarterly against the current provider documentation, because both endpoints and retention controls change.
None of this is glamorous either, but it is the difference between a stack you can operate and a stack you can only hope about. The logs are where that choice gets made, one schema decision at a time.