Blog

Tool Calling in Production: Schema Design, tool_choice, Parallel Calls, and the Loop

Tool calling is easy to demo and hard to keep boring: schemas need to be enforceable, turns need a documented number of calls, and the loop needs a plan for failures. Here is the production view.

Tool Calling in Production: Schema Design, tool_choice, Parallel Calls, and the Loop

Tool calling is not hard to start; it is hard to keep boring. The demo works in an afternoon, and then production asks the harder questions: will the schema actually be enforced, how many calls can one turn contain, what happens when parallel calls arrive, how do results get stitched back, and who handles a failing tool. Both major vendors document those answers precisely, and the differences between them are where integrations get interesting. This guide follows the function calling best practices that hold up outside the demo: the loop anatomy, schema requirements, the tool_choice parameter, parallel calls, result handling, and the cross-vendor semantics that decide how you write the orchestration code. It pairs with the structured outputs guide from earlier this week: that article is about what the model is allowed to say, this one is about what it is allowed to do.

The tool use round trip, in three layers

OpenAI's function calling documentation establishes the vocabulary that makes the rest of this article tractable, in three layers. Tools are functionality we give the model: a function defined by a JSON schema, and the docs are precise that a function is one kind of tool, alongside custom tools (free-form text in and out) and the platform's built-in tools. Tool calls are the model's request to use one of those tools, arriving as a special kind of response. Tool call outputs are what we generate: the result of actually running the tool, referenced back to a specific call so the conversation stays coherent.

The documented flow is five steps, and reading it once is the best defense against spaghetti orchestration: make a request with the tools available; receive a tool call; execute your code with the provided arguments; make a second request carrying the tool output; receive the final response, or more steps if the model needs them. Anthropic's overview describes the same shape from the other side: Claude decides to call a tool based on the request and the tool's description, returns a structured call your application executes, and your result comes back for the final answer. The whole design goal is stated in one line of OpenAI's guide: sending back the tool definition, the original prompt, the model's call, and the output together is what lets the model finish the job.

Function calling best practices start with the schema

The single most valuable line in the OpenAI documentation: setting strict: true ensures function calls reliably adhere to the function schema, instead of being best effort, and the recommendation is to always enable it. That is an upgrade from the bad old days of asking nicely, but this strict mode function schema has hard requirements, because under the hood it leverages the structured outputs machinery:

  1. additionalProperties must be set to false for each object in parameters.
  2. All fields in properties must be marked as required.

Optional fields are expressed the structured-outputs way, by adding null as a type option rather than omitting the key from required. A schema that misses these requirements is rejected with details about the missing constraints, which is the good failure mode.

The trap is the default. If you omit strict, the behavior depends on the API: on the Responses API, the service attempts to normalize your schema into strict mode when possible, and falls back to non-strict, best-effort function calling if it cannot; the fallback is visible because the response tool shows strict: false. Chat Completions requests remain non-strict by default. The operational reading, quoting the structure of the docs: do not assume "no strict written" means best effort, and do not assume "strict written" means guaranteed, because the API surface and a silent fallback both change the answer. Log the effective strict state of the tools in your responses and alert when it disagrees with intent.

tool_choice and parallel tool calls

By default, the model decides when and how many tools to use; the tool_choice parameter lets you override that. The documented options: "auto" (the default, zero, one, or multiple functions), "required" (one or more functions), a forced function call targeting exactly one named function, an allowed-tools list that restricts the model to a subset while leaving the full tool list in place (so prompt caching still sees a stable prefix), and "none", which imitates passing no functions at all. The subset variant is more useful than it sounds: same tools on the wire, smaller decision space per turn.

On parallel tool calls, the facts matter more than the vibes. On supported models beginning with GPT-5, functions can be called in parallel when built-in tools are also available, with the caveat that built-in tools cannot be included in a parallel function-call batch. When you want exactly zero or one call per turn, set parallel_tool_calls to false. Two documented footnotes belong in your design review. First, if you are using a fine-tuned model and it calls multiple functions in one turn, strict mode is disabled for those calls. Second, the gpt-4.1-nano-2025-04-14 snapshot can sometimes emit multiple tool calls for the same tool when parallel calls are enabled, and the docs recommend disabling the feature there. The general principle: parallelism buys throughput and costs determinism; any tool with side effects should either be idempotent or run with parallel calls disabled.

Round-trip results and error handling

The return path is specified more casually than the request path, and that is where teams improvise badly. The result you send back in the call-output message should typically be a string, and the format is yours to choose: JSON, error codes, plain text; the model interprets it. The one structural requirement is the reference back to the specific call, so parallel or sequential calls cannot be mixed up. For tools that return images or files, an array of the corresponding objects is accepted instead of a string.

Two practices turn this from a demo into an operation. First, treat tool failures as readable results: when the API you called behind the tool returns an error, feed a concise failure description back as the tool output instead of crashing the loop or silently retrying; the model can often route around a well-described failure, and the transcript shows the real story. Second, log the round trip like an engineer, not a spectator: tool name, call reference, duration, retry count, and the request ID; the field list for that logging, including the vendor request identifiers worth capturing, is covered in the logging guide on this site. And when the loop itself trips on HTTP error codes, the error code guide is the triage path; agent loops are exactly where 401s and 429s hide inside retries, which is why the guardrails in the CI guide apply to a web backend as much as to a pipeline.

Non-JSON outputs: custom tools and grammars

Not every tool wants JSON. OpenAI documents custom tools that accept and return free-form text, and for cases where the output must follow a specific shape without JSON wrapping, a grammar can be attached to a custom tool, with two supported syntaxes: Lark and regex. The grammar machinery is a context-free grammar, and the docs are direct about its costs: limit the grammar to the rules and patterns the tool actually needs, because a grammar that is too complex can cause an API error, and complex grammars tend to require iteration on the grammar, the prompt, and the tool description together. Two documented details worth internalizing: use a single bounded terminal for free text between anchors, because a lexer matches greedily and splitting free text across rules loses control; and remember that terminals and rules are distinct layers, with regex syntax following the Rust regex crate syntax rather than Python's re module. Grammars are powerful, and like most powerful things, their maintenance bill scales with complexity.

Token and cost accounting

Tool definitions are not free, and the docs say why: functions are injected into the system message, so callable function definitions count against the context limit and are billed as input tokens. The documented mitigations are practical: limit the number of functions loaded up front, shorten descriptions where possible, and use tool search so deferred tools load only when needed. This is the same cost discipline as the rest of your integration, and the attribution habits that make it manageable, tagging usage by caller and feature, are covered in the usage attribution guide. A tool list that grows monotonically is a silent cost increase; prune it like you prune dependencies.

Cross-vendor: the same loop, different semantics

Anthropic's tool use overview reads like a translation of the same concepts, with two differences worth planning for. On the request side, you pass a tool with an input_schema; on the response side, Claude returns one or more tool_use blocks naming the tool and its arguments, with a stop_reason of "tool_use"; your code executes, and a second request carries tool_result blocks that pair back through tool_use_id. The pairing identifier, not positional ordering, is the contract.

The second difference is execution location. Anthropic splits client tools, which include your custom tools plus Anthropic-defined schemas like bash and text_editor and run in your application, from server tools like web_search, web_fetch, code_execution, and tool_search, which run on Anthropic's infrastructure and return results you see directly. Server tools need no execution harness, but the docs note an interplay rule when a server tool shares a parallel batch with a client tool, so read the stop-reason documentation before mixing the two in one turn. For parallelism control, the same tool_choice block accepts a flag to disable parallel tool use, documented as asking for at most one tool call per turn.

The orchestration lesson across vendors: OpenAI's loop pairs results to calls by explicit call references, and Anthropic's pairs by tool_use_id; both are ID-based contracts, and neither tolerates a loop that assumes ordering. Write the orchestration layer against IDs, and the same logic ports between providers. If you reach models through one compatible endpoint, the shape checks in the compatibility checklist apply to tool calls as much as to text, and the endpoint context for your service lives in its documentation.

The production checklist

  1. Bring every schema into strict compliance: additionalProperties: false everywhere, every property in required.
  2. Verify the effective strict state per API surface, and alert when a fallback produces strict: false.
  3. Decide, per endpoint, how many tool calls one turn may contain; do not inherit the default by accident.
  4. Make side-effecting tools idempotent, or disable parallel calls for them.
  5. Return failures as readable results, and log tool name, call ID, timing, retries, and request ID.
  6. Keep tool definitions pruned, and use tool search deferral when the catalog grows.
  7. Write the loop against call IDs, not ordering, so provider swaps stay cheap.
  8. Re-run the tool-call regression on model upgrades, following the migration habits from the deprecation playbook when models rotate.

Tool calling rewards teams that treat it as a protocol with documented edges, not as magic. Read the two guides, respect the ID-based contracts, and the loop runs boring in the best way: unattended, observable, and boring.