Putting an AI coding agent pipeline together is not hard to start and easy to get wrong, because the difficulty is not in making the tool run. It is in three things: invoking it non-interactively, getting credentials into the CI environment without leaking them, and defining permission boundaries that hold when nobody is watching. This guide covers those three, plus the accountability and triage habits that keep a pipeline healthy, and it deliberately stays on the unattended case; the local interactive setup is its own guide, and this article does not repeat it.
Headless Claude Code and codex exec: the two entry points
Claude Code's non-interactive entry is the -p flag, also written --print, and the same entry point powers Claude Code CI/CD usage and one-off scripted runs alike. Per the provider's documentation on running it programmatically, adding -p to a command runs it without the interactive UI, and the common companions are --continue for continuing conversations, --allowedTools for pre-approving tools, and --output-format for structured output. The exit code is meaningful by design: success exits zero, failures exit non-zero, so a pipeline can branch on the result instead of scraping text. Structured output formats include plain text and JSON, with streaming variants for progressive consumers; the exact list is maintained on the reference pages, and it is worth checking there, because both CLIs evolve quickly and this article describes the shape, not the day's exact flag table.
One Claude Code feature deserves special attention in CI: bare mode. The documentation notes that bare mode never reads OAuth credentials or the system keychain, and expects ANTHROPIC_API_KEY in the environment instead, which is exactly the shape a pipeline wants. A minimal unattended call can pre-approve a single tool so the run cannot stall on a prompt, for example asking for a summary with only the read tool approved.
Codex CLI's non-interactive entry is codex exec. Its non-interactive mode documentation describes it as the way to run Codex from scripts and CI jobs without the interactive TUI, and the mechanics are pipeline-friendly: progress streams to stderr while only the final agent message goes to stdout, so redirecting or piping the result is straightforward. Flags such as --json and --ephemeral shape the output and avoid persisting session files on the runner. The same page's guidance for unattended runs is explicit: use pre-set sandbox and approval settings rather than inheriting interactive behavior.
Credentials in CI: secrets injection and rotation
The first rule has no exceptions worth the risk: the API key enters through the CI secret store and is exposed to the single step that needs it. It does not belong in the repository, in committed config files, in build artifacts, or in log output, and a pipeline that echoes its environment for debugging is a credential leak waiting for an audit.
The most common CI failure has a boring explanation. If a command works on a laptop and returns 401 in CI, the request never had the credential, or it had a different one: the secret was not injected into that job, the variable name does not match what the tool reads, or the secret is scoped to a different environment. Treat "works here, fails there" as an injection problem until proven otherwise.
Each tool wants credentials in its own documented way. For Claude Code in bare mode, that is an environment variable, as described above. For Codex, the non-interactive documentation covers the options directly: codex exec reuses saved CLI authentication by default, which is rarely what a fresh runner has; for CI it describes providing credentials explicitly, including an API-key approach that can be scoped to a single run through CODEX_API_KEY set inline rather than as a job-wide variable, and it points to dedicated automation tooling for hosted pipelines. The docs are also candid about the sharp edges, including the warning that token-based login workflows are not appropriate for public repositories. Read the current page before choosing your method, because this is the surface where stale blog posts do the most damage.
Rotation follows one order, and getting it backwards causes the outage it was meant to prevent: add the new credential, verify it in the pipeline, and only then revoke the old one. The full change checklist, including how configuration surfaces interact when both secrets briefly exist, is in the provider switching guide.
Permission boundaries for an AI coding agent pipeline
In an interactive session, the agent asks before doing something surprising. In a pipeline, there is no one to ask, which converts every permission question into a configuration decision you make in advance. This is the part of the setup that deserves the most care, and neither provider offers a way to skip it: the documentation's guidance is to define sandbox and approval behavior explicitly for unattended runs.
The building blocks are consistent across both tools. Claude Code's --allowedTools pre-approves a specific set of tools, and its permission surface distinguishes allow rules from deny rules, so a scoped rule can permit reading while denying a class of shell commands. Codex documents sandbox modes on a spectrum: read-only, where the agent can inspect but not modify without approval; workspace-write, where edits and routine commands stay inside the workspace boundary; and full access, which removes the filesystem and network boundaries entirely. The middle mode is the default for local work; the last one is the one to enable only deliberately, and only where you accept what it means.
Three principles carry across both tools and every pipeline. First, grant the smallest writable scope the job can function with: if the task is a review, the job needs reading, not writing, and if it needs writing, it rarely needs the whole machine. Second, prefer read-only and reasoned defaults over convenience flags; a run that fails because it was not allowed to do something is a report, while a run that was allowed to do everything is an incident waiting for a trigger. Third, treat dangerous capabilities as explicit opt-ins that appear in the pipeline file and therefore in code review, not as ambient defaults. Nothing here makes an unattended agent safe by itself; it makes the risky parts visible and reviewable, which is the achievable goal.
Keeping the pipeline's usage visible and bounded
An unattended pipeline is also an unattended bill, and the two habits that keep it boring are attribution and reuse. Tag every job and repository so each run can be told apart when usage is reviewed; the labeling scheme and why it matters are covered in the usage attribution guide. And remember that CI is a repeat machine: a retry loop on a failing job is also a loop of new calls, and long, repeated contexts in pipeline runs are exactly the profile that prompt caching helps with, as covered in the caching guide. Neither habit is exotic; both are cheap to add at the start and annoying to retrofit.
CI-specific triage, in order
When a pipeline misbehaves, the failures cluster into a predictable sequence. Work through it in order, because later layers are invisible until earlier ones are ruled out.
- 401, unauthorized. The credential layer, and almost always injection: no secret in that job, wrong variable name, or a secret that was rotated while the pipeline kept the old copy. Verify the injection before touching anything else.
- 404, not found. The endpoint layer: the base URL is wrong, typically missing the path suffix the endpoint expects. If your endpoint uses the
/v1convention, confirm the exact value once in a terminal rather than squinting at YAML. - Model not found. The model layer: the model ID in the pipeline does not exist on the endpoint. Check the live model list rather than trusting an ID that was correct last quarter.
- Permission denied, non-interactively. The guardrail layer: the run hit a boundary you set, with no interactive prompt to approve the exception. This is the design working as intended; widen the specific permission, do not remove the boundary.
The same symptoms outside CI have a full triage guide of their own, the error code reference, and the endpoint details for whatever platform you run against are documented in the platform's own docs.
Where to go next
The reading path this article sits in: local interactive setup first, then key rotation, then this unattended case, then the cost instruments and the error reference above. If you are wiring the agent into a specific editor or tool on a developer machine next, the configuration surfaces involved are the same ones the provider switching walkthrough maps out.
Make the pipeline boring on purpose: narrow permissions, injected secrets, tagged usage, and a triage order everyone follows. Unattended should describe the mode, not the outcome.