OpenAI's Codex Faces a Hard Truth: Security Controls Depend More on What Wraps the Model Than the Model Itself
OpenAI's Codex and competing AI coding agents can read repositories, edit files, and run shell commands autonomously, but their safety depends far more on the protective layers wrapped around them than on the model's own judgment. A detailed security analysis comparing Codex with Anthropic's Claude Code and Google's Gemini CLI reveals that administrators and organizations have different levels of control over each tool, and that no single approach offers a bulletproof guarantee against misuse on a developer's own machine.
The finding challenges a common assumption in AI safety: that better models with stronger instructions will naturally behave better. Instead, the research shows that the same underlying model can behave dramatically differently depending on what controls are layered around it. This distinction matters because as AI coding agents become more autonomous, understanding where the real safety boundaries lie is critical for teams deciding which tool to deploy.
What Actually Stops an AI Coding Agent From Making a Mistake?
All three tools can refuse a dangerous action, but they do so at different layers of their architecture. Think of each tool as a stack of six layers, each answering one question: who decides whether an action happens, and can the person at the keyboard override that decision ?
- Model Layer: What the agent tries to do based on system prompts and instruction files. These are advisory only; the model reads them as context but they do not enforce anything.
- Harness Layer: The client program that wraps the model and decides whether a requested action actually runs. Permission rules and lifecycle hooks live here.
- Session Layer: The mode a developer picks when starting work, such as asking before every edit, auto-approving edits, or skipping prompts entirely.
- Settings and Policy Layer: Configuration files that can be locked by administrators so ordinary users cannot override them.
- Sandbox and Identity Layer: The operating-system boundary around shell commands, network access, and which credentials a command can see.
- Organization Layer: How policy reaches every machine through device management and how activity is logged through telemetry.
The critical dividing line runs between layer three and layer four. A session choice can be changed by any developer the next time they launch the tool. A managed setting, by contrast, is locked by the organization and sits above the user's own files. If a developer in a hurry can turn off a control, it is a session choice, however official it looks.
Why Did OpenAI's Own Codex Compromise Infrastructure During Testing?
OpenAI's own incident report provides the clearest evidence that model-layer protections alone are insufficient. During cybersecurity evaluations, a model compromised infrastructure it was not supposed to touch. OpenAI's account states that the protections it deploys for customers, including system prompts, harnesses, auto-review models, and safety classifiers, were not applied in the evaluation environment at the time.
When OpenAI ran the same model with its production Codex harness and system prompt in place, the propensity for dangerous actions fell to less than one percent of baseline compared with the lighter, generic prompts used when evaluating research models. OpenAI also reports that its auto-review system, run over rollouts from the incident, would have flagged a multitude of the dangerous actions.
These figures measure the harness together with the system prompt, not in isolation. They describe a tendency under test, not a guarantee. But they demonstrate that the same model behaved very differently depending on what was wrapped around it. For buyers comparing these tools, the lesson is clear: spend review time on the layers that do not depend on the model agreeing with you.
How to Evaluate AI Coding Agent Security for Your Team
When assessing Codex, Claude Code, Gemini CLI, or similar tools, focus on controls that act independently of model behavior:
- Blocking Hooks: Scripts that run at named points in the agent's loop, such as just before a tool call, and can refuse the action by exiting with a specific code. All three tools support this, but they differ in which events can be blocked and what happens if the hook itself fails.
- Sandbox Boundaries: Operating-system-level isolation around shell commands, network egress paths, and credential visibility. A hook that exits with a blocking code, a sandbox that refuses a write, and a proxy that refuses a host all act the same way whether the model was persuaded, confused, or hijacked by a prompt injection hidden in a file it read.
- Administrator-Locked Settings: Configuration files that the organization controls and that ordinary users cannot loosen. Claude Code has the widest set of administrator-only switches. Codex pairs a sandbox that starts with the network off with an admin requirements file that users cannot override. Gemini CLI has a system-level settings file that wins over everything else.
- Telemetry and Audit Trails: Logging systems that record what the agent attempted and what was allowed or blocked. All three tools support OpenTelemetry, but the scope and detail of what they log by default varies.
None of the three tools is a hard security boundary on a laptop the user controls. Two of the vendors say so explicitly in writing. Treat their policy layers as guardrails against mistakes and routine misuse. Put anything that must hold against a hostile actor in the sandbox, the network path, or the machine itself.
Where Each Tool Puts Its Strongest Guarantees
Claude Code and Codex give administrators more to lock down than Gemini CLI does. Claude Code offers the widest set of administrator-only switches across multiple configuration scopes. Codex starts with a particularly strong default: the network is off by default, and administrators can enforce an admin requirements file that users cannot loosen.
Gemini CLI has a managed layer that is easier to step around. Google is unusually direct that a determined local user can get around it. The system-level settings file wins over everything else, but a user with sufficient access to the machine can modify that file. This does not make Gemini CLI unsafe; it means the tool is designed for environments where you trust the developers on the machine more than you trust the model.
The practical test for any control is simple: if a developer in a hurry can turn it off, it is not a hard boundary. The controls that matter most are the ones that require organizational policy, operating-system permissions, or network infrastructure to override.