Claude Code vs Codex CLI is a choice most developers are making from a leaderboard, and the leaderboard is out of date. On the September 2026 revision of Terminal-Bench 2.1, Codex CLI scored 79.1% and Claude Code 70.1%. Neither score tests the models these tools ship today, so picking by that table can waste your money. What follows walks through the models each agent actually defaults to this week, the benchmarks that disagree, the two very different approaches to sandboxing, what a subscription buys on each side, and where both vendors are pushing the agent next — so you can match one to your budget and your risk.
Two agents, one terminal
Claude Code comes from Anthropic. Its GitHub repository had roughly 148,700 stars and 25,100 forks by late September 2026. Codex CLI comes from OpenAI. It is written in Rust and licensed Apache-2.0, and sits at around 127,400 stars and 19,900 forks.
Both tools install where developers already work. Claude Code installs with a curl script, Homebrew or WinGet, and runs in the terminal, inside IDEs, or when you mention Claude on a GitHub issue or pull request. Codex CLI runs locally on macOS, Linux and Windows, plugs into VS Code, Cursor and Windsurf, and its exec command runs it without a human in a CI pipeline.
Reach, in other words, is close. Claude Code leads on stars. The differences that decide the question sit underneath the install instructions.
The models under the hood this week
On September 22, 2026, Anthropic made Claude Opus 5.5 the default Opus model in Claude Code, shipped in version 2.1.280. Anthropic says Opus 5.5 matches its larger Fable 5.1 model on most work while costing 40% less than Opus 5 — $4 per million input tokens and $20 per million output tokens, with a 1-million-token context window by default.
OpenAI answered at DevDay, on September 29 and 30. The Codex CLI docs now show GPT-6.1 Sol as the default model. Sol has a 1.05-million-token context window and up to 128,000 output tokens, and costs $2 per million input tokens and $10 per million output tokens — about a fifth of the price of OpenAI's flagship, GPT-6 Astra.
Side by side, the context windows are nearly the same. Sol is half the price of Opus 5.5 on both input and output. If you pay per token, that gap is the first number that matters.
Benchmarks that disagree
On the September 2026 revision of Terminal-Bench 2.1, Claude Code with Opus 4.6 jumped from 58.0% to 70.1% — the largest single-version gain of any agent and model pair tracked there. On the same leaderboard, Codex CLI with GPT-5.3-Codex scored 79.1%. Codex leads by 9 points.
Then look at the model names. Opus 4.6 and GPT-5.3-Codex are not what either tool ships by default this week. The leaderboard describes the last generation. It says nothing about Opus 5.5 or Sol.
A blind test by comparison blogger Blake Crosley paints a messier picture. He ran both tools across the same 12 task categories. Codex CLI caught an SSRF vulnerability, thanks to its sandboxed network restrictions, and Claude Code missed it. Claude Code, running Opus, caught a timing side-channel bug that Codex missed. Each tool had blind spots the other covered.
Developers argue about this too. On a Hacker News thread about a week of using Codex more than Claude, the top comment says Codex avoids word vomit and comments that turn into dead context noise. Two other commenters in the same thread say the opposite: that Codex writes overly complex solutions and Claude simplifies code better. No single score settles it.
Kernel wall versus hooks
Here is the design difference that matters most for safety. Codex CLI enforces its sandbox at the operating system level: Seatbelt on macOS, Landlock and seccomp on Linux. You pick one of three modes through the permissions command — read-only, workspace-write, or danger-full-access.
Claude Code takes another route. It relies on an application-layer hook system with roughly 31 lifecycle events, plus a sandboxed Bash tool. That trades a smaller kernel-level trust boundary for deeper programmability: you can script almost anything that happens around the agent.
Anthropic is betting on the model as well as the walls. It reports that Opus 5.5 shows 85% fewer boundary circumvention attempts in automated behavioral audits, and says the model resists prompt injection better than Opus 5. That is a claim about how the model behaves. A kernel sandbox is a rule the agent cannot talk its way around.
So the trade is legible. If you want hard walls, Codex has them built in. If you want to program every step of the agent's life, Claude Code gives you the hooks.
What a subscription buys
Pricing looks identical at the entry level. Claude Code comes with Claude Pro at $20 a month, or $17 a month billed annually, and shares one usage pool with Claude.ai chat. Codex CLI comes with ChatGPT Plus at $20 a month.
Above that, the ladders differ. Anthropic's Max plans run from $100 a month for 5x usage to $200 a month for 20x. OpenAI's Pro tiers cost $100, $200 or $500 a month, and the $500 tier adds Astra Ultrafast access. ChatGPT Plus and Business users face five-hour local message caps that vary by model — for example from 350 to 3,000 messages. Pro subscribers have no five-hour limit on Codex usage.
For teams, Anthropic's Team Standard costs $20 per seat a month billed annually, or $25 monthly; Team Premium costs $100 per seat a month billed annually, or $125 monthly. Both include Claude Code. OpenAI's Business plans start at $20 per user a month, billed annually. At the team entry point, the two are level.
The race to the cloud
Both companies are moving the agent off your laptop. At DevDay on September 29, OpenAI gave Codex reusable, persistent cloud environments that follow you across laptop, phone and browser. It also added voice control for starting and directing CLI tasks, and Codex Security Cloud, which scans GitHub repositories and prepares fixes. Access is rolling out to Plus and Pro subscribers, and to Business, Enterprise, Edu and Healthcare workspaces.
Anthropic shipped its piece a week earlier. Claude Code version 2.1.280 added native Slack integration for cloud sessions, with a working indicator, a Stop button and thread titles. It also fixed the Remote Control diff view, so you can review cloud-run changes from another device.
Anthropic's own engineers show where this is heading. In a video published September 2, 2026, which passed 310,000 views by month's end, Claude Code engineers say they now do most of their coding through cloud sessions, giving Claude goals rather than step-by-step instructions. A year earlier, every action needed a manual permission prompt.
The tools are also close enough that switching feels cheap. Hacker News commenters argue you can swap Codex for Claude Code pretty easily — their opinion, not a measurement. But matching entry prices do make it easy to test both.
Claude Code vs Codex CLI: which agent fits you
Here is the verdict. If you pay per token, Codex CLI with GPT-6.1 Sol is the cheaper engine, at half of Opus 5.5's price. If you want a sandbox enforced by the kernel, Codex has it by design. That makes it the pick for cost and for hard safety walls.
If you want to script the agent's whole lifecycle, Claude Code's roughly 31 hook events are the deeper toolkit. Anthropic says Opus 5.5 matches its larger Fable model on most work at 40% less than Opus 5 — Anthropic's claim, not an independent test.
On benchmarks, Codex leads Terminal-Bench 2.1 by 9 points, but with last-generation models. The blind test showed each tool catching bugs the other missed, so for security review, running both is not paranoia.
On a $20 plan, the price is the same. The difference is the limits: ChatGPT's five-hour caps on one side, and on the other a pool Claude Code shares with chat. Pick the wall you trust, then pick the plan.
Sources
- anthropics/claude-code
- openai/codex
- Introducing Claude Opus 5.5
- OpenAI gives Codex reusable cloud environments that work across devices
- Claude code is falling behind Codex not because of token cost, but because of Opus 5.
- Claude Code Agent Allegedly Deletes 48,000 Files in 103 Seconds
- Terminal-Bench 2.1
- A week of using Codex more than Claude
- Release v2.1.280 · anthropics/claude-code
- Codex CLI docs
- How the Claude Code team uses Claude Code
- Meet the all new Codex Cloud
- Tell HN: Codex Is Down [fixed]
- Codex CLI vs Claude Code 2026: Architecture, Pricing, and China Access
- OpenAI Launches GPT-6.1 Sol At DevDay
- Claude Pricing
- ChatGPT Pricing
- Engineer says Claude Code has made his job "soul-sucking" as workers spend 12-hour days pressing enter