Two launches, one tie
On 22 September Anthropic released Claude Opus 5.5 with a table showing it at 66.4% on Terminal-Bench 4.0, a test of agents doing real work in a terminal, against 57.9% for OpenAI's GPT-6 Astra. The same day Artificial Analysis, which runs the benchmark itself, published its own figure for Opus 5.5: 59.6%, level with Astra.
Nobody's number was false. Anthropic ran the test in its own setup, and its footnotes say it used the xhigh effort setting for Opus 5.5 and high for Astra, although OpenAI's model page lists two settings above high. Artificial Analysis ran both models at xhigh in its own test setup. Change who runs the test and how, and a clear win becomes a tie.
That is the right way into Claude Code against Codex this autumn. The two products have grown into near copies of each other: the same $20, $100 and $200 price ladder, the same spread across terminal, editor, desktop app and cloud, the same parallel sessions and support for the Model Context Protocol. What separates them now is smaller and more practical: which model each hands you by default, how far apart its cheap and expensive options sit, how much it tells you about your limits, and how it treats the machine it runs on. On those terms Claude Code gives you the stronger default model. Codex gives you a wider range of prices, limits you can read in advance, an open-source CLI and better footing on Windows.
The products have converged
The old shorthand, Claude Code in your terminal and Codex in OpenAI's cloud, no longer holds. Anthropic's Claude Code documentation lists a terminal CLI, extensions for VS Code and JetBrains, a desktop app and a web version whose cloud sessions keep running after you close the laptop. OpenAI's Codex documentation lists a CLI, an IDE extension, the desktop app and Codex cloud, which runs tasks in parallel and hands back each one as it reaches a reviewable result.
The plans converged too. Both companies fold the coding agent into the general subscription rather than selling it separately, and the tiers line up almost exactly:
| Monthly price | Anthropic (Claude Code) | OpenAI (Codex) |
|---|---|---|
| $0 | Free: no Claude Code | Free: GPT-6 Luna in the desktop app, subject to rollout |
| $8 | No equivalent | Go: same Codex access as Free |
| $20 | Pro | Plus |
| $100 | Max 5x (five times Pro) | Pro 5x (five times Plus) |
| $200 | Max 20x (twenty times Pro) | Pro 20x, closed to new sign-ups since 10 September |
| Per user, teams | Team standard seat, $20 billed annually or $25 monthly | Business, $20 billed annually or $25 monthly |
Sources: Claude pricing, Anthropic's Max plan help page and OpenAI's Codex pricing page.
The one asymmetry is at the top. On 10 September OpenAI stopped taking new sign-ups and upgrades for its $200 tier, citing demand for Astra, TechCrunch reported. Existing subscribers keep it and the $100 tier is unaffected, but if you are not already on OpenAI's top plan, you cannot buy it this month. Anthropic's Max 20x is on sale.
The default model is the biggest difference
Set the two model lineups side by side and they match exactly at the top, then pull apart below it.
| Tier | Anthropic, per million tokens in / out | OpenAI, per million tokens in / out |
|---|---|---|
| Top | Claude Fable 5.1: $10 / $50 | GPT-6 Astra: $10 / $50 |
| Workhorse | Claude Opus 5.5: $4 / $20 | GPT-6 Sol: $2 / $10 |
| Mid | Claude Sonnet 5: $2 / $10 | No separate tier |
| Small | Claude Haiku 4.5: $1 / $5 | GPT-6 Luna: $0.10 / $0.50 |
| Context window | 1M tokens (Haiku 4.5: 200K) | 1.05M tokens |
Figures from Anthropic's models overview and OpenAI's API pricing and model pages. OpenAI charges more for prompts beyond 272K input tokens.
What matters day to day is which row each tool starts you on. Claude Code's model configuration page says the default now resolves to Opus 5.5 on Pro, Max, Team, Enterprise and the API; before version 2.1.280, Pro users defaulted to Sonnet 5. Opus 5.5 runs at medium effort unless you change it. Codex starts at a preset its models page calls Sol Light: GPT-6 Sol at its lightest reasoning setting.
So an untouched Claude Code session runs a model that costs twice as much per token as the one an untouched Codex session runs, and the independent numbers suggest you get something for the money. Artificial Analysis measured Opus 5.5 at 59.6% on Terminal-Bench 4.0 and GPT-6 Sol at 43% to 44%, both at maximum effort. A Codex user who wants Opus-class results picks Astra, which costs two and a half times as much as Opus 5.5 per token and, on Plus, comes with 5 to 45 messages per five-hour window.
Both companies fence off their most expensive model. Anthropic's help center says Fable 5.1 is outside Pro's usage limits and available there only through paid usage credits; on Max, you can spend up to half your weekly allowance on Fable models. OpenAI includes Astra on Plus but gives it the smallest allowance of any model.
At the bottom the gap runs the other way. GPT-6 Luna costs a tenth of Claude Haiku 4.5 per token, and a Plus subscriber gets 350 to 3,000 Luna messages per five hours. Luna is not a strong coder on its own: Artificial Analysis scores it 41 on its Coding Agent Index, against 57 for Sol. It is the model for renames, lint fixes and other edits that need speed more than judgment. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow Opus 5.5 "in the coming weeks". Until they ship, Codex owns the cheap end.
What the limits look like
OpenAI publishes its Codex limits as numbers. The pricing page gives ranges per model for each five-hour window: on Plus, 15 to 150 messages with Sol, 5 to 45 with Astra and 350 to 3,000 with Luna, with Pro 5x and Pro 20x scaling those up. The ranges are wide because a message that reads half a repository and reasons for minutes costs far more than one that fixes a typo, but they give a buyer something to plan against.
Anthropic publishes multiples instead. Max 5x gives five times Pro's per-session usage and Max 20x twenty times, and Claude Code draws on the same limits as everything else you do in Claude. There is no published message count for Pro, so a week of real use is the only way to find out whether it is enough.
Both companies sweetened launch day the same way. Anthropic raised five-hour limits on Pro, Max, Team and seat-based Enterprise plans and added a rate-limit reset you can save and spend when you choose. OpenAI's announcement gave Plus, Pro and Business users a "banked reset".
Cloud work is not extra capacity on either side. Claude Code's cloud documentation says cloud sessions share rate limits with the rest of your account, parallel sessions use them up proportionately, and there is no separate charge for the VM. OpenAI's pricing page says local messages and cloud chats share the plan's allowance. If you prefer to pay per token, both run on API billing: Claude Code with an Anthropic Console account, Codex with an API key at standard rates.
How each one treats your machine
An agent that asks permission for every shell command is tiring to use; one that asks for none is a risk. Both tools resolve this the same way, with an operating-system sandbox that fences off files and network so the agent can run routine commands unattended. The implementations differ where it counts for some teams.
Claude Code's sandbox uses macOS's built-in Seatbelt framework and, on Linux and WSL2, two extra packages. Its documentation says native Windows is not supported: Claude Code itself runs on Windows, but to get the sandbox you run it inside WSL2. Codex's sandbox uses Seatbelt on macOS, bubblewrap on Linux and WSL2, and a native Windows sandbox in PowerShell. Its defaults let the agent edit files inside the project and run routine commands, and make it stop and ask before going further, including onto the network. For a Windows shop that does not use WSL, that is a practical reason to prefer Codex.
In the cloud, each Claude Code session runs in an isolated virtual machine that Anthropic manages, with network access limited by default, git credentials kept outside the sandbox, and GitHub required for cloning and opening pull requests. Codex cloud runs tasks in isolated environments you configure per repository, and you can start them from GitHub pull requests, GitLab merge requests and issues, Linear issues or Slack threads. Claude Code covers similar ground through Slack, GitHub Actions and GitLab CI/CD, routines that fire on a schedule or on GitHub events, and an auto-fix mode that watches a pull request and pushes fixes for failed checks and review comments.
For parallel work on one machine, Claude Code's background agents move each session into its own git worktree before editing, so several can read one checkout without overwriting each other. Codex supports worktrees too, and its Ultra setting splits a single task among subagents.
Claude Code has the deeper kit for packaging repeatable work: CLAUDE.md for standing project instructions, skills for workflows a team shares, and hooks that run shell commands before or after the agent acts. Codex reads AGENTS.md and supports MCP servers, as Claude Code does. Claude Code will also read an existing AGENTS.md, so one instruction file can serve both agents in the same repository.
One difference is structural. The Codex CLI is open source, developed in public on GitHub. A team that wants to read, audit or patch the agent running on its developers' machines can do that with Codex.
Reading the launch tables
Go back to the table that opened this piece. OpenAI's own Astra announcement on 3 September reported 57.9% on Terminal-Bench 4.0 against 55.8% for Claude Fable 5.1. Three weeks later Anthropic's table had Opus 5.5 at 66.4%, with Astra run at high. Artificial Analysis, running both at xhigh, found them level at 59.6%. Note that Opus 5.5 at the same xhigh setting scored almost seven points lower in the independent run than in Anthropic's, so the test setup matters as much as the effort dial. Each vendor's figure is a real result under conditions the vendor chose; the independent run is the only one of the three that held the conditions constant for both models.
Two more reasons to hold launch tables loosely. First, benchmark runs use high or maximum effort, and your tool probably does not: Claude Code runs Opus 5.5 at medium by default and Codex starts at Sol Light. Second, cost per task depends on how many steps an agent takes as well as its price per token. Artificial Analysis measured GPT-6 Sol at $2.99 per task on its Coding Agent Index, at maximum effort running inside Codex itself; it has not published a comparable Claude Code figure for Opus 5.5.
Anthropic's most eye-catching claim, that one tester finished a 680,000-line code migration in under a day, is a single unnamed example. It may be true. It is not something to plan a quarter around.
The comparison that settles it for a team is its own. Take five tasks from your backlog, run each through both tools at their default settings for a week, and count what merged, how long review took, and whether anyone hit a limit.
Which fits which work
For a long, tightly coupled change in one codebase, where step nine depends on what happened at step three (a framework migration, a refactor across dozens of files), start with Claude Code. Its default model is the stronger of the two defaults, and background agents in separate worktrees let you split off side tasks without disturbing the main job.
For a steady flow of small, well-specified tickets, or mechanical changes across many files, Codex is the better value. Luna and Sol are cheap enough to throw at volume, and the published limits tell you in advance how far a plan will stretch. The same holds if your developers work on native Windows or your security team wants to read the agent's source.
If your team repeats the same jobs every week (a release checklist, a dependency audit, a report built from the same exports), Claude Code's skills, hooks and routines make those repeatable in a way Codex does not yet match.
Running both is common and cheap to set up. Put shared instructions in AGENTS.md, give each agent its own branch or worktree, and review everything as a pull request.
Search and ads teams face the same split. A coding agent is the right tool for site-wide template, markup and redirect changes; the numbers that justify those changes should come from code reading Search Console and Google Ads exports, with a person approving the result. That division is how AIO Copilot runs SEO and Ads work for businesses and agencies. It is priced per site; tell us how many you manage and we email the price.
Both tools now start at $20 and both $20 tiers include a model near the top of the independent tables. A week on each at the lowest paid tier will tell you more than any launch table, and with one AGENTS.md file serving both, the second trial costs little beyond the subscription. Buy the $100 or $200 tier only after you have hit the $20 limits doing real work.
Frequently Asked Questions
Is Claude Code or Codex better in September 2026?
Neither wins everywhere. Claude Code starts you on Claude Opus 5.5, which independent tests by Artificial Analysis place well ahead of GPT-6 Sol, the default in Codex. Codex offers a wider price range, published usage limits, an open-source CLI and a native Windows sandbox. Pick by the work you do most, then test both on your own repository.
How much does Claude Code cost?
Claude Code is included in Claude Pro at $20 a month, Max 5x at $100 and Max 20x at $200, and in Team and Enterprise seats; the Free plan does not include it. Through the API, the default model, Claude Opus 5.5, costs $4 per million input tokens and $20 per million output tokens.
How much does Codex cost?
Codex comes with OpenAI's paid plans rather than as a separate subscription: Plus at $20 a month, Pro at $100 (5x) or $200 (20x), and Business at $20 per user a month billed annually. New sign-ups for the $200 tier have been paused since 10 September 2026. Through the API, GPT-6 Sol costs $2 and $10 per million input and output tokens, Luna $0.10 and $0.50, and Astra $10 and $50.
What is the default model in Claude Code and in Codex?
Claude Code defaults to Claude Opus 5.5 on Pro, Max, Team, Enterprise and API accounts, at medium effort. Codex starts at a preset its documentation calls Sol Light, which is GPT-6 Sol at its lightest reasoning setting. Both let you switch models and effort per session.
Can I use Claude Code and Codex in the same repository?
Yes. Codex reads an AGENTS.md instruction file and Claude Code can read AGENTS.md as well as its own CLAUDE.md, so one set of project instructions can serve both. Give each agent its own branch or git worktree so their edits do not collide.
Never miss an update
Get the latest AI and SEO strategies delivered to your inbox.