Codex vs Claude Code: the same three bugs, on one machine
Both agents fixed the same three real open-source bugs on one machine, judged by the maintainers' tests: results, time, tokens, and what else decides it.
On October 9, 2026 we gave both agents the same three real bugs from open-source projects, fixed upstream in the ten days before, on the same machine, and judged the work by each project’s own tests. Both fixed all three. Claude Code was faster on the two small bugs and slower on the hard one, where it went further than asked: it found the same bug in a second place and put its fix behind a preview flag, citing the project’s stability rules. Codex’s fixes were the more conservative, one of them the maintainers’ own, and its sandbox got in the way of running two of the three test suites. On this evidence the agents are close; what decides between them is mostly the plan you already pay for, the machine you work on, and how you like to work.
The practical differences: Claude Code comes with paid Claude plans from $20 a month, or runs on a Console API key billed per token, with no free plan, while OpenAI says every ChatGPT plan includes Codex, with Plus at $20 the cheapest that both of its pages put in the terminal; Codex’s CLI is open source and Claude Code isn’t; Claude Code’s Remote Control lets your phone drive a session on any machine, a Linux server included, while Codex’s phone connection goes through its desktop app on a Mac or Windows PC. And you don’t have to choose: both run side by side on one machine.
The run
Each task was a fix merged between September 29 and October 6, 2026, so neither model could have learned it in training: an issue in more-itertools, a bug in Commander.js reported in its fix’s own words (it had no issue), and an issue in Black, the Python formatter. Each project was checked out at the commit just before the fix, with its history removed so the fix couldn’t be looked up, its dependencies installed, and one copy for each agent. Each agent got the same prompt, the bug report and a request to fix it, add tests and keep the suite passing, and worked alone with no network: Claude Code 2.1.295 with Opus 5.5 at xhigh effort, non-interactive, edits allowed and its sandbox on; Codex 0.160.0 with GPT-6.1 Sol at xhigh, in its workspace-write sandbox with approvals off. All six runs went at once on that machine. Afterwards we ran the maintainers’ own test from the fix, which neither agent saw, and the full suite, outside both sandboxes.
| Task | Claude Code: time, tokens | Codex: time, tokens |
|---|---|---|
| more-itertools | 71 s 326k in, 5k out | 99 s 218k in, 3k out |
| Commander.js | 105 s 386k in, 10k out | 359 s 779k in, 11k out |
| Black | 964 s 5.4M in, 51k out | 530 s 3.0M in, 13k out |
Both agents’ fixes passed the maintainers’ test and the full suite on all three tasks: more-itertools’ ichunked(…, 0) raising on a non-empty input (a small Python bug), a Commander.js child process reusing its parent’s IPv6 inspector port (a small JavaScript bug), and Black’s --preview leaving a line over the length limit (a hard bug in a code formatter). The tokens are each tool’s own count of what the model read (mostly the same context re-read from cache) and wrote, and the two count differently, so compare them within a column more safely than across it. Both runs were on subscriptions, so their cost was a share of each plan’s allowance rather than a bill.
What each one did
The small Python bug. Both wrote the same two guard lines the maintainers did, and both added tests; Claude Code also wrote the changelog entry the maintainers wrote, and Codex documented the behaviour in the function’s docstring instead. Nothing to choose between them.
The JavaScript bug. Codex’s fix is the maintainers’ fix: the same pattern for a host in square brackets, tested across the four inspector flags. Claude Code first ran the failing command against Node’s real inspector, found a third form that failed too (an IPv6 address without brackets, which Node accepts), and split host from port at the last colon, as Node itself does, so its fix covers more than the maintainers’. Part of Codex’s extra four minutes went on its sandbox: the project’s tests start child processes, which lost their output inside it, so Codex wrote a temporary adapter outside the project to get all the tests through.
The hard bug. Both found the same cause in Black’s line splitter, a fallback that is skipped as soon as any hidden parenthesis on the line has been used, which the new dictionary formatting trips, and then wrote three different fixes, counting the maintainers’. Codex changed one condition and added a test file, and, as the maintainers did, corrected an existing test that had the bug recorded as the expected output; its sandbox refused to open local sockets, so Black’s daemon tests failed and its async tests hung there, and Codex said plainly that it could not finish the full run. Claude Code checked whether the bug reached Black’s stable style, found it did, with an await expression, and so put its fix behind a new preview flag, saying Black’s stability policy keeps a change to stable output out of a fix like this; it showed its test failing without the fix, reformatted Black’s own source and test files at three line lengths before and after to show nothing stable changed, and ran the full suite in its sandbox, where two caching tests failed the same way on the untouched checkout. That took sixteen minutes to Codex’s nine, and seven files where the maintainers changed three.
What three bugs can’t tell you
This is three tasks with one attempt each, on one day’s versions, at one effort setting. Both makers ship new models and releases every few weeks, and we ran both at xhigh, where Claude Code’s documented default effort for Opus 5.5 is medium, and OpenAI documents none for Codex. What it does show is each agent’s manner on real work, and that on bugs of this size the difference between them was in how far they went and how their sandboxes behaved, not in whether they got there.
The differences that decide it
| Claude Code | Codex | |
|---|---|---|
| Cheapest plan | Claude Pro, $20 a month ($17 billed annually), or a Console API key billed per token; the free plan has no Claude Code | OpenAI’s help article says every ChatGPT plan, Free and Go ($8) included; its pricing page gives those two the desktop app only; Plus ($20) is in the CLI on both |
| Plans above | Max at $100 or $200, monthly | Pro at $100, $200 or $500 |
| Models | Claude only; Opus 5.5 by default on every paid plan, Fable included on Max | OpenAI’s, GPT-6.1 Sol recommended; local open models with --oss |
| Open source | No: “All rights reserved” | The CLI, under Apache 2.0; the IDE extension and the cloud aren’t |
| Instruction file | CLAUDE.md; reads AGENTS.md when there’s no CLAUDE.md (its docs) | AGENTS.md; reads CLAUDE.md only if you add it as a fallback name (its docs) |
| Out of the box | Auto mode: a second model reviews actions; the OS sandbox is off until you turn it on | An OS sandbox: writes in the project only, no network, asks before crossing |
| From your phone | Remote Control, from the terminal on any machine | Through the ChatGPT desktop app on a Mac or Windows PC (remote connections) |
| PR review on GitHub | Team and Enterprise only, billed per review | Plus and above |
| Training, personal plans | Unless you turn the account’s training setting off | Unless you turn off “Improve the model for everyone,” which OpenAI’s data controls FAQ says covers Codex tasks on a personal plan |
The rest they share: a terminal CLI, editor extensions (VS Code and its forks, and JetBrains through its own AI chat for Codex), a desktop app, a hosted cloud version, a phone app, MCP servers, skills in the same open format, plugins, hooks, subagents, a scripted mode with JSON output, SDKs, and an allowance shared with the maker’s chat app, with a five-hour window on the entry plans and weekly limits on top. The plans themselves are Claude’s and ChatGPT’s guides, and the two clouds are Claude Code on the web and Codex cloud.
Using both
The two are built to share a repository. Write your instructions once in AGENTS.md, which Codex reads and Claude Code reads when there’s no CLAUDE.md, or keep a CLAUDE.md that imports it with @AGENTS.md. Keep skills in .agents/skills/ with a link from .claude/skills/, which the skills guide tested on both. And each can bring the other’s setup across: Claude Code’s /import codex and Codex’s /import copy instructions, MCP servers and skills across, and Codex’s brings Claude Code’s settings too.
Both need a computer that stays on for work that outlasts your laptop’s lid. Signed in with your own plans, both agents run on one always-on machine, each in its own terminal session. Everpod’s developer pod is that machine ready made: a cloud computer that is yours, always on, with your pick of Claude Code, Codex, OpenCode and Pi installed, reached only through your own private network, from $24 a month. A developer pod is one machine that stays yours: what you installed, the servers you left running, Docker and your work in progress are where you left them tomorrow.