Claude Code subagents: make one, what it costs, and Codex's
What a subagent is, the file that defines one, the built-in ones, when Claude delegates and when Codex does, and three runs of the same search with and without one.
A Claude Code subagent is a separate worker that Claude starts for a side task: it gets its own context window, its own instructions, the tools and model you choose for it, and hands back only its result, so the searching and reading it did never fill your main conversation. Make one by saving a Markdown file in .claude/agents/ (this project) or ~/.claude/agents/ (every project) with a name, a description of when to use it, and its instructions as the body. Claude hands work to a subagent on its own when a task matches the description, or when you name it. Its requests count against the same usage limits as the rest of your session, so the saving comes from giving it a cheaper model, not from delegating as such.
Codex has subagents too, on by default, defined as TOML files in .codex/agents/, with one large difference: it starts them when you, an AGENTS.md file or a skill asks for parallel agents, and on its own only at its highest reasoning setting.
What a subagent is, and what it isn’t
Anthropic’s subagents page gives the reason to use one: “when a side task would flood your main conversation with search results, logs, or file contents you won’t reference again: the subagent does that work in its own context and returns only the summary.” Underneath it is one tool, Agent (called Task before version 2.1.63): the main session writes a task, the subagent works on it, and “the parent doesn’t see the subagent’s intermediate tool calls or outputs, only that final result.”
A subagent starts fresh. It doesn’t see your conversation, the files Claude has already read or the skills it has already used; it gets its own instructions, your CLAUDE.md files, a snapshot of the git status and the task message Claude writes for it. So a rule it must follow, such as “ignore the vendor folder”, has to be in what you ask Claude to delegate. The exception is a fork, started with /subtask, which inherits the whole conversation and still returns only its result.
Its neighbours, in Anthropic’s terms: a skill is reusable instructions loaded into whatever context uses it, while a subagent is an isolated worker, and the two combine (a subagent can preload skills). Agent teams are separate Claude instances that message each other and share a task list, where subagents report back to the session that started them. And two commands with similar names are something else: /agents inside a session now only reminds you to ask Claude or edit the folders (most tutorials still show the wizard it had up to version 2.1.197), and claude agents in your terminal manages background sessions, not subagents.
Making one
A subagent is a Markdown file whose frontmatter configures it and whose body becomes its system prompt. Only name and description are required. This is the one we used in our test:
.claude/agents/key-finder.md
---
name: key-finder
description: Finds the functions in this Python library whose signature takes a parameter the caller names. Use when asked which functions accept a given parameter.
tools: Read, Grep, Glob
model: haiku
---
You search this library's source for public functions whose signature includes
the parameter the caller names. Check the signatures themselves, not just
mentions in docstrings. Return one line per function: the function's name, then
the file and line of its definition. Nothing else.The description is what Claude matches tasks against, and Anthropic suggests a phrase such as “use proactively” in it if you want Claude to reach for the subagent unasked. The other fields that change what a subagent does:
| Field | What it does |
|---|---|
tools, disallowedTools | An allowlist and a denylist; leave both out and it gets every tool a subagent can have |
model | haiku, sonnet, opus, a full model ID, or inherit; left out, the main conversation’s model |
effort, maxTurns | Its reasoning effort, and how many turns before it stops and returns what it has |
permissionMode | Its own permission mode, ignored when the main session runs in auto mode or with edits accepted or permissions bypassed |
skills, mcpServers, hooks | Skills loaded in full at its start, MCP servers only it can use, hooks only while it runs |
memory | A folder of notes it keeps between sessions, for you, the project or this machine |
isolation: worktree | Runs it in a temporary git worktree, removed if it changed nothing |
background | Keeps it in the background even when Claude asks for the foreground |
Where the file lives decides who gets it. In order of priority: organisation-wide managed settings, then a --agents JSON flag for one session, then the project’s .claude/agents/ (the nearest one to where you started), then your ~/.claude/agents/, then plugins. So a project’s subagent overrides a personal one of the same name, which is the opposite of how skills resolve. Claude Code watches both folders, so an edited or new file is used at the next delegation without a restart. You can also just ask Claude to write one: Anthropic’s own example is “Create a personal code-improver subagent in ~/.claude/agents/ that scans files … Make it read-only and have it use Sonnet.”
The ones built in
| Subagent | Model | For |
|---|---|---|
| Explore | the main conversation’s | Searching and reading code, read-only; Claude sets how thorough |
| Plan | the main conversation’s | Research for a plan, read-only, used in plan mode |
| general-purpose | the main conversation’s, unless you set a subagent model | Multi-step work and code changes, with every tool |
| statusline-setup, claude-code-guide | Sonnet, Haiku | Setting up the status line; answering questions about Claude Code |
Explore and Plan skip your CLAUDE.md and the git snapshot to stay quick. To run exploration on a cheaper model, Anthropic’s advice is to define your own subagent named Explore with model: haiku, which replaces the built-in one.
Handing work over
There are three levels of asking. Name the subagent in your prompt (“use the code-reviewer subagent to…”) and Claude decides whether to delegate; @-mention it and it is guaranteed to run for that task; or start a whole session as it with claude --agent key-finder. Whenever Claude delegates, it writes the task the subagent receives: in our run below, our one-line question became a paragraph-long brief spelling out what “public” and “named exactly key” meant.
In an interactive session, subagents run in the background by default since version 2.1.232, so you keep working while they do; a non-interactive claude -p run waits for them. Up to 20 can run at once, and by default a subagent can start its own, three levels deep. A finished subagent is not gone: ask Claude to resume it and it carries on with its whole history, even after a restart if you resume the same session, except Explore and Plan, which are one-shot. Everything stays inside the one session, though, which is why running many agents from one for days takes separate sessions rather than subagents.
What it costs: the same search three ways
Anthropic is plain that subagents don’t make work free: each “sends its own requests, which count toward the same usage limits as your main conversation,” and “to spend less on them, choose a smaller model.” We measured what that means on October 9, 2026, with Claude Code 2.1.295 on Opus, in a fresh copy of the open-source more-itertools library (about 16,800 lines of Python). The question was the same each time: which public functions accept a parameter named key, with the file and line of each.
| Run | Time | Main conversation’s context at the end | Subagent | Claude Code’s cost figure (API rates) |
|---|---|---|---|---|
| No subagent (Agent tool switched off) | 45 s | 27,758 tokens | none | $0.29 |
| “Use the Explore subagent” | 150 s | 26,885 tokens | Explore on Opus, 11 turns, ending at 51,516 tokens | $0.81 |
| “Use the key-finder subagent” | 84 s | 27,818 tokens | key-finder on Haiku 5.5, 4 turns, ending at 29,905 tokens | $0.22, of which Haiku $0.01 |
All three found the same 14 functions (the run without a subagent also counted a class, and the other two named it as a close case). What the numbers show, for one question run once each: on a search this size, delegating saved the main conversation nothing, because in both delegated runs the main session checked the subagent’s answer with a quick search of its own before replying, and ended with about as much in its context as the run that did the whole search itself. Explore cost the most: it runs on the main conversation’s model, it read more, and in this non-interactive run its shell commands were refused, so it searched with its read-only tools instead, which may have added turns. The subagent on Haiku was the cheapest of the three. The runs drew on a subscription, so the dollar column is Claude Code’s own report at API rates, useful only to compare the runs with each other. A subagent earns its place when what it would read is far bigger than what it reports, such as a sweep of a large codebase or a pile of logs, and when its model is cheaper than yours; for a quick, targeted lookup, Anthropic’s own advice is to stay in the main conversation.
Codex’s subagents
OpenAI’s subagents page says current Codex releases “enable subagent workflows by default,” with nothing to switch on. The difference that matters most is who decides: Codex delegates only when you ask in so many words (“spawn two agents”, “one agent per point”) or an AGENTS.md file or a skill asks for it, and on its own only at its highest reasoning setting, Ultra. Claude Code delegates whenever a task matches a description.
- Defining one: a TOML file in
.codex/agents/or~/.codex/agents/withname,descriptionanddeveloper_instructions, plus any setting a Codex config takes, such asmodel = "gpt-6-luna", the cheaper model OpenAI suggests for lighter subagent work, orsandbox_mode = "read-only". The built-ins aredefault,workerandexplorer. - Watching them:
/agentswitches between their threads in the terminal; like Claude Code’s,codex agentsis something else, a list of sessions. - What one sees: OpenAI’s docs don’t say. Codex’s source does: with the current models, a subagent starts with the parent’s whole conversation unless told otherwise, like a Claude Code fork rather than a fresh subagent, runs three at a time by default, and can start its own. It works in the same folder as its parent, with no worktree of its own.
We tried Codex 0.160.0 on GPT-6.1 Sol on October 9, 2026, in the same repository. Asked to split the search between two subagents, one per file, it started both, waited for them and returned the same 14 functions in 83 seconds, the parent’s spawn calls asking for each to start with its whole conversation. A second test made the difference plain: given a code word in the prompt and told to start a subagent whose task left the word out, Codex’s subagent answered with the word; Claude Code’s general-purpose subagent, asked the same way, answered that it didn’t know it. Two things to know before scripting it. With codex exec --ephemeral, which keeps no session file, each spawn logged “collab spawn failed” with “no rollout found for thread id,” and one such run waited until we stopped it; without --ephemeral the same requests worked. And asked in the docs’ own phrasing to “have the key_finder agent do the search,” with a project file defining key_finder on GPT-6 Luna, Codex started a subagent named key_finder but without that agent type, so it ran on the parent’s model and settings, in a trusted folder as in an untrusted one. Asked a third time for a subagent “with agent type key_finder,” it did the same. In our runs the file’s settings never took effect, so check that a custom agent’s model and settings applied before you rely on them.
When to use one
Anthropic’s list is short. Use a subagent when the task produces output you don’t need in your main conversation, when you want to limit what tools or permissions the work gets, and when the work is self-contained enough to come back as a summary. Stay in the main conversation when the task needs back-and-forth, when several phases share a lot of context, for a quick targeted change, and when speed matters, since a fresh subagent has to find its bearings first. And give a subagent you use often a cheaper model: in our runs, that was the only setting that made delegating cost less.