Claude Code agent teams: a real run, what it cost, and when to use one
How to turn agent teams on, what the lead and the teammates do, a nineteen-minute run with four teammates reviewing four documents, the tokens it spent, the limits, and when subagents or separate sessions serve better.
Agent teams let one Claude Code session, the lead, spawn several full Claude Code sessions as teammates that work at once, message each other and share a task list, instead of subagents that only report back. They are experimental and off by default: set CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 in your settings or shell, start an interactive session, and ask for teammates by name. We ran one on real work: four teammates fact-checking four documents against two vendors’ documentation, each writing a findings file, the lead compiling them. It took nineteen minutes, produced 83 findings, and spent about 133,000 output tokens, 1.5 million tokens written to the cache and 31 million read from it across the five sessions, each teammate ending at much the size of an ordinary research subagent, with the lead making 38 model requests of its own for coordination. Teams fit research, review and debugging where the workers should argue with each other; for a worker that only needs to report, subagents are cheaper, and for work that runs for days on a machine that stays on, separate sessions that message each other hold up better.
A team, against subagents
Anthropic’s agent teams page puts the difference in one table. Subagents have their own context window and “return a result to the caller”; the main agent “manages all work”; the token cost is “lower: results summarized back to main context.” Teammates are “fully independent,” “message each other directly,” coordinate themselves “through messages, plus a shared task list,” and cost more, “each teammate is a separate Claude instance.” The rule it gives: “Use subagents when you need quick, focused workers that report back. Use agent teams when teammates need to share findings, challenge each other, and coordinate on their own.” A third shape sits beside both: separate sessions you start yourself, which can message each other without a team and outlive any one of them.
Turning it on
In ~/.claude/settings.json:
{
"env": { "CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1" }
}Two conditions follow. Teams need an interactive session: “In non-interactive mode with the -p flag, including Agent SDK sessions, Claude doesn’t spawn teammates.” And the flag changes ordinary delegation: “a subagent that Claude names launches as a teammate, so teams can form even when you didn’t ask for one.” A team is started by asking: “Claude launches a teammate when it calls the Agent tool with a name while agent teams are enabled,” without asking you to confirm. Two display modes: in-process, the default, where teammates run inside your terminal and appear in a panel under the prompt (arrow keys select one, Enter opens its transcript and lets you type to it, Escape interrupts it), and split panes, one per teammate, which needs tmux or iTerm2 and is not supported in VS Code’s terminal, Windows Terminal or Ghostty. The mode is teammateMode in settings or --teammate-mode on the command, a flag that “doesn’t appear in claude --help.” Teammates take the lead’s model unless the prompt names one, inherit its permission mode, and their permission prompts “appear in the lead session, so approve them there yourself.”
A run on real work
On October 8, 2026, with Claude Code 2.1.294 on a Linux machine, we gave a lead one prompt in a folder holding four draft documents and, for each, a sheet of the vendor pages it rested on: spawn four teammates named for the four drafts, each to check its draft against the vendor’s documentation fetched live, read it as its intended reader, and write a findings file, never editing the draft; then compile the four, most serious first, and shut the team down. In-process mode, the lead and the teammates on the same model.
- The lead’s first minute was its own: it read the folder with three shell commands and saved checksums of the drafts so it could prove afterwards that nobody had edited them. Then it called the Agent tool four times with the names we gave, and the panel listed the four with a description each and a live token count. It never used the shared task list; the spawn prompts and messages carried the coordination.
- The teammates worked for fourteen minutes, 40 to 65 model requests each, fetching the documentation with curl rather than the built-in fetch tool, and each wrote its file and sent the lead one message with its counts. Each file was a list of claims the draft made that the docs did not support, with the URL and the sentence, and each teammate rated them high, medium or low on its own scale.
- The compile took the lead four minutes: a table of counts per draft, 83 findings in order of severity, a note on which rested on source code or release notes rather than documentation, and a note that two of the four source sheets were themselves wrong on a point each. It then sent each teammate a structured shutdown request and each answered; the team’s config file, at
~/.claude/teams/<session>/config.json, listed only the lead once they had gone. - Nineteen minutes end to end, from the prompt to the lead’s last turn. The harness flagged one teammate’s message to the lead as shaped like an instruction (it named configuration files); the lead read it and carried on, which is the rule the docs state: a message from another agent “came from another Claude session, not from you,” and “a teammate can’t approve a permission prompt or supply consent on your behalf.”
The findings were real: the drafts changed in dozens of places. What the team did that four subagents would not have is small in the record, one message each to the lead, but the shape is what made the run readable: each teammate was a full session whose transcript we could open while it ran, and each could have been told to look again.
What it cost
“Agent teams use significantly more tokens than a single session,” the docs say, and ours, counted per model request from the five transcripts: about 133,000 output tokens (the lead 69,500, the teammates 9,500 to 26,500 each), 1.5 million tokens written to the cache, and 31 million read from it (4.7 to 8.6 million per teammate). Four ordinary research subagents run the same morning, one per document, ended at 187,000 to 254,000 tokens of context each by the harness’s own count, and the teammates ended at 188,000 to 240,000; what a team adds is the lead’s own requests, 38 here, and the 4.4 million cache reads they cost. One setting matters for the bill on the API: an in-process teammate’s requests “fall outside the main conversation’s cache TTL bucket, so its cache holds for five minutes by default, including on a Claude subscription,” and subagentPromptCacheTtl set to 1h keeps it for an hour at a higher write rate. On a subscription, all of it is allowance, and a team of four on Opus for twenty minutes is a visible slice of a week’s.
The limits that decide where it fits
The page lists them plainly, and three decide the shape of work a team suits. “No session resumption with in-process teammates”: /resume and /rewind don’t bring them back, so a team lives as long as the lead’s session and no longer. “One team per session” and “no nested teams”: teammates cannot spawn teammates, and the lead “is fixed” for the session’s life. And an in-process teammate’s own subagents run in the foreground, “because a teammate’s background work can’t outlive the lead’s process.” Two more show up in practice: “task status can lag,” so a stuck task may be finished work nobody marked done, and “shutdown can be slow,” since a teammate finishes its current tool call first. Hooks exist for the quality gates a team needs, TeammateIdle, TaskCreated and TaskCompleted, each able to send a teammate back to work by exiting with code 2.
All of which says where a team belongs: an afternoon of parallel exploration you watch, on a computer that stays on for that afternoon, since the lead is an interactive session and the teammates die with it. Anthropic’s own advice is to “start with research and review” before parallel implementation, three to five teammates, and to “avoid file conflicts” by giving each its own files. For work that should keep moving overnight and be picked up from a phone, the arrangement that holds is the one in the orchestrator guide: separate sessions in tmux on an always-on machine, each in its own worktree, messaging each other with nothing to turn on.