In a September 24, 2026, DEV Community post, an author writing as jidonglab reported that workflow subagents accounted for 48% of their Claude Code cost in one month of usage, even though output tokens were just 0.9% of their total tokens. Those figures describe one developer’s workload—not a typical cost split—and the post does not specify the month’s start and end dates.
The gap is a reminder that output tokens are only one part of usage accounting. Input context, cache writes and reads, tool definitions, and tool results can all contribute to costs. Whether subagents are expensive for you depends on how many you launch, the work and context each receives, and the models and rates involved.
As an Amazon Associate I earn from qualifying purchases.
What the 48% and 0.9% figures actually measure
jidonglab analyzed local Claude Code JSONL session transcripts, classifying assistant messages as main-thread or subagent activity using the isSidechain field. The author summed input, cache-creation input, cache-read input, and output tokens for each group, then applied the relevant model rates to estimate cost.
In that account-specific calculation, subagents were responsible for 48% of Claude cost. Separately, output tokens made up 0.9% of all tokens. The 0.9% figure is a share of token volume, not a share of cost. The post does not publish raw transcripts or an independent audit, and it does not establish that other users will see similar proportions.
#1 Best Overall
The author described a workload involving large audits and research fan-outs, and expected subagent costs to be lower for developers who mostly make single-file edits. The calculation also depends on the transcript fields present in the author’s observed version, how sidechain activity is classified, and the model-rate mix used.
Why subagents can consume cost without producing many output tokens
A subagent’s answer is only the visible end of its work. It may also receive instructions and context, call tools, and process tool results before returning a short finding. The input and tool-related usage can therefore be substantial even when the final response is brief.
jidonglab estimated that each subagent in their setup began with about 51,000 tokens of context. The author attributed that starting context to items such as system instructions, tool schemas, project instructions, memory, and skill listings. This is an estimate from that setup, not a Claude Code-wide baseline or evidence that every request repeats an identical full context.
Free tools Windows power users keep installed
One-click scans. No signup required.
In general, the more work agents do, the more opportunity there is for accumulated conversation and tool results to add to their input. Anthropic’s pricing documentation explains that tool definitions and tool results contribute to input accounting, and that input, cache writes, cache hits, output, and some tool usage can have distinct cost treatment. Prompt caching can make repeated input cheaper, but it does not make it free; the actual charge depends on the model and route.
What the author’s other usage figures show
The same post reported that spending was concentrated in a relatively small part of the author’s activity. These are observations from that account and period, not general thresholds:
- 45 sessions costing more than $100 accounted for 79% of the author’s cost.
- Requests above 400,000 tokens accounted for 54% of main-session cost.
Those observations suggest that reviewing unusually large sessions and requests may be more useful than focusing only on the length of an agent’s final answer. They do not show that a $100 session or a 400,000-token request is inherently wasteful, or that either number predicts another user’s bill.
Rank #4
How to investigate your own subagent costs
Use the same basic distinction as the author—main-thread versus subagent activity—but treat your transcript fields and billing setup as specific to your environment. Claude Code’s setup and access routes vary; Anthropic’s setup documentation describes access through Console, Claude plans, and enterprise arrangements.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →- Choose a representative period. Record the dates, models, and access or billing route involved. A month with a large audit or research project may not represent routine work.
- Inspect local session transcripts. Identify how the transcript version marks sidechain or subagent messages. The author used
isSidechain; do not assume that field or its meaning is unchanged in every version. - Separate usage categories. Keep input, cache-creation input, cache-read input, and output distinct where your records expose them. Do not treat total tokens or output-token share as a direct proxy for cost.
- Apply the rates that match your models and route. Rates and plan treatment can differ. An estimate based on the wrong model mix or pricing basis will not explain the actual charge.
- Compare main-thread and subagent activity. Look for expensive sessions, large requests, repeated context, and tool-heavy work. Investigate what drove a spike before concluding that delegation alone caused it.
Ways to reduce avoidable subagent work
jidonglab suggested the workflow changes below as ways to control usage. The author explicitly said they had not run a clean before-and-after month under these rules and would not claim a savings percentage. Treat them as operational ideas to evaluate against your own work, not as quantified or independently tested savings.
Best Value
- Set a ceiling on agents per workflow. Require a concrete reason before exceeding it, rather than launching more agents by default.
- Batch small tasks. Give one agent a coherent set of related, small items instead of assigning a separate agent to every tiny task.
- Match the model and effort to the work. Consider a lower-cost model or effort setting for mechanical collection and formatting, reserving deeper reasoning for review and synthesis where appropriate.
- Trim irrelevant context. Keep global instructions focused and load project-specific material only when it is useful to the task.
- Hand off when the topic changes or history grows unwieldy. Write a concise note containing the relevant findings and next steps, then start a fresh session rather than carrying unrelated history forward.
- Answer simple lookups directly. If delegation adds no useful independent work, handle the lookup in the main session.
How to read the result
The case study does not show that subagents are always inefficient or that output-token reduction alone will control a bill. It shows that, in one developer’s workload and cost calculation, subagent activity was a large expense despite output tokens being a very small share of total tokens. For another user, the balance will depend on agent count, task independence, starting and accumulated context, tool use, cache treatment, and model rates.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




