Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk4 min

Why a Claude Code Log Sum Can Double Even After Message-ID Deduplication

Deduplicating repeated Claude Code message IDs is only one step. The remaining total may reflect placeholder output counts, cumulative resumed-session usage, subagents, or real context processing.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Deduplicating repeated message IDs fixes one common overcount, but it does not guarantee that a total parsed from Claude Code logs is correct. Other causes include placeholder output counts, cumulative results from resumed sessions, mismatched subagent scope, or genuinely heavy context processing. Without the Claude Code version, source format, and representative records, the remaining discrepancy cannot be assigned to one cause.

First, identify which total you are calculating

“Usage” can mean several different things: output from one response, a query’s completed token usage, a resumed session’s accumulated spend, all models in an agent tree, or a local estimate of cost. Those figures are not interchangeable.

As an Amazon Associate I earn from qualifying purchases.

Question Relevant source Important qualification
How much output was reported as a response finished? Agent SDK result message usage Use the completed result rather than treating each assistant message’s output count as final.
How much did each model use, including subagents? SDK modelUsage or model_usage in the same cost guide The result’s top-level usage covers the main loop, not subagents.
What was the output progress during streaming? Documented message_delta usage events in the cost guide Use the stream’s documented events, not repeated snapshots as if each were new usage.
What does a local session transcript show? The specific Claude Code transcript and parser The SDK guide does not establish that every local JSONL version exposes identical fields or snapshot semantics.
What was billed? An authoritative billing record An SDK cost estimate uses a client-side price table and may differ from billed cost when prices or billing rules differ.

Why deduplicating message IDs may not be enough

Repeated records can represent one response

Anthropic’s Agent SDK guide says that when Claude uses multiple tools in one turn, the messages in that turn can share an ID; count that shared response ID once when accumulating per-step usage. The guide states: “When Claude uses multiple tools in one turn, all messages in that turn share the same ID, so deduplicate by ID to avoid double-counting.” This is SDK guidance, not proof that repeated rows in every local transcript have the same meaning. Group candidate rows by the documented identifier and inspect their usage fields. Do not merge distinct IDs merely because their text looks alike.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Assistant output counts can be placeholders

In the SDK, an assistant message’s output_tokens can be a placeholder based on what the API reported when the message started. For a completed SDK query, use the result message’s usage; use modelUsage or model_usage when you need per-model usage or whole-tree accounting. For streaming progress, use the documented message_delta usage events. These distinctions are described in Anthropic’s Agent SDK cost guide.

Resumed-session results may include earlier usage

A result from a call that resumes a session can include spend from earlier in that session. Adding every result together can therefore count earlier usage more than once. For a resumed session total, use the latest applicable result rather than summing cumulative results. Streaming-input mode has its own running-total and reset rules; aggregate it according to those documented boundaries.

Main-loop usage and whole-tree usage have different scope

The SDK result’s top-level usage covers the main loop and excludes subagents. modelUsage or model_usage is the documented route to whole-tree token accounting. Before combining a parent total with child traces, check whether the parent already includes that work; otherwise the scopes may overlap.

A corrected count can still be high

Claude Code sends conversation history and project context with later turns. A parser can remove duplicate representations correctly and still show substantial usage because the session processed a large or growing context. That is different from a counting error. Anthropic’s explanation of Claude Code usage is in its cost guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical way to investigate the remaining discrepancy

  1. Record the source and version. Note whether the data came from streamed Agent SDK messages, an SDK result, a local session transcript, or an export rebuilt by another tool. Record the Claude Code and SDK versions; local transcript semantics are not established as identical across formats.
  2. Inspect repeated IDs and usage objects. For SDK parallel-tool messages, count a shared response ID once. Check whether the rows repeat the same usage snapshot. Do not deduplicate separate response IDs solely on similar content.
  3. Verify the field used for output. If the sum reads assistant-message output_tokens, it may be adding start-of-response placeholders. For a completed SDK query, compare with the result message’s usage; for streaming, use the documented delta events.
  4. Check the aggregation boundary. Determine whether each result is an independent query total or a cumulative resumed-session result. Avoid adding successive cumulative snapshots; follow streaming-input reset boundaries where applicable.
  5. Check subagent coverage. Decide whether the desired number is main-agent-only or whole-tree usage. Use the relevant SDK field and avoid adding parent and child figures until their scopes are known.
  6. Compare billing only to billing records. A local estimate is not an authoritative billed amount; SDK estimates can diverge when prices or billing rules differ.
  7. Keep a small redacted sample. Retain representative IDs and usage fields alongside the version details. Transcripts may contain prompts, tool outputs, URLs, credentials, and personal information; redact secrets before sharing them for parser debugging.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What is known about duplicate-ID prevalence

A third-party analysis by Frederick Douglas Pearce reported duplicate assistant message IDs in 986 of 1,047 files (94%) in the author’s measured corpus, with naive row summation inflating that corpus’s total by 1.99×. Those are corpus-specific results, not an Anthropic statistic, a platform-wide prevalence estimate, or evidence about any particular reader’s file: the author’s analysis.

Anthropic’s separate compliance-session API documentation instructs clients to deduplicate listed sessions by session ID and messages by message ID. It describes captured transcripts as reconstructed from API calls and notes that content may be unavailable or truncated in specified circumstances. That guidance applies to the compliance API; it should not be generalized to every local Claude Code JSONL format.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.