October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

LLM Token Counter Errors: Two Common Traps and a Runnable Demo

A local token counter can use the wrong encoding or miss request structure. Compare encodings with a runnable Python demo and choose a count that matches what you need.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A simple LLM token counter can be wrong in two different ways: it may use an encoding that does not match the target model, or count visible text while leaving out the structured parts of the request. The Python demo below compares encodings for a text string; it is useful for illustrating tokenization, but it is not an exact count of every API request.

Why a token count can disagree

Tokens are units produced by a tokenizer, not fixed chunks of characters or words. The result depends on the encoding, the text—including its language and spelling—and the surrounding text. A count or token ID from one encoding should not be assumed to transfer to another model.

As an Amazon Associate I earn from qualifying purchases.

There is a separate issue: an API request can contain more than the visible text. Roles, message boundaries, tool definitions, schemas, images, and files may contribute to the input count. Counting only each message’s content string therefore may not count the complete request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Failure mode 1: using the wrong encoding

A counter that fixes one encoding and later gets reused for another model may return a plausible but mismatched text count. OpenAI’s Cookbook shows selecting the encoding for a model with tiktoken.encoding_for_model(model), rather than assuming a single encoding fits all models. The Cookbook also cautions that its method for estimating message counts is not a permanent guarantee. See OpenAI’s token-counting example and the OpenAI Help Center explanation of tokens.

Runnable encoding comparison

Install the package, then run this code to compare local text counts for the same Japanese string:

python -m pip install tiktoken
import tiktoken

text = "お誕生日おめでとう"
for name in ("p50k_base", "cl100k_base", "o200k_base"):
    encoding = tiktoken.get_encoding(name)
    print(f"{name}: {len(encoding.encode(text))} tokens")

The published Cookbook example gives 14 tokens for p50k_base, 9 for cl100k_base, and 8 for o200k_base. These are the results for that example string, not a general conversion rule or a guarantee for other text. The snippet counts only the string; it does not measure API billing or a complete chat request.

Failure mode 2: counting text instead of the request

Chat requests have structure. OpenAI’s documentation says that input-token counts include formatting tokens used to represent that structure, such as message roles and boundaries. Depending on the request, tools, schemas, images, or files can also matter. A local loop over visible message text does not necessarily include these elements.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For OpenAI Responses requests, the input-token counting endpoint accepts supported input forms and accounts for request formatting. To count before sending, pass the same supported input structure you intend to send; a plain string count answers a narrower question. The official guide’s current example is:

from openai import OpenAI

client = OpenAI()
count = client.responses.input_tokens.count(
    model="gpt-6-astra",
    input="Tell me a joke.",
)
print(count.input_tokens)

This is the model name and API usage shown in the official token-counting guide; model availability and APIs can change. The endpoint is provider-specific, so its availability and supported input formats should not be assumed for other services.

For Hugging Face chat models

Use the target model tokenizer’s chat template so the conversation is rendered in the format that model expects. If you tokenize the rendered template separately, set add_special_tokens=False when the template already includes the required special tokens; otherwise, tokenization can add duplicates. See the Hugging Face chat templating documentation.

Choose a counting method for the question you have

Method What it counts Best use and limitation
Raw text with a chosen encoding A text string encoded with the selected tokenizer Useful for inspecting or estimating text when the encoding matches the target model; it omits request structure not included in the string.
Model-aware local tokenizer Text tokenized using the target model’s tokenizer or encoding Preferable to a fixed, unrelated encoding for local text counts, but it may not reproduce all request formatting or provider-side behavior.
Chat-template tokenizer A conversation rendered with the target open model’s chat format Preserves model-specific formatting; avoid adding special tokens a second time if the template already supplies them.
Request-level counting endpoint Supported structured input and its formatting Useful for a pre-send count when the provider offers it; support depends on the provider and accepted request forms.
Returned API usage Usage reported after a call Use it to inspect actual reported usage, rather than treating a visible-text estimate as the result. Generated output cannot be predicted from input text alone, and output totals may include tokens that are not visible in the returned text.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret estimates

Character-to-token and word-to-token ratios are rough estimates, not substitutes for tokenization. OpenAI’s Help Center gives about four characters per token and about three-quarters of a word per token as rough English estimates, while noting that the relationship varies with text and language. Those ratios should not be used as exact counts, especially for languages or text that tokenize differently.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For the question “How many tokens is this?”, first decide whether “this” means a text string, a formatted conversation, or the full structured request. Use the target model’s tokenizer for local text inspection; use a request-level counter for supported structured input when available; and check returned usage after a call if you need the API’s reported usage. No local pre-send count should be presented as an exact match for every API request.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.