Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk4 min

What Does One AI Token Actually Cost?

API token prices depend on the model and billing category. Calculate input, output, cached tokens and separate tool charges to estimate what a request will cost.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universal price for one AI token. For API use, the price depends on the provider, model, token category and service mode—and your request’s input, output, caching and tool usage. Providers usually quote rates per million tokens, so estimate the whole request rather than treating a token as a fixed-dollar unit.

How to calculate the cost of an API request

Apply the selected model’s rate to each usage category separately, divide each product by one million when the rate is quoted per million tokens, then add any separately billed tools or services:

Estimated request cost = (input tokens × input rate + cached input tokens × cached input rate + output tokens × output rate) ÷ 1,000,000 + separately billed tool or service charges

Use the exact categories on the model’s rate card. Some providers distinguish cache reads from cache writes or charge for cache storage; some count reasoning tokens as output. Do not assume every input token is cached or that providers count images, audio, video and other modalities alike.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example using GPT-6 Sol rates

OpenAI lists GPT-6 Sol at $2.00 per million standard input tokens, $0.20 per million cached input tokens and $10.00 per million output tokens for short context. A request with 10,000 standard input tokens and 1,000 output tokens, with no cached input or separately billed services, would cost $0.03 at those listed rates: ($2.00 × 10,000 + $10.00 × 1,000) ÷ 1,000,000. This is an illustration of the calculation, not a promised invoice. Check the current OpenAI API pricing page for the applicable model, context and service-mode row.

Published API rates show why a token has no single price

These USD list-price examples are snapshots, not a provider-neutral average or a like-for-like comparison of quality or workload. Effective charges can vary with endpoint, tier, contract, geography, discounts and rate dates.

Provider and model Rate category Listed price per million tokens Scope
OpenAI GPT-6 Sol Standard input $2.00 Short context; check current model and service-mode row
OpenAI GPT-6 Sol Cached input $0.20 Short context; check current model and service-mode row
OpenAI GPT-6 Sol Output $10.00 Short context; check current model and service-mode row
OpenAI GPT-6 Astra Input $10.00 Flagship table, short context
OpenAI GPT-6 Astra Cached input $1.00 Flagship table, short context
OpenAI GPT-6 Astra Output $50.00 Flagship table, short context
Anthropic Claude Opus 4.5 API Standard Global Input $5.00 May 27, 2026 list-price document; cache writes and hits are distinct categories
Anthropic Claude Opus 4.5 API Standard Global Output $25.00 May 27, 2026 list-price document
Anthropic Claude Opus 4.5 API Batch Input $2.50 May 27, 2026 list-price document
Anthropic Claude Opus 4.5 API Batch Output $12.50 May 27, 2026 list-price document
Google Gemini 3.7 Flash paid Standard Input $0.75 through Dec. 31, 2026; $1.50 starting Jan. 1, 2027 Scheduled rates; separate context-caching and storage charges may apply
Google Gemini 3.7 Flash paid Standard Output $3.75 through Dec. 31, 2026; $7.50 starting Jan. 1, 2027 Scheduled rates; separate context-caching and storage charges may apply

See the current OpenAI API pricing, Anthropic Claude API pricing and Gemini API pricing for the applicable rate rows and effective dates before budgeting.

What can change your total

Input and output mix

Input and output rates can differ substantially. Calculate them separately; multiplying all conversation tokens by the input rate can understate or overstate the cost.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Caching

Reused prompt prefixes may qualify for a lower cached-input rate, but cache writes or storage can have separate charges. OpenAI describes automatic prompt caching for supported models on prompts longer than 1,024 tokens; that does not mean every token in every request is cached. Check the provider’s prompt caching documentation and rate card.

Processing mode

Some models offer discounted batch or lower-priority processing, while faster or priority modes can cost more. Eligibility and rates vary by model; compare the specific mode you will use, rather than assuming a discount applies to every request. OpenAI’s pricing documentation describes applicable service modes.

Context length and processing region

Long requests or regional endpoints may have different rates. OpenAI says GPT-6 Astra requests above 272K input tokens are charged at twice the input and cache rates and 1.5 times the output rate for the full request. Its pricing documentation also lists a 10% uplift for eligible regional-processing and FedRAMP endpoints. These conditions apply to the stated model and endpoint cases, not API requests generally; confirm the current terms on OpenAI API pricing.

Tools and other modalities

Images, audio, video, search grounding and other tools may use different billing rules or add charges. Gemini’s pricing page lists separate grounding and tool fees. Check whether retrieved content is included in token billing for the specific tool and endpoint.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tokenization and reasoning

The same text can produce different token counts on different models, and models may generate different amounts of output or reasoning. A lower price per token therefore does not guarantee a lower cost for a completed task. OpenAI recommends testing representative work and comparing total usage and cost in its usage guidance.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Estimate and verify your own API cost

  1. Choose the exact setup. Record the provider, model, endpoint and service mode. Confirm that the rate card applies to your account and region.
  2. Collect usage by category. Use the request response or provider dashboard to find input, output, cached input and any other reported usage categories.
  3. Apply each rate separately. Multiply each category’s token count by its matching rate. Divide by 1,000,000 when the rate is per million tokens.
  4. Add non-token charges. Include separately billed tools, cache storage and modality fees where applicable.
  5. Check conditions on the rate. Verify context thresholds, regional processing, batch eligibility, account terms and effective dates.
  6. Test representative tasks. Compare the total cost to complete the same kind of work on each candidate model—not only the visible answer or input-token price.
  7. Reconcile the estimate. Compare your estimate with actual usage in the provider dashboard or request response. OpenAI documents both account-level dashboard review and request-level usage inspection in its usage guidance.

How to compare providers fairly

Use the same representative workload and compare the costs that actually apply to it. Include model capability for the task, input and output rates, cache reads and writes, context thresholds, service-mode eligibility, region and contract terms, and separately billed tools or modalities. A single input-price column cannot establish which provider will be cheapest for your work.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.