Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk5 min

Understanding Tokens per Second: A Practical LLM Benchmark Guide

LLM TPS is not a universal score. Define the timing window and workload, then measure both per-request responsiveness and aggregate throughput at realistic concurrency.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universal “good” tokens-per-second (TPS) score for a large language model. A useful result must say what was counted, which part of the request was timed, and whether the rate describes one request or all requests combined. For interactive use, pair generation speed with time to first token (TTFT) and full response time; for capacity planning, measure aggregate throughput at a stated concurrency and latency limit.

What does tokens per second measure?

TPS is a rate, not a single standardized benchmark. It can mean generated output tokens per second for one request, or the total output tokens per second produced across concurrent requests. Some tools also differ on whether timing starts before or after the first token. A TPS figure without those definitions is difficult to interpret.

As an Amazon Associate I earn from qualifying purchases.

For example, Ollama describes its per-request output rate as generated output tokens divided by generation time after the initial wait. That is a project-specific methodology, not a universal standard. NVIDIA likewise notes that benchmarking tools can define token metrics differently. (Ollama’s methodology; NVIDIA’s benchmarking concepts.)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Per-request output TPS: the generation pace of one response stream. It does not convey how long the user waited for the first token or how many simultaneous users the service can handle.
  • Aggregate output throughput: total output tokens produced per second across the requests running concurrently. This is a system-capacity measure, not the speed experienced by any one user.

Which speed metrics matter to a user?

Interactive inference has a startup phase and a generation phase. The prompt is processed before the model begins producing its response; then output tokens are generated one by one. A long prompt can increase the wait before the first visible token, while a longer requested answer generally takes more time to finish.

  • Time to first token (TTFT): time from sending the request until the first content token arrives. Depending on where the measurement is taken, it can include queueing, prompt processing, and network delay. NVIDIA’s client-side description includes all three.
  • Time per output token (TPOT) or inter-token latency (ITL): the average interval between generated tokens after the first token. The exact calculation varies by tool. NVIDIA describes its GenAI-Perf calculation as generation time divided by output-token count minus one, excluding TTFT.
  • End-to-end latency: time from sending a request to receiving its final token. It captures the whole wait, but its treatment of queueing and transport depends on the measurement tool.

NVIDIA summarizes TTFT as “the time it takes to process the prompt and generate the first token.” (NVIDIA Developer; Databricks endpoint benchmarking.)

Keep units explicit: TPS is tokens per second; TTFT, TPOT, and ITL are usually expressed in milliseconds or seconds. If you convert an average interval such as TPOT into a rate, state that the rate is its reciprocal and whether the first-token wait is excluded. For example, an interval in seconds per token can be inverted to tokens per second, but that still does not include TTFT unless the timing definition says it does.

How many tokens per second is a good speed for an LLM?

There is no evidence-backed universal threshold. A result that feels responsive for a short interactive answer may be unsuitable for a workflow with long outputs, and a high aggregate rate may coexist with slow individual responses under concurrency. The target depends on the model, prompt and output lengths, serving setup, workload, and the latency constraints of the application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Databricks recommends maximizing throughput within an application’s latency budget. Google Cloud’s accelerator benchmarking guidance similarly frames inference comparisons around workload-specific performance and latency constraints. Choose the metric that matches the decision: user-facing responsiveness for interactive use, or total output capacity at an acceptable latency for batch processing and service sizing. (Databricks; Google Cloud.)

How to benchmark LLM inference speed

  1. Define the decision. Decide whether you are evaluating an interactive model, sizing an endpoint, comparing local accelerators, or estimating batch capacity. Select success metrics accordingly. NVIDIA distinguishes performance benchmarking from load testing at scale; Databricks presents throughput in relation to a latency budget.
  2. Specify a representative workload. Use prompts that reflect the intended task and record input-token and output-token lengths or their distributions. For a fair comparison, hold the model and version, tokenizer, precision or quantization, serving stack, streaming mode, and generation settings constant. Input length affects prefill and TTFT; output length affects generation duration.
  3. Warm up and repeat. Record the benchmark tool and methodology, how warm-up was handled, the number or approach of repeated runs, and whether reported results are a mean, median, or percentile. NVIDIA’s benchmarking guide organizes the process around warm-up, use-case sweeps, and analysis; consult the exact tool documentation for version-specific options. (NVIDIA benchmarking guide.)
  4. Measure one stream and a concurrency sweep. A single-request test characterizes one response stream. Repeat with increasing numbers of concurrent requests to reveal aggregate throughput, queueing, and the effect on individual-request latency. More parallel requests can raise total throughput until capacity limits are reached, while increasing latency.
  5. Report the metric set, not a headline number. Include per-request output TPS or TPOT, TTFT, end-to-end latency, aggregate output throughput, concurrency, and success or error rate. Include p50 and, when the sample size supports it, a tail percentile such as p95 or p99. Averages alone can hide slow requests.
  6. Stop at the service constraint. For an interactive service, identify the point at which the chosen latency requirement is exceeded and report the throughput that was sustainable within that requirement. Google Cloud describes increasing concurrency until a P99 latency SLA is violated, then recording sustained throughput. (Google Cloud accelerator benchmarking.)
  7. Disclose the scope. State whether results are independently measured, vendor-published, or produced by your own test. Provider measurements can include network path and service load; sequential and concurrent tests answer different questions. One run or a vendor headline is not a universal model or hardware specification.

What to include in a benchmark report

  • Model and version, tokenizer, precision or quantization, and serving configuration.
  • Prompt task and input/output token lengths or distributions.
  • Generation settings and whether streaming was enabled.
  • Tool and version, timing definitions, warm-up, repetitions, and summary statistic.
  • Concurrency and whether token throughput is per request or aggregate.
  • TTFT, TPOT or ITL, end-to-end latency, output throughput, and errors, with units and relevant percentiles.
  • Hardware or provider context and any latency target used to determine sustainable throughput.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why two TPS figures may not be comparable

Before comparing results, check that both use the same numerator and timing window. A rate based on output tokens after TTFT is not the same as one that includes the initial wait or counts input and output tokens. Also check whether the number is for one request or aggregated across concurrent requests.

Then compare workload and conditions: model, prompt and answer lengths, serving and generation settings, streaming, concurrency, and measurement location. Finally, compare responsiveness and tail behavior alongside throughput. A system that produces more total tokens per second may still deliver a slower response to an individual user or violate a service’s latency target. Google Cloud recommends fixed-model normalization and latency constraints for inference comparisons. (Google Cloud.)

Speed alone says nothing about answer quality. A benchmark is useful only when it answers a specific question about a defined workload and service requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.