October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

How to Benchmark Inference Throughput per GPU for AI Agents

Measure AI agent inference with representative multi-turn workloads, a documented serving setup, warm-up, and a concurrency sweep. Report system TPS and latency before showing any per-GPU average.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Benchmark a representative agent workload against a documented inference-serving setup, warm up the service, and increase concurrent load until throughput levels off. Report total system output tokens per second (TPS), latency, concurrency, and the full configuration. If you divide TPS by the GPU count, label it a per-GPU average—not single-GPU performance or scaling efficiency.

What to measure—and what “per GPU” means

Inference throughput is the amount of model output a serving system produces over time. For an AI agent, the result depends not just on the GPU but also on the model, serving software, batching and parallelism settings, request mix, and the agent’s pattern of turns and tool use. A useful benchmark therefore reports the whole system and workload, not just a tokens-per-second figure.

As an Amazon Associate I earn from qualifying purchases.

NVIDIA defines total system TPS as output-token throughput across simultaneous requests. Its AIPerf definition calculates output tokens over the interval from the first request to the final response; configured warm-up can be excluded. If the system produces T output tokens per second across G GPUs, the simple normalized figure is T ÷ G TPS per GPU. Keep T and G beside that average.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This division is useful for a rough comparison of similarly configured systems, but it does not measure what one GPU would achieve alone. Multi-GPU parallelism, communication, batching, and system design affect the aggregate result. It is also not scaling efficiency: that would require comparing performance at different GPU counts under an appropriately controlled setup.

#1 Best Overall
ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950
  • System Compatibility Note: This 2-slot card measures 271 x 112 x 39 mm and requires a single 12V-2x6-pin power connector. Please verify chassis and PSU compatibility before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • Professional Intel Arc Pro B70 GPU: Built on the Intel Xe2-HPG architecture, it features 32 Xe cores and 256 XMX engines, designed to accelerate AI, rendering, and complex visualization workloads.
  • Massive 32GB GDDR6 VRAM: Equipped with 32GB of high-speed GDDR6 memory on a 256-bit bus, running at 19 Gbps, which allows for handling large AI models and complex datasets locally.
  • High-Performance Engine Clock: Delivers an engine clock of 2540 MHz, providing the compute power needed for demanding professional applications and AI inference.

Build a workload that resembles your agent

Start by describing the requests the inference service actually receives. A short, fixed prompt and response may be convenient, but it can miss the variable context and repeated generations of an agent session.

Capture the workload shape

  • Model and tokenizer: record the model name and version and the tokenizer used to count tokens.
  • Input and output lengths: record their distributions, not only a single average. Include how input context grows across turns.
  • Turns and tools: record the turn-count distribution and the tool-use pattern. If representative multi-turn or coding/tool traces are available, use them or derive a documented workload profile from them.
  • Generation settings: document decoding and sampling settings so another run does not silently use a different generation workload.

The September 28, 2026 AgentPerfBench preprint argues that single-turn chat tests and fixed input/output lengths can miss realistic agent workloads. It describes profiles based on empirical per-turn input length, output length, and turn-count distributions, and argues for testing load through hardware saturation. Treat those recommendations as recent research, not a universal benchmark standard. The authors report more than 3,000 benchmark results and more than 140,000 per-kernel Nsight Compute profiling records across four GPU platforms and 11 model architectures; those are figures reported by the preprint’s authors.

Rank #2
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

Fix and document the serving configuration

Keep the serving stack constant when comparing runs. Record the GPU model and count, serving engine and version, precision or quantization, parallelism and batching settings, model-serving configuration, and network placement. Also record the concurrency or request-arrival policy and the measurement duration or interval used by the tool. A change in any of these can change both throughput and latency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA documents AIPerf as a client-side benchmarking tool for OpenAI-compatible inference services. Its guide recommends placing the client on the same host when network latency is not part of the test. If client and server are on different hosts, report that arrangement: end-to-end latency can include network time, so the result is not directly equivalent to a same-host run.

Rank #3
GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card
  • Powered by Radeon AI PRO R9700 - Supercharge you workflow with the cutting-edge RDNA 4 Architecture and 2nd-gen AI Accelerators.
  • 32GB GDDR6 with 256-bit memory bus - Tackle larger, more complex projects without limits.
  • PCIe Gen 5 - Unlock lightning-fast data transfers with PCIe Gen 5 support.
  • GIGABYTE TURBO Fan Cooling System - Indented metal cover and blower fan increase airflow intake, while the vapor chamber, all copper heat sink, and metal frame offer efficient heat dissipation. Optimized airflow design allows for easy multi-GPU scalability.
  • Double Ball Bearing Fan - Delivers superior heat resistance and rotational efficiency for better performance and a longer lifespan compared to conventional sleeve fans.

Run a warm-up and sweep the load

  1. Start the service with the configuration you intend to compare. Save the model, serving, and hardware settings. Do not change them midway through a sweep.
  2. Warm up the service before collecting measured results. AIPerf’s documented example includes a warm-up before its sweep. Preserve the tool’s configuration and the warm-up settings with the results.
  3. Run a concurrency sweep. Test representative deployment loads, then increase concurrency far enough to observe whether throughput levels off. NVIDIA advises using concurrency for most benchmarks and notes that throughput can saturate while latency continues to rise.
  4. Save the run artifacts. AIPerf’s example exports JSON and CSV results and includes a latency-throughput plot. Preserve the artifacts and the command or configuration used to generate them so the run can be checked or repeated.

Do not report only the best-looking load point. Concurrency and request rate both control offered load; state which policy you used. A throughput number without its load and latency can conceal whether the system delivered a useful response time or was already queueing heavily.

Report the metrics together

Use a metric set that shows aggregate capacity and the experience of an individual request. Definitions can vary between tools, so identify the implementation and do not assume identically named metrics are interchangeable.

Rank #4
ASRock Intel Arc Pro B60 Creator 24GB Graphics Card, Workstation GPU, Xe2-HPG, 2400MHz, 24GB GDDR6 192-bit, PCIe 5.0, 4X DP 2.1, Blower
  • System Compatibility Note: 2-slot card, 271x112x39mm, single 8-pin power, 200W TDP. Verify chassis clearance and PSU capacity before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • 24GB GDDR6 on 192-Bit Bus: Massive 24GB memory with 456 GB/s bandwidth – ideal for LLMs, AI inference, 3D rendering, and generative design.
  • Intel Xe2-HPG Architecture: Built on Intel's next-gen architecture with 20 Xe cores and 160 XMX engines for AI acceleration (197 INT8 TOPS).
  • PCIe 5.0 Support: PCI Express 5.0 x16 interface for maximum bandwidth with the latest workstation platforms.
Metric What it describes How to interpret it
Total output tokens per second (TPS) Aggregate output-token throughput across simultaneous requests. AIPerf measures output tokens over the interval from the first request to the final response; configured warm-up can be excluded. Report this as system throughput, along with the GPU count and load.
TPS per user A single-client perspective: output sequence length divided by end-to-end latency for a request. Do not confuse it with aggregate system TPS.
Time to first token (TTFT) Time from query submission until the first received output token, when the response contains content. Shows how long a user waits before generation begins.
Inter-token latency (ITL) or time per output token (TPOT) Average time between consecutive output tokens. Check whether the tool includes TTFT in its definition; AIPerf excludes it.
End-to-end latency Time from query submission until the complete response, including queueing, batching, and network latency. Relate it to the deployment’s response-time requirement.
Requests per second (RPS) Successful requests completed per second over the benchmark interval. Useful beside TPS when comparing workloads with different response lengths.

Include averages and relevant tail percentiles when the tool provides them, and say which statistics you report. For agent workloads, distinguish a model-serving measurement from a full agent-task measurement: the latter may include orchestration and tool interactions, while the inference metrics describe requests to the model service. State what is inside the timed path rather than treating the two scopes as interchangeable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose a useful operating point

Plot a user-facing latency metric against total system TPS, with each point labeled by concurrency. TTFT, ITL, end-to-end latency, or TPS per user can serve as the latency axis, depending on the question. Choose the point that meets the deployment’s latency budget, then report its throughput and load. The top TPS point alone is not necessarily useful if it comes with unacceptable latency.

Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

Normalize and compare without overstating the result

For each run, retain the total-system figure and configuration. If a per-GPU average helps the reader, show the arithmetic explicitly, for example: system TPS ÷ GPU count = average TPS per GPU. Do not omit GPU type, GPU count, or parallelism configuration next to the result, and do not describe the quotient as a single-GPU benchmark.

Before comparing GPUs or serving systems, align these dimensions or disclose the differences:

  • Model and model version, tokenizer, and generation settings.
  • GPU model and count, plus parallelism configuration.
  • Serving framework and version, precision or quantization, and batching settings.
  • Input/output length distributions, agent turn counts, and tool-use pattern.
  • Concurrency or request-arrival policy and measurement interval.
  • Latency metric, target, and whether the client’s network path is included.
  • Total-system throughput, any per-GPU calculation, and the load point at which each result was obtained.

Standardized evaluations and deployment-specific tests answer different questions. MLPerf offers standardized inference evaluations across model architectures and scenarios; a trace-based workload can better reflect a particular agent deployment. Neither a custom result nor a standardized submission should be generalized beyond its stated workload and system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, NVIDIA reports up to 3.7× higher throughput for Vera Rubin NVL72 than GB300 NVL72, and 99% scaling efficiency for a 288-GPU GB300 NVL72 submission, as vendor-reported MLPerf Inference v6.1 results retrieved from MLCommons on September 16, 2026. Those figures describe the named submitted systems and workloads; they are not a general GPU comparison or a conversion formula for an agent workload.

Keep backend metrics identifiable

AIPerf’s server-metrics reference maps throughput, latency, queue, and cache metrics across Dynamo, vLLM, SGLang, TensorRT-LLM, and Triton. When reporting server-side counters, preserve the backend-specific metric names and definitions. Similar labels do not establish that two backends measure the same interval or event.

A compact report format

A reproducible report can use this structure:

  • Workload: model/version and tokenizer; input/output length distributions; turns, context growth, and tool pattern; generation settings.
  • System: GPU model/count; serving engine/version; precision or quantization; parallelism and batching; client placement.
  • Procedure: warm-up configuration; concurrency or arrival policy; sweep values; measured interval; saved tool configuration and artifacts.
  • Results: total TPS and RPS; TTFT, ITL/TPOT, and end-to-end latency with stated summary statistics; concurrency for each point.
  • Normalization: if shown, total TPS divided by GPU count, explicitly labeled a per-GPU average and presented beside the total.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.