October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

Trace One Tensor from Model Math to LLM Serving Cost

A tensor's shape and FLOP count do not determine its serving cost. Follow a decoder activation through data movement, GPU execution, KV-cache capacity, deployment, and workload-based cost accounting.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A tensor has no fixed serving cost. Its impact depends on the operations applied to it, the bytes moved, how those operations run on particular GPUs, and how the serving system schedules requests. Here is an illustrative trace of one activation through a decoder-only Transformer, from matrix multiplication to memory capacity and cost accounting.

What tensor are we tracing?

Use a hypothetical decoder-only Transformer with hidden width 4,096, 32 attention heads, and 32 layers. Assume BF16 activations and weights, a batch of one prompt containing 512 tokens, and no grouped-query attention. These are illustrative dimensions—not a specification for a particular model.

As an Amazon Associate I earn from qualifying purchases.

At one layer, let X be the residual-stream activation entering the query projection. Its shape is [batch, sequence, hidden], or [1, 512, 4096]. At two bytes per BF16 value, it occupies 4 MiB. The mathematical query projection is Q = XW, where W has shape [4096, 4096]; the result Q has shape [1, 512, 4096], also 4 MiB.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For each output value, the projection combines 4,096 input values with 4,096 weights. Across all 512 positions and 4,096 output channels, that is 8,589,934,592 multiply-accumulates. If one multiply-add is counted as two FLOPs, the operation is about 17.18 GFLOPs. That counting convention is described in NVIDIA’s GPU Performance Background User’s Guide; it is a way to count mathematical work, not a runtime prediction.

#1 Best Overall
MINISFORUM MS-02 Ultra Workstation Mini PC, Intel Core Ultra 9 285HX (24C/24T, up to 5.5GHz), PCIe 5.0 x16, 32GB RAM 1TB SSD,USB4 v2 80Gbps, Dual 25GbE+10GbE+2.5GbE, Wi-Fi 7, 350W PSU
  • High-Performance AI Processor:The MS-02 Ultra features an Intel Core Ultra 9 285HX (24C/24T, up to 5.5 GHz, 13 TOPS NPU), delivering fast and efficient performance for AI inference, algorithm development, and media workloads. A PCIe x16 expansion slot supports desktop-class GPU upgrades for advanced model training and accelerated computing tasks. It's ideal for creators, engineers, and teams handling intensive parallel workloads.
  • 4 × M.2 PCIe 4.0 + 4 × DDR5 SODIMM slots:Four DDR5 SODIMM slots support up to 256 GB of memory, while ECC helps maintain data integrity in mission-critical environments. Four PCIe 4.0 M.2 slots support up to 24 TB of storage, supporting RAID 0/1/5/10, combining high-speed performance with data protection. It allows for the creation of independent scratch disks, media libraries, and project drives, providing high-throughput for production workflows.
  • PCIe & USB 4.0 v2: Up to three PCIe slots can be equipped, including a dual-slot x16 GPU. The main slot supports PCIe 5.0, meeting the needs of high-bandwidth creative and computing workloads. USB 4.0 v2 (80Gbps) supports high-bandwidth external storage and displays.
  • Ultra-fast Networking: Wi-Fi 7 further enhances wireless performance with next-generation speeds and low-latency stability. Intelligent bandwidth switching optimizes throughput in different network environments, ensuring optimal performance for enterprise or local networks. Dual 25GbE ports (providing up to approximately 3.125 GB/s bandwidth, about 25 times faster than traditional 1GbE), enabling seamless large-scale file transfers and parallel computing. 10GbE and 2.5GbE ports, with support for Intel vPro technology, ensure enterprise-grade remote management and deployment flexibility.
  • Server-grade thermal architecture: Utilizing a dedicated CPU/GPU airflow design, equipped with a 6-pipe dual-fan cooler, it maintains stable performance even under sustained loads, delivering up to 140W Turbo power while maintaining a 100W TDP, and operating with noise levels as low as 36 dB. An integrated 350W power supply ensures stable and reliable output for demanding computing tasks and fully loaded extended configurations.

How much data moves, and what does arithmetic intensity tell us?

A simple lower-bound traffic estimate for this prompt projection counts one read of the 32 MiB weight matrix, one read of the 4 MiB input, and one write of the 4 MiB output: 40 MiB total. Under those assumptions, 17.18 GFLOPs divided by 40 MiB is roughly 410 FLOPs per byte. This estimate assumes weights are fetched once and does not include extra traffic from intermediate tensors, padding, cache behavior, or other layers and operators.

For one decode position instead of 512 prompt positions, the same projection performs about 33.55 MFLOPs. If the weight matrix must be read and the 8 KiB input and 8 KiB output are read or written, traffic is about 32 MiB; the ratio is roughly 1 FLOP per byte. The weight matrix dominates this small operation, while the 512-position case reuses those weights across many more rows. Actual traffic depends on whether data is cached or reused and on the chosen implementation.

Arithmetic intensity—operations per byte moved—helps explain whether an operation is more likely limited by math throughput or memory bandwidth. NVIDIA’s guide frames performance limits as math bandwidth, memory bandwidth, or latency. Its V100-era FP16 examples classify a linear layer with 4,096 outputs and 1,024 inputs as arithmetic-limited at batch 512 (315 FLOPs/byte) and memory-limited at batch 1 (1 FLOP/byte). Those examples illustrate the batch-size effect under their stated assumptions; they are not predictions for every GPU or model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How does the framework turn this operation into GPU work?

The equation Q = XW describes the model math. It does not dictate a single GPU kernel. A framework may lower the operation to a matrix-multiplication kernel, combine it with neighboring operations through fusion, or compile a larger region. The resulting performance depends on the GPU, data types, matrix sizes, kernel implementation, and how much work is available to run in parallel.

Rank #3
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
  • Launch and latency: For small workloads, the time to launch work and other fixed overheads can matter as much as the arithmetic.
  • Parallelism and occupancy: A kernel needs enough independent work to keep GPU execution units busy. Small batches or awkward dimensions can leave capacity unused or create tail effects.
  • Fusion and compiler boundaries: Combining operations may avoid extra launches or intermediate reads and writes. But graph breaks can limit what a compiler can optimize together. PyTorch’s Llama 2 inference report describes graph breaks associated with unsupported operations and distributed collectives: PyTorch’s report.
  • Communication: When work spans devices, moving intermediate results between them adds time that a FLOP count does not capture.

Why are prompt prefill and token decode different workloads?

Prompt prefill

During prefill, the model processes the prompt’s existing tokens. In this example, the query projection handles 512 positions together, allowing each weight to be reused across those rows. The arithmetic-intensity estimate above is for that projection alone; attention and other layer operations also contribute to prefill time and memory traffic.

Autoregressive decode

During decode, the model generates tokens sequentially: each new token depends on the preceding context. Rather than recomputing keys and values for all earlier tokens at every step, inference commonly keeps them in a key/value (KV) cache. In this example, the cache for one sequence at 512 positions is approximately 8 MiB per layer: 512 positions × 4,096 values for each of K and V × 2 bytes. Across 32 layers that is about 256 MiB per sequence, before allocator overhead or implementation-specific storage. At 4,096 positions, the same assumptions imply about 2 GiB per sequence.

Rank #4
Sale
Apple 2026 MacBook Pro Laptop with Apple M5 Max chip with 18-core CPU and 40-core GPU: Built for AI, 16.2-inch Liquid Retina XDR Display, 48GB Unified Memory, 2TB SSD, Wi-Fi 7; Silver
  • FAST RUNS IN THE FAMILY — The 16-inch MacBook Pro with the M5 Pro or M5 Max chip brings next-generation speed and powerful on-device AI to personal, professional, and creative tasks. With all-day battery life, double the starting storage,* and a breathtaking Liquid Retina XDR display, it’s pro in every way.*
  • BUCKLE UP — Along with a next-generation CPU, faster unified memory, and up to 2x faster SSD storage,* M5 Pro and M5 Max feature a more powerful GPU with a Neural Accelerator built into each core, delivering faster AI performance and on-device training capabilities. So you can blaze through demanding workloads at mind-bending speeds.
  • BUILT FOR AI — Apple silicon, and every major component that powers it, is designed to run demanding on-device AI workloads like LLM inference and training. And Apple Intelligence helps you write, express yourself, and get things done effortlessly with groundbreaking privacy protections at every step.*
  • ALL-DAY BATTERY LIFE — MacBook Pro delivers the same exceptional performance whether it’s running on battery or plugged in.*
  • MACOS RUNS APPS FAST — All your go-to apps run lightning fast in macOS, including built-in apps like FaceTime and Messages. Plus, built-in virus protection and free software updates help keep your Mac running smoothly and securely.

Prompt length, generated length, batch size, and concurrency therefore change both the amount of work and the active cache. Variable sequence dimensions also complicate execution. PyTorch/XLA describes bucketing or padding prompts and using fixed-shape KV-cache updates to manage dynamic shapes: PyTorch/XLA’s inference report. These are implementation techniques, not requirements for every serving stack.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A published result belongs to its full setup. PyTorch and IBM Research contributors reported 29 ms/token in 2023 for a single-user Llama 2 70B configuration on eight NVIDIA A100 GPUs; the reported experiment used a 512-token input and generated 50 tokens. That is a configuration-specific report, not a speed guarantee for another model, batch, hardware setup, or serving target (report details).

Best Value
MINISFORUM MS-S1 MAX Mini AI Workstation PC, AMD Ryzen AI Max+ 395 (16C/32T),RDNA3.5 GPU,128GB LPDDR5x RAM 2TB SSMINI PC, Dual M.2 PCIe 4.0,PCIe x16 Slot, USB4 V2(80Gbps)& Dual 10GbE, 320W PSU,Wi-Fi 7
  • 【High-Performance APU】The MS-S1 MAX features an AMD Ryzen AI Max+ 395 APU, integrating a Zen 5 architecture CPU (up to 5.1GHz, 16C/32T, 64M L3 Cache), an RDNA 3.5 GPU, and an NPU (50 TOPS). The total system output is 126 TOPS. It provides powerful parallel computing capabilities for demanding AI workflows. It is ideal for running local LLMs, multimodal models, and computationally intensive tasks
  • 【128GB UMA Memory】Equipped with up to 128GB of LPDDR5x-8000MT/s unified memory, it enables the CPU and GPU to access a shared, high-bandwidth memory pool with extremely low latency. Ideal for large-scale AI inference, 3D workloads, and complex timelines in video editing. It eliminates traditional VRAM bottlenecks, ensuring smoother data transfer during high-intensity computations. The UMA design maximizes performance stability under high loads
  • 【Flexible Expansion】The MS-S1 MAX features USB4 V2 (up to 80Gbps), dual 10GbE LAN, HDMI 2.1 (up to 8K60), a full-length PCIe x16 expansion slot, and dual M.2 slots supporting up to 16TB RAID 0/1. Wi-Fi 7 provides stronger signal coverage and a more stable wireless experience. The slide-out design facilitates upgrades and maintenance. It easily adapts to personal, studio, or rack-mount enterprise environments
  • 【High-Efficiency Cooling System】Utilizing an aerospace-grade aluminum alloy chassis, copper base plate, six heat pipes, dual turbine fans, and advanced PCM thermal conductive material, it maintains stable cooling performance even under continuous load. This system supports 130W continuous power and 160W peak power operation, with a built-in 320W power supply. It boasts multiple global certifications including CCC, FCC, UL, CE, and UKCA, ensuring stable and reliable operation in various environments
  • 【Cluster Design】Two MS-S1 MAX units can be configured as a dual-unit cluster to run a large 235B Q4 model locally, achieving an output speed of 10.87 tok/s. Supporting 2U rack deployment, multiple MS-S1 MAX units can be cascaded into a distributed cluster to create a high-efficiency AI computing center. A cluster of four MS-S1 MAX units successfully ran a DeepSeek-R1 671B Q4 large model. A reserved cluster power-on interface allows for unified start-up and shutdown
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What changes when the model is deployed across GPUs?

First account for model weights, runtime allocations, and the KV cache for the active workload. A model may fit by weights alone yet lack room for the cache and concurrent requests. For the illustrative model, the 32-layer KV-cache estimate is 256 MiB per 512-position sequence; real models can differ substantially because of width, layer count, cache dtype, and attention layout.

If a model or its working set does not fit on one GPU, deployment may distribute work. Tensor parallelism splits operations across GPUs, commonly within a node; pipeline parallelism assigns different layers to different stages and can span nodes. Both introduce communication and topology considerations. vLLM’s deployment guidance discusses parallelism choices, and its logs can report KV-cache token capacity and an estimated maximum concurrency. Those estimates help assess fit; they are not a monetary bill or a guarantee of achieved throughput: vLLM parallelism and scaling.

Deployment choice What it distributes What to account for
One GPU Weights and execution stay on one device if they fit. Usable memory for weights, active KV cache, runtime needs, and the desired request concurrency.
Tensor parallelism Parts of operations are spread across GPUs. GPU topology and communication between devices, alongside memory and workload fit.
Pipeline parallelism Layers are assigned to different stages or devices. Communication between stages and how the request workload uses those stages.

Compare deployment options at the workload you expect to serve—not from peak FLOPs alone. Relevant measures include time to first token (TTFT), inter-token latency, throughput at target concurrency, usable memory and KV-cache headroom, GPU count and interconnect, and quality constraints imposed by numeric format. A system that looks faster at one batch size or prompt length may not be the better fit at another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When can you turn tensor work into serving cost?

Only after defining the workload and the cost basis. A FLOP estimate does not supply a price, and there is no general dollar cost per token established here. For a provider, use the applicable dated machine rate; for owned hardware, use a stated amortization method and period. Then relate that cost to useful completed work under the service-level objective.

  • Workload: model, input and output length distributions, batch size, concurrency, and request mix.
  • Capacity and performance: GPU configuration, memory headroom, utilization, throughput, TTFT, and inter-token latency.
  • Cost basis: machine or internal cost, billing unit, time period, and how idle capacity is treated.
  • Unit of comparison: cost per request or per generated token, with the workload and latency target stated.

Use measured serving data for that configuration to calculate cost per useful request or token. If utilization falls, fixed machine cost is spread over fewer completed requests; if longer contexts or concurrency force more GPUs, the cost basis changes. The model equation alone cannot settle either outcome.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.