Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

MLCommons released MLPerf Inference v5.0 on April 2, 2025, with 17,457 performance results from 23 organizations. The round added tests for large language models, graph neural networks and automotive perception, and drew particular attention to NVIDIA Hopper and Blackwell systems, AMD Instinct MI325X accelerators and Intel CPU-only inference. These are historical results: MLPerf has since published v5.1 and v6.0.

The headline is not a universal ranking of AI hardware. Each score belongs to a specific model, scenario, accuracy target and system configuration. For a useful comparison, match those details before comparing numbers.

# Preview Product Price
1 MX3 M.2 AI Accelerator MX3 M.2 AI Accelerator $169.00

What MLPerf Inference v5.0 measured

MLPerf Inference is a system-level benchmark: it measures how quickly a complete system processes inputs and produces outputs using specified trained models. Results can reflect the accelerator, host CPU and memory, interconnect, software and kernels, precision, batching, serving implementation, power configuration and accuracy target—not just peak chip performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MLCommons describes the suite as an architecture-neutral way to produce reproducible performance information across datacenter and edge systems. That makes the results useful evidence, but not a promise that a different model or production service will perform the same way. MLCommons’ v5.0 announcement reported 17,457 results submitted by 23 organizations.

#1 Best Overall
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

The release followed a submission deadline of February 28, 2025. Its results tables are organized by workload and scenario rather than by a single overall winner. The official v5.0 comparison tables let readers inspect individual system submissions.

Four additions in v5.0

V5.0 introduced four benchmark workloads or variants:

  • Llama 3.1 405B Instruct: A test of systems serving a very large, 405-billion-parameter language model.
  • Llama 2 70B Interactive: A more responsiveness-focused language-model scenario, with constraints that include time to first token (TTFT) and time per output token (TPOT).
  • RGAT: A graph neural network workload based on the Illinois Graph Benchmark Heterogeneous dataset. MLCommons describes the dataset as containing 547,306,935 nodes and 5,812,005,639 edges. MLCommons’ RGAT overview explains the benchmark.
  • Automotive PointPainting: An edge workload for 3D object detection using camera and lidar-related processing. MLCommons’ PointPainting overview covers its automotive context.

The suite also included established workloads such as ResNet50, RetinaNet, BERT, DLRM-v2, 3D-Unet, GPT-J, Stable Diffusion XL, Llama 2 70B and Mixtral-8x7B. See the v5.0 benchmark documentation for workload definitions and submission details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why Llama 2 70B attracted attention

MLCommons said Llama 2 70B had the highest submission rate of any workload in this round, overtaking ResNet50. Compared with a year earlier, submissions increased 2.5 times, the median submitted score doubled, and the best score was 3.3 times faster than in Inference v4.0. These figures describe changes across benchmark rounds; they do not mean every deployment will achieve those gains.

It is important to distinguish the scenarios. Offline measures throughput in a batch-oriented setting with less restrictive latency demands. Server measures throughput while meeting a service latency constraint. Interactive places greater emphasis on responsiveness, including how soon the first token arrives and how quickly subsequent tokens are produced. A high offline throughput score alone cannot tell you whether a chatbot will feel responsive.

The rise of Llama 2 70B submissions indicates where benchmark participants were directing optimization effort. It does not establish that this model—or generative AI generally—is the most important workload for every buyer.

What ServeTheHome highlighted

ServeTheHome’s April 2, 2025 coverage focused on the visible competition among NVIDIA, AMD and Intel. It described a round with many NVIDIA submissions, prominent Hopper systems including H200, emerging Blackwell results involving B200 and GB200, and Grace-based platforms. It also noted AMD MI325X results, Intel Xeon submissions focused heavily on CPU-only inference, and Google TPU Trillium’s appearance in the official results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That is useful platform context, but “NVIDIA dominated” should be read as a description of submission volume and visibility—not as a claim that NVIDIA won every workload. Nor is there one score for a vendor: GPU generation, accelerator count, host processor, node count, topology, power and software can all differ between submitted systems.

Vendor or platform What appeared in the round How to interpret it
NVIDIA Hopper systems including H200; Blackwell B200 and GB200 results; Grace-based systems. Compare a specific system and workload. These platforms differ in generation, accelerator count, memory, host architecture, interconnect and power envelope.
AMD Instinct MI325X results, including single-node and multi-node submissions. ServeTheHome described some results as in the general performance range of H200 systems for particular comparisons. That is not enough to claim a blanket match: workload, scenario, accuracy, precision and system size must align.
Intel Xeon 6980P and Xeon 6700P-family submissions, with an emphasis on CPU-only inference. These results address a different deployment choice from a multi-GPU server. Intel’s “only server CPU on MLPerf” wording is best understood as referring to CPU-only submissions, not as saying systems with AMD or NVIDIA server CPUs were absent.
Google TPU Trillium, also identified as TPU v6e, was among the newly represented processors. Its appearance does not imply broad coverage across every workload. Check the tables for the particular test and configuration.

MLCommons also listed six newly available or soon-to-ship processors represented in the round: AMD Instinct MI325X, Intel Xeon 6980P, Google TPU Trillium, NVIDIA B200, NVIDIA Jetson AGX Thor 128 and NVIDIA GB200. Their inclusion is not a statement about current availability, pricing or suitability for a specific deployment.

Why CPU-only inference still matters

CPU inference can be relevant for smaller models, cost-sensitive services, existing server fleets, edge or data-sovereignty requirements, and workloads that do not use an accelerator efficiently enough to justify buying one. Those are deployment considerations, not proof that CPU and GPU scores are interchangeable. Compare systems against the same application requirements, including latency, throughput, power and scale.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Datacenter, edge and accuracy are not interchangeable

The v5.0 suite separates datacenter and edge categories. The documentation says all v5.0 benchmarks except BERT apply to the datacenter category; the edge category excludes DLRM-v2, Llama 2 70B, Mixtral-8x7B and RGAT. An edge result should not be ranked directly against a datacenter result: power limits, memory, latency, form factor, thermal conditions, connectivity and real-time requirements can differ.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some workloads have normal and high-accuracy variants: BERT, Llama 2 70B, GPT-J, DLRM-v2 and 3D-Unet. The documented accuracy thresholds are at least 99% of reference-model accuracy for the default requirement and at least 99.9% for high accuracy. Because accuracy requirements can affect speed, compare like-for-like variants, not just the largest throughput number. See the benchmark definitions and category rules for details.

DeepSeek-R1 was not a v5.0 benchmark

ServeTheHome noted that NVIDIA and AMD discussed DeepSeek-R1 performance in related vendor material. Those figures were not official MLPerf Inference v5.0 results. Treat them as separate vendor-provided claims, not as scores in the v5.0 tables. They may use different precision, model implementations, prompt and output lengths, or latency targets. In particular, a vendor’s DeepSeek-R1 FP4 figure is not directly comparable with an official MLPerf Llama 2 70B or Llama 3.1 405B result.

How to compare two v5.0 results

Before drawing a conclusion, check the following in the official result tables:

  1. Version and workload: Compare v5.0 with v5.0, and the same benchmark with the same benchmark.
  2. Category and scenario: Keep datacenter and edge results separate; match server, offline or interactive scenarios.
  3. Accuracy: Match the normal or high-accuracy variant.
  4. System configuration: Note system name, submitter, accelerator model and count, host CPU, memory, number of nodes and interconnect where listed.
  5. Power and scale: Check power results where available, and consider whether the system size fits your intended deployment. A large multi-node score may require more capital, power, networking and operational effort than a smaller service needs.
  6. Submission status: Distinguish available results from preview or otherwise qualified entries. Published results can change; MLCommons maintains a results change log, which records subsequent invalidations of some v5.0 preview results that did not receive required validation submissions.

Even a like-for-like benchmark comparison does not tell you which system is cheaper or easier to operate. MLPerf scores do not by themselves establish purchase price, cloud hourly cost, delivery time, support cost or total cost of ownership.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the results can—and cannot—tell a buyer

Use v5.0 to narrow questions, not to skip workload testing. Start with your production model and serving pattern: request volume, concurrency, prompt and output lengths, latency objectives, precision and accuracy needs. Then look for a benchmark that resembles that workload and compare systems under matching conditions. A result on a 405-billion-parameter language model says little by itself about a small vision model, a recommendation system or a custom fine-tuned model.

Also account for software. Benchmark scores reflect the submitted software stack and optimization choices; your inference engine, kernels, quantization, batching policy and model implementation may differ. A throughput-focused configuration can be a poor fit if your priority is fast first-token response. A large system can deliver excellent aggregate throughput but be uneconomical or unnecessarily complex for a smaller service.

Finally, v5.0 is a dated snapshot, not the current benchmark round. MLCommons subsequently released Inference v5.1 and v6.0. For current comparisons, consult the MLPerf Inference release archive and use v5.0 when you specifically need to understand the April 2025 results.

Quick Recap

Bestseller No. 1
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.