Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

AMD’s Ryzen AI 9 HX 375 did generate local-LLM tokens faster than Intel’s Core Ultra 7 258V in AMD’s published testing. AMD reported up to 27% higher token-generation performance, up to 50.7 tokens per second on Llama 3.2 1B Instruct, and up to 3.5× faster time to first token on larger models.

That is useful evidence for buyers interested in running models locally, but it is not a universal processor verdict. The October 2024 comparison came from AMD, used two specific laptops, paired a higher-tier 12-core/24-thread Ryzen chip against a lower-tier Core Ultra 7 258V, and did not include a directly comparable Intel Vulkan GPU-offload result.

Quick verdict

AMD has the stronger result in the specific comparison that matters here: five local language models tested in LM Studio 0.3.4 on Windows 11. AMD says the Ryzen AI 9 HX 375 led the Core Ultra 7 258V across all five models and achieved a peak advantage of up to 27% in tokens per second.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

However, the fairest conclusion is narrower: AMD’s benchmark shows that the HX 375 platform can be faster for the tested local-LLM workloads. It does not prove that every Ryzen AI 9 HX 375 laptop will outperform every Core Ultra 7 258V laptop. Memory configuration, cooling, sustained power, software backend, model size, quantization, and battery mode can all change the outcome.

#1 Best Overall
Sale
havit HV-F2056 Laptop Cooling Pad for 15.6-17 Inch Laptops, Black
  • Ultra-Portable: Slim, portable, and light weight allowing you to protect your investment wherever you go
  • Ergonomic Comfort: Doubles as an ergonomic stand with two adjustable height settings
  • Optimized for Laptop Carrying: The metal mesh provides your laptop with a stable laptop carrying surface
  • Ultra-Quiet Fans: Three ultra-quiet fans create a noise-free environment for you
  • Extra Usb Ports: Extra USB port and power switch design allows for connecting more USB devices. Warm Tips: The packaged cable is USB to USB connection. Type C connection devices need to prepare an Type C to USB adapter

For local AI experimentation, the AMD system is the more compelling choice when its laptop has adequate memory and cooling. For general laptop buying, compare the complete machine—not just the processor name.

AMD’s published benchmark is the primary source for the figures below. Tom’s Hardware also noted that the processors were not equivalent in product tier.

What “generates tokens faster” actually means

When a local chatbot answers, several stages affect how responsive it feels:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Prompt processing: The time required to ingest your question and any conversation history.
  • Time to first token: The delay before the first generated piece of text appears.
  • Tokens per second: The rate at which the model produces output after generation begins.
  • End-to-end response time: The combined effect of prompt processing, first-token latency, generation speed, and application overhead.

A high tokens-per-second number does not automatically mean a faster complete answer. Long prompts and large context windows can make prompt processing dominate the initial delay. AMD reported results for both output throughput and time to first token, but the figures should not be treated as a complete measure of every local-AI experience.

What AMD tested

AMD compared two retail laptop platforms using the following setup:

Rank #2
Sale
ChillCore Laptop Cooling Pad, RGB Lights Laptop Cooler 9 Fans for 15.6-19.3 Inch Laptops, Gaming Laptop Fan Cooling Pad with 8 Height Stands, 2 USB Ports - A21 Blue
  • 9 Super Cooling Fans: The 9-core laptop cooling pad can efficiently cool your laptop down, this laptop cooler has the air vent in the top and bottom of the case, you can set different modes for the cooling fans.
  • Ergonomic comfort: The gaming laptop cooling pad provides 8 heights adjustment to choose.You can adjust the suitable angle by your needs to relieve the fatigue of the back and neck effectively.
  • LCD Display: The LCD of cooler pad readout shows your current fan speed.simple and intuitive.you can easily control the RGB lights and fan speed by touching the buttons.
  • 10 RGB Light Modes: The RGB lights of the cooling laptop pad are pretty and it has many lighting options which can get you cool game atmosphere.you can press the botton 2-3 seconds to turn on/off the light.
  • Whisper Quiet: The 9 fans of the laptop cooling stand are all added with capacitor components to reduce working noise. the gaming laptop cooler is almost quiet enough not to notice even on max setting.
Component AMD system Intel system
Laptop HP OmniBook Ultra 14 ASUS Zenbook S14 UX5406SA
Processor Ryzen AI 9 HX 375 Core Ultra 7 258V
Memory 32 GB at 7,500 MT/s 32 GB at 8,533 MT/s
Operating system Windows 11 Pro 24H2 Windows 11 Pro 24H2
Virtualization-based security Enabled Enabled
LM Studio Version 0.3.4 Version 0.3.4
CPU threads used 12 8

The primary test used a fixed sample prompt, Q4_K_M quantization, and averages from three runs. The models were:

  • Meta Llama 3.2 1B Instruct
  • Meta Llama 3.2 3B Instruct
  • Microsoft Phi 3.1 4K Mini Instruct
  • Google Gemma 2 9B Instruct
  • Mistral Nemo 2407 13B Instruct

These are practical laptop-sized models, ranging from very small models to a 13B model. They do not represent every local-LLM workload. Larger models, longer contexts, different quantization formats, later LM Studio releases, and other llama.cpp backends may produce different results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The headline results

AMD reported three headline outcomes:

  • Up to 27% higher tokens-per-second performance for the Ryzen AI 9 HX 375 in the main comparison.
  • Up to 50.7 tokens per second on Llama 3.2 1B Instruct using 4-bit quantization.
  • Up to 3.5× faster time to first token on larger models.

The phrase “up to” matters. It identifies the best reported case, not an average advantage across all five models. The available material does not establish that the AMD system was 27% faster in every test, nor does it justify converting that peak number into a general performance uplift.

AMD also reported additional gains when using GPU offload and Variable Graphics Memory on its own system. In a separate comparison involving different software paths, AMD reported 8.7% higher performance in Phi 3.1 and 13% higher performance in Mistral 7B Instruct 0.3 against Intel AI Playground. Those figures should not be merged with the primary LM Studio comparison because the application and comparison method differed.

Why AMD may have led

The Ryzen AI 9 HX 375 is a substantial processor on paper. AMD lists it with:

Rank #3
Sale
Kootek Laptop Cooling Pad Cooler Stand with 5 Quiet Fans for 12"-17" Laptop
  • Whisper-Quiet Operation: Enjoy a noise-free and interference-free environment with super quiet fans, allowing you to focus on your work or entertainment without distractions.
  • Enhanced Cooling Performance: The laptop cooling pad features 5 built-in fans (big fan: 4.72-inch, small fans: 2.76-inch), all with blue LEDs. 2 On/Off switches enable simultaneous control of all 5 fans and LEDs. Simply press the switch to select 1 fan working, 4 fans working, or all 5 working together.
  • Dual USB Hub: With a built-in dual USB hub, the laptop fan enables you to connect additional USB devices to your laptop, providing extra connectivity options for your peripherals. Warm tips: The packaged cable is a USB-to-USB connection. Type C connection devices require a Type C to USB adapter.
  • Ergonomic Design: The laptop cooling stand also serves as an ergonomic stand, offering 6 adjustable height settings that enable you to customize the angle for optimal comfort during gaming, movie watching, or working for extended periods. Ideal gift for both the back-to-school season and Father's Day.
  • Secure and Universal Compatibility: Designed with 2 stoppers on the front surface, this laptop cooler prevents laptops from slipping and keeps 12-17 inch laptops—including Apple Macbook Pro Air, HP, Alienware, Dell, ASUS, and more—cool and secure during use.
  • 12 cores and 24 threads, using four Zen 5 cores and eight Zen 5c cores
  • Boost speeds up to 5.1 GHz
  • A default 28 W TDP and configurable 15–54 W range
  • Radeon 890M integrated graphics with 16 graphics cores
  • Up to 55 NPU TOPS and up to 85 total platform TOPS
  • Support for LPDDR5X memory up to 8,000 MT/s as specified by AMD

By comparison, the Intel laptop used a Core Ultra 7 258V with eight CPU threads selected for the test. The HX 375 therefore had more available CPU threads, and it is a higher-tier part than the 258V. That product-tier mismatch is one reason Tom’s Hardware described the matchup as less than equal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Other factors may also have contributed:

  • Different CPU designs and compiler optimizations
  • Different power limits and cooling systems
  • Differences in LM Studio or llama.cpp runtime behavior
  • Memory bandwidth and latency
  • Integrated-GPU and Vulkan implementation differences
  • The AMD system’s use of 12 threads versus eight on the Intel system

The published benchmark cannot isolate the contribution of each factor. It is best understood as a comparison of two complete laptop platforms, not a controlled processor-only experiment.

The fairness question: AMD won, but the test has important limits

The comparison has strengths. Both systems used Windows 11 Pro 24H2, 32 GB of memory, enabled VBS, the same LM Studio version, the same model family of tests, and a stated three-run averaging method. It is also relevant to buyers choosing between actual laptops.

But it was not perfectly balanced:

  • The processor tiers differed. AMD’s premium HX 375 was compared with Intel’s Core Ultra 7 258V, not Intel’s highest-tier Lunar Lake processor.
  • AMD supplied the testing. The result is vendor-reported and has not been established as an independent, multi-sample benchmark by the cited material.
  • Memory speeds differed. The Intel system used faster memory: 8,533 MT/s versus 7,500 MT/s on AMD.
  • Thread settings differed. AMD used 12 CPU threads, while Intel used eight, reflecting the systems’ configurations but affecting reproducibility.
  • GPU-offload treatment was asymmetric. AMD did not include Intel’s Vulkan GPU-offload result in the primary comparison because that result reportedly performed worse than Intel’s CPU-only mode.

The memory point is particularly important. Local LLM generation can be bandwidth-sensitive, especially with larger models. Integrated-GPU inference also shares system memory. The fact that AMD still led despite slower memory is notable, but it does not remove the need to disclose the difference.

Nor does excluding the Intel Vulkan result prove that Intel hardware is incapable of effective GPU-accelerated inference. It means that AMD’s selected LM Studio/Vulkan test did not provide a directly comparable Intel GPU-offload result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Razer Laptop Cooling Pad Adaptive Smart, Intelligent Fan Control
  • SMART COOLING — From idle to full load, keep the laptop running smoothly with our first laptop cooling pad that changes fan speeds automatically to manage system temperatures based on the settings
  • AIRTIGHT PRESSURE CHAMBER — Included foam seals ensure no cool air leakage and works in tandem with a long lifespan 140 mm brushless fan that spins up to 3000 RPM to significantly reduce CPU, GPU, and surface temperatures
  • WORKS WITH MOST LAPTOPS — Whether you've got an ultra-portable 14″ laptop or an 18″ powerhouse, choose between three magnetic frames that maximize cool air pressure and circulation
  • PRESET & CUSTOM FAN CURVES — Keep the system cool in any scenario with our recommended presets or calibrate the fan to adjust for noise level or desired internal temperature via Razer Synapse
  • 3-PORT USB TYPE A HUB — From webcams to controllers to drawing tablets, plug in more devices to the laptop without solely relying on its native USB ports

CPU, GPU, and NPU: which part is doing the work?

“AI PC” specifications can be confusing because local inference may use several compute paths:

  • CPU-only inference is broadly compatible and often the fallback for local models.
  • Integrated-GPU inference can improve throughput when the application and backend support Vulkan, DirectML, or another suitable path.
  • NPU inference can be efficient for supported AI operations, but only when the application, model format, drivers, and operators work with the NPU backend.
  • Hybrid inference can divide work across CPU, GPU, and other accelerators.

AMD’s test was principally a CPU and LM Studio comparison, with separate GPU-offload testing. It should not be described as proof that the HX 375’s NPU generated the measured tokens.

AMD lists the HX 375’s NPU at up to 55 TOPS and the total platform at up to 85 TOPS. Those numbers describe theoretical AI-compute capability under specified conditions; they do not translate directly into tokens per second. TOPS is not a substitute for a measured local-inference benchmark.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the result means for common use cases

Small chatbots and quick questions

For 1B to 3B quantized models, both systems should be viable local-chat machines. AMD’s reported 50.7 tokens per second on Llama 3.2 1B suggests a particularly fluid experience on the tested HX 375 laptop, but the practical difference may be less important if both systems already respond comfortably.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Coding assistants

Coding workloads can involve long prompts, large conversation histories, and repeated short responses. Token-generation speed matters, but prompt-processing speed and context capacity may matter just as much. A 32 GB laptop is preferable to a 16 GB configuration when the editor, browser, operating system, and model must coexist.

Best Value
Sale
llano V12 Gaming Laptop Cooling Pad Laptop Cooler Laptop Cooling Fan Stand
  • New Upgraded Version-Cooling Gets Quiet and Quicker: Equipped with 5.5-inch large diameter turbo booster fan, combined sealed foam, ensures perfect cooling effect, 360 degrees all-round dynamic cooling, your laptop can reduce temperature by 44°C in 90 seconds (CPU+GPU), even during 4K rendering or AAA gaming, operating noise ≤70dB
  • 3-Port USB Hub and Precise Control: The V12 laptop cooler features three USB 2.0 ports, turning your cooling pad into a central workstation hub. This solves the problem of limited ports on modern laptops, allowing you to connect a high-speed mouse, keyboard, and hard drive simultaneously. Meanwhile, the scroll wheel allows for easy, instantaneous, and precise airflow adjustment. ⚠️ For peripherals only – NOT for charging devices
  • Integrated Dust Filtration Extends Laptop Lifespan: This cooler fan features a high-density, removable dust filter that effectively protects the internal fans of laptops. It captures hair and debris, preventing them from entering the vents, thus solving the common "pressure cooker" dust buildup problem, significantly extending the lifespan of your expensive gaming PC and reducing costly professional cleaning
  • Soothing Controllable RGB: RGB light bar on the laptop cooling pad for PC, with 10 modes and 4 light colors collection, matches your PC gears accessories for amazing synergy even in dim room. Intuitive Touch-Mute Button adjusts RGB lighting with a single finger, minimizing distractions. Configured memory function, the laptop cooler RGB eliminates repeated selections and brings itself alive when power on
  • User-Friendly Design: Featuring a reinforced chassis design, it's suitable for heavy-duty laptops from 15.6 to 19 inches. Three adjustable tilt angles (3°/12°/15°) allow you to customize your viewing height (scenario), directly alleviating neck and shoulder fatigue during long gaming or work sessions (pain point solved), ensuring maximum comfort and better posture

7B to 13B models

These models place greater pressure on memory capacity and bandwidth. Q4 quantization reduces the memory requirement, but model files, runtime overhead, context, and other applications still consume RAM. AMD’s lead in the cited test is relevant, but sustained performance will depend heavily on the laptop’s power and cooling limits.

Long-context summarization

Long prompts can make time to first token and prompt processing more important than the eventual output rate. A benchmark based on a fixed sample prompt cannot predict every long-context workflow.

Battery-powered use

Performance can fall away from AC power as laptops reduce processor power limits. A machine that wins a plugged-in benchmark may not maintain the same advantage on battery. Firmware settings, fan profiles, and thermal limits should be checked in independent reviews of the exact laptop.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quiet office use

The HX 375 can be configured across a broad 15–54 W range. A thin implementation running quietly at a low sustained power limit may not match a short benchmark burst from a better-cooled configuration. Higher throughput can also mean more fan noise.

Buying advice: compare the laptop, not the badge

  1. Choose at least 32 GB for serious experimentation. Sixteen gigabytes may run small quantized models, but leaves less room for Windows, a browser, an editor, and long contexts.
  2. Check the exact memory configuration. Laptop RAM is commonly soldered. Capacity and speed may be impossible to upgrade later.
  3. Inspect sustained power and cooling. The HX 375’s advertised maximum boost and configurable power range do not guarantee identical performance across laptops.
  4. Confirm the backend. CPU, Vulkan, DirectML, Intel-specific, and other paths can deliver different results or have different model compatibility.
  5. Check software versions. AMD’s comparison used LM Studio 0.3.4. Later LM Studio, llama.cpp, driver, and firmware updates may change performance.
  6. Consider model size first. If your real goal is running much larger models, more RAM, dedicated VRAM, or a desktop may matter more than a modest CPU advantage.
  7. Account for storage. Local model files can occupy many gigabytes, particularly when keeping multiple quantizations.
  8. Do not pay a large premium for one peak figure. AMD’s 27% result is useful evidence, but it came from a narrow vendor-supplied comparison.

The HP OmniBook Ultra 14 is the AMD test platform named in the benchmark, while the ASUS Zenbook S14 UX5406SA is the Intel platform. Regional SKUs can differ in memory, firmware, display, power behavior, and availability, so the model name alone is not enough.

LM Studio is the most directly relevant application for readers who want to reproduce the published software path, but current results should not be assumed identical to AMD’s LM Studio 0.3.4 measurements.

Final verdict

AMD wins the cited local-LLM test. The Ryzen AI 9 HX 375 generated tokens faster than the Core Ultra 7 258V in AMD’s October 2024 LM Studio comparison, with a reported peak advantage of up to 27%, a result of up to 50.7 tokens per second on Llama 3.2 1B, and up to 3.5× faster time to first token on larger models.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

But the evidence supports a platform-specific conclusion, not a universal ranking. The AMD chip had more cores and threads and occupied a higher product tier; the Intel system had faster memory; the GPU-offload comparison was incomplete; and the test came from AMD rather than an independent lab. Buyers should treat the result as a strong indication that the HX 375 is attractive for local AI—not as a guarantee that every HX 375 laptop will beat every Core Ultra 7 258V machine.

Quick Recap

SaleBestseller No. 1
havit HV-F2056 Laptop Cooling Pad for 15.6-17 Inch Laptops, Black
havit HV-F2056 Laptop Cooling Pad for 15.6-17 Inch Laptops, Black
Ergonomic Comfort: Doubles as an ergonomic stand with two adjustable height settings; Ultra-Quiet Fans: Three ultra-quiet fans create a noise-free environment for you
$27.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.