Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Yes—some GeForce RTX 5090 and NVIDIA RTX PRO 6000 Blackwell systems have reportedly failed to reset after GPU passthrough workloads. The problem is associated with KVM/QEMU and VFIO, not ordinary gaming. After a virtual machine shuts down, reboots, or releases the GPU, the card may fail PCIe Function-Level Reset (FLR) and remain unusable until the host is rebooted.
The short version
- Affected context: KVM/QEMU virtual machines using VFIO PCI passthrough.
- Reported GPUs: GeForce RTX 5090 and RTX PRO 6000 Blackwell.
- Trigger: VM shutdown, reboot, forced stop, startup, or GPU reassignment.
- Typical symptom: VFIO cannot complete the PCIe reset, leaving the device inaccessible.
- Recovery: Often a host reboot; more severe, separate GPU failures can require a full power cycle.
- Status: CloudRift and community reports say NVIDIA acknowledged or reproduced the issue, but the public material reviewed does not establish a universal official fix.
This is not evidence that every RTX 5090 or RTX PRO 6000 is defective. It is best described as a reported reset or reinitialization failure in certain Blackwell GPU-passthrough configurations.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card | $1,772.53 | Buy on Amazon |
| 2 |
|
ASUS TUF Gaming GeForce RTX 5090 32GB GDDR7 OC Edition Gaming Graphics Card | $6,899.99 | Buy on Amazon |
What the bug does
In a typical passthrough setup, the host assigns a physical GPU to one guest VM through VFIO. When the guest stops, the hypervisor must reset the card so it can be safely reused. A successful PCIe FLR returns the device to a clean state without restarting the host.
In affected systems, that reset may never complete. CloudRift documented errors including:
#1 Best Overall
- Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
- Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
- Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
- 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
- Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
vfio-pci: not ready 65535ms after FLR; giving up
Other reported symptoms include:
libvirt: error : internal error: Unknown PCI header type '127'
and failures involving PCIe link retraining or a transition from the low-power D3cold state back to the active D0 state. The host may remain online, but VFIO, QEMU, libvirt, or the next VM cannot safely use the GPU.
That is why descriptions such as “the GPU is bricked” are too broad. In many reports, rebooting the host restores the card. The practical problem is that a reset failure can interrupt unrelated workloads and defeat unattended VM recycling.
What CloudRift reported
CloudRift described production systems in which RTX 5090 and RTX PRO 6000 cards became unresponsive after VM use or during VM startup and shutdown. The company reported that affected nodes required a complete host reboot before the GPU could be reassigned and offered a $1,000 bug bounty while investigating the failure.
Its comparison testing reportedly did not reproduce the same behavior on H100, B200, or RTX 4090 systems. That is useful comparative evidence, but it does not prove those models are immune to every reset problem. CloudRift’s conclusions also were not an independently audited population-wide failure study. Read CloudRift’s original report.
Tom’s Hardware separately reported the passthrough failure and the requirement for a host reboot.
Which products are involved?
The strongest public reports concern:
- GeForce RTX 5090.
- RTX PRO 6000 Blackwell, including workstation-oriented configurations.
NVIDIA’s usual branding is RTX PRO 6000 Blackwell, rather than “RTX 6000 Pro.” Do not confuse it with the older Quadro RTX 6000, RTX 6000 Ada Generation, or unrelated RTX 6000 entries in vGPU documentation.
The RTX PRO 6000 family includes different editions, including Workstation Edition and Server Edition. Their VBIOS, cooling, firmware, support status, and intended virtualization paths may differ. NVIDIA’s vGPU documentation lists supported RTX PRO 6000 Blackwell Server Edition configurations, but that support documentation is not proof that generic KVM/VFIO FLR failures have been universally fixed. See the Linux KVM vGPU release notes and vGPU documentation updates.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Is this a gaming or ordinary workstation problem?
Not primarily. The documented incident involves GPU passthrough, VM lifecycle events, and device reset or reassignment. It does not establish a broad problem with ordinary bare-metal Windows or Linux gaming.
Other RTX 5090 and RTX PRO 6000 reports involve different symptoms, including Windows hibernate/resume failures, GSP timeouts, inference crashes, or full-chip resets. Those may involve related firmware or power-management components, but they should not be treated as the same defect without evidence.
For example, a separate NVIDIA forum listing includes RTX 5090 hibernate/resume reports, while another report describes an RTX PRO 6000 sustained-inference failure. These are separate failure modes.
What is known about the root cause?
No single public root cause has been proven. The evidence is consistent with an interaction involving one or more of:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute- PCIe Function-Level Reset handling.
- Secondary bus reset behavior.
- D3cold-to-D0 power-state transitions.
- GPU firmware or GSP state surviving VM teardown incorrectly.
- Interactions among the Blackwell firmware, NVIDIA driver, VFIO, motherboard firmware, and PCIe topology.
A Proxmox discussion reports errors such as “Unable to change power state from D3cold to D0, device inaccessible.” One participant said NVIDIA had reproduced the problem and was considering a fix. That is forum testimony, not a publicly identified NVIDIA engineering advisory. See the Proxmox discussion.
Mitigations administrators can test
1. Move beyond the 575-series driver branch
CloudRift says some users reported improvement with drivers in the 580-or-newer series. That should be treated as a reported workaround, not a guaranteed fix. The available report does not define one minimum driver, operating system, firmware version, or complete test matrix.
Before upgrading production hosts, test the exact combination of GPU, VBIOS, host kernel, guest driver, hypervisor, and VM lifecycle operations.
Rank #2
- AI Performance: 772 AI TOPS
- OC mode: 2580 MHz Default mode: 2550 MHz(Boost clock)
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- SFF-Ready Enthusiast GeForce Card
- Axial-tech fans feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
2. Avoid repeated reassignment
The lowest-risk architecture is to bind the GPU to VFIO at boot and dedicate it to one VM for the host’s entire uptime. Avoid repeatedly detaching and reattaching the card between guests, especially in an unattended multi-tenant service.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThis reduces reset events but does not guarantee safety: the GPU may still need to reset when the dedicated VM shuts down or reboots.
3. Test D3 power-management changes
Proxmox users have suggested disabling idle D3 handling with:
disable_idle_d3=1
The exact implementation depends on the Proxmox and Debian release, bootloader, kernel command-line configuration, and device-binding method. Do not add it blindly to a production system. D3cold has been associated with some reports, but that does not prove it is the root cause or that disabling the state fixes every system.
4. Consider guest DRM modesetting changes
One Proxmox user reported success in a particular Linux guest after adding:
options nvidia-drm modeset=0
and rebuilding the initramfs:
update-initramfs -u
The report had not established long-term stability. This may also affect display handling, framebuffer initialization, Wayland, or other graphics features. Treat it as an anecdotal, configuration-specific workaround—not an NVIDIA-approved solution. See the reported workaround.
5. Plan for host-level recovery
If the device is genuinely inaccessible, repeated reset commands may hang. Avoid repeatedly running commands such as nvidia-smi -r against a wedged GPU. Document whether a soft reboot, hard reboot, or complete PSU power cycle is required, and make sure unrelated services can tolerate that recovery path.
How to diagnose an affected host
Capture evidence before rebooting where possible:
dmesg -T | grep -Ei 'vfio|flr|pcie|nvidia|xid|d3cold|reset'
lspci -nnk
nvidia-smi -q
On Proxmox, also record:
pveversion -v
uname -a
Record the GPU model and board variant, VBIOS and driver versions, host kernel, Proxmox version, guest operating system, QEMU/libvirt versions, PCIe slot and topology, and whether the GPU was bound to VFIO at boot or dynamically detached.
Note the event that preceded the failure: guest shutdown, reboot, migration, forced stop, host suspend, or reassignment. Also record whether the guest was under sustained compute load and whether the GPU entered D3cold.
Recommended Free Tools
Should you buy or deploy these GPUs?
| Use case | Assessment |
|---|---|
| Bare-metal gaming | This passthrough report alone provides no reason to reject the RTX 5090. |
| Single VM with rare reboots | Potentially acceptable after testing, provided a host reboot is tolerable. |
| Proxmox passthrough with frequent resets | High caution; validate repeated shutdown and startup cycles first. |
| Multi-tenant GPU cloud | Prefer validated data-center hardware or a supported NVIDIA vGPU configuration. |
| Professional workstation without VM reassignment | Evaluate the workstation workload separately from the passthrough bug. |
| Production inference with strict uptime | Require long-duration testing and a tested recovery procedure. |
The RTX 5090 remains a plausible choice for bare-metal use or a dedicated passthrough VM when downtime is manageable. It is a poor fit for a service that depends on frequent, automated GPU recycling unless the exact platform has passed reliability testing.
The RTX PRO 6000 Blackwell is more appropriate for professional and enterprise workloads, but its professional branding does not make arbitrary KVM/VFIO passthrough automatically reliable. Confirm the exact edition and supported virtualization route. NVIDIA’s vGPU platform is a different deployment model from consumer-card passthrough.
For revenue-generating infrastructure, H100, B200, or another validated data-center platform may offer a safer operational design. CloudRift reported no reproduction on H100 and B200 in its comparisons, but that is not a guarantee of immunity from unrelated reset failures. The trade-off is substantially higher cost, power, cooling, and infrastructure requirements.
What has not been proven
- That every RTX 5090 or RTX PRO 6000 is affected.
- That bare-metal gaming is broadly affected.
- That the problem is definitely a physical hardware defect.
- That D3cold is the sole root cause.
- That driver 580 or newer permanently fixes all systems.
- That H100, B200, or RTX 4090 cannot experience other reset failures.
- That NVIDIA has published a universal fix for this specific FLR problem.
Bottom line
The reported RTX 5090 and RTX PRO 6000 Blackwell issue is serious for GPU virtualization because a failed reset can turn a routine VM lifecycle event into host downtime. It is not a blanket warning against buying either GPU. Treat dynamic KVM/VFIO reassignment as unvalidated until repeated reset testing succeeds, and choose supported enterprise virtualization hardware when multi-tenant availability matters.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

