I'm troubleshooting an increasingly frequent black-screen and hard-lock problem with an ASUS TUF RTX 4090 OC. During games, the monitor suddenly loses signal while audio, fans, and system lighting continue running. The display does not recover, so I have to force the PC off. A similar issue has occasionally happened when waking from sleep.
The system uses an Intel Core i9-13900K, ASUS ROG Strix Z790-E Gaming WiFi, 64 GB DDR5 currently at 4800 MT/s, an ASUS TUF RTX 4090 OC, and a Corsair RM1000e 1000 W PSU. Windows 11, the motherboard BIOS, and NVIDIA drivers are up to date. Resizable BAR is enabled and the GPU negotiates at PCIe Gen4 x16.
Logs show previous DXGI device-hung/device-removed errors and nvlddmkm TDR activity. More concerningly, Windows has recorded over 5,700 WHEA Event 17 warnings from PCI Express Root Port bus 0, device 1, function 0, with PCIe Completion Timeout status codes. There is no normal crash dump or blue screen; after a forced restart, the system reports Kernel-Power Event 41 with bugcheck code 0.
I have already updated the BIOS and graphics driver, disabled overlays, changed game display modes, enabled V-Sync, capped the frame rate, reduced graphics settings, tested PCIe Gen4 and Gen3, tried a 60% GPU power limit, and tested without GPU Tweak III. Temperatures and reported GPU throttling look normal, but the failures continue.
Possible causes include the RTX 4090, its 16-pin power connection or cable, the motherboard or CPU-connected PCIe path, the PSU, Armoury Crate, or a driver/firmware interaction. I'm considering reseating and inspecting the GPU and power connector, trying another compatible cable, testing a lower chipset-connected PCIe slot, uninstalling Armoury Crate, and trying another PSU. Without a second high-end PC, what is the most reliable way to isolate the faulty component?
2 Answers
Run separate OCCT tests for the CPU and memory rather than only stressing the GPU. The 13th- and 14th-generation Intel platforms have had stability problems, and high-capacity or four-stick memory configurations can make them harder to spot. Make sure the board is on a BIOS version containing Intel’s voltage and stability fixes, then check whether the errors appear during CPU or memory testing.
How old is the processor, and were all of the BIOS updates addressing the Intel stability issues installed? The CPU has been in the system since 2023, and the motherboard is now on BIOS 3202 with AI overclocking disabled. Since the problem remains after that update, the BIOS is less likely to be the only cause, though CPU, memory, and platform stability tests are still worthwhile.
The system was updated to BIOS 3202, but the black-screen failure still happened afterward, so I’m checking the CPU and memory separately as the next step.

I’m running those tests now and will compare the results with the PCIe warnings.