I'm troubleshooting an intermittent black-screen and hard-lock problem with an ASUS TUF RTX 4090 OC. During games, the monitor suddenly loses signal while audio, fans, and lighting may continue working. The system sometimes fails similarly when waking from sleep, and I have to force it off and restart. The problem started about three months after the system was built and has recently become more frequent.
The system uses an Intel Core i9-13900K, ASUS ROG Strix Z790-E Gaming WiFi, 64 GB DDR5 currently running at 4800 MT/s, an ASUS TUF RTX 4090 OC, a Corsair RM1000e 1000 W PSU, and a Samsung 49-inch 5120×1440 ultrawide monitor. Windows 11, the motherboard BIOS, and NVIDIA drivers are up to date, and Resizable BAR is enabled.
Logs show DXGI device-hung or device-removed errors, nvlddmkm TDR activity, and more than 5,700 WHEA Event 17 warnings. Most have the same PCIe Completion Timeout signature on bus 0, device 1, function 0, involving Intel device A70D. The latest failure occurred during a game without producing a normal application crash or blue screen; I had to force a restart, resulting in Kernel-Power Event 41 with bugcheck code 0.
I have already updated the BIOS and graphics driver, disabled overlays, tried different display modes, enabled V-Sync, capped the frame rate, reduced the graphics load, forced PCIe Gen4 and Gen3, tested a 60% GPU power limit, and tried running without GPU Tweak III. Temperatures and GPU telemetry look normal. The problem persists at Gen3 and reduced power, and the original PCIe errors appeared before GPU Tweak III was installed.
Possible causes include the GPU, its 16-pin power connection or cable, the motherboard or CPU-connected PCIe path, the PSU, Armoury Crate, or a driver and firmware interaction. I'm considering reseating and inspecting the GPU and power connector, trying another compatible cable, testing a lower chipset-connected PCIe slot, uninstalling Armoury Crate, and borrowing another PSU. I may also pursue an ASUS warranty claim.
Has anyone encountered this combination of signal loss, continued audio, DXGI and nvlddmkm errors, and thousands of PCIe Completion Timeout warnings? What is the most effective way to isolate the faulty component without access to another high-end system?
2 Answers
Run separate OCCT tests for the CPU and memory rather than testing everything at once. The 13th- and 14th-generation Intel platforms have had stability problems, and high-capacity memory or multiple DIMMs can make them harder to diagnose. Make sure the motherboard is running a BIOS version that includes the Intel voltage and stability fixes, then check whether the errors appear during CPU or memory testing.
How old is the processor, and were all of the BIOS updates addressing the Intel stability issues installed? Since the failure can come from the CPU’s PCIe controller or platform instability, confirming the BIOS history and testing the CPU and memory is important before blaming the graphics card. A current BIOS with automatic overclocking disabled does not completely rule out a CPU or motherboard problem, but it is a useful baseline.
The 13900K has been installed since the system was built in 2023. I updated the Z790 motherboard to BIOS 3202, with AI overclocking disabled, but the black-screen problem continued afterward.

I’m running those tests now and will compare the results with the PCIe errors.