RTX 5090 loses display signal and sends fans to full speed under heavy load

0
0
Asked By MellowCedar47 On

I'm troubleshooting an intermittent desktop failure where both monitors suddenly go black and enter standby while the PC remains powered on and the case fans ramp to maximum speed. The system does not recover, although audio can continue playing and keyboard shortcuts may still work. I have to hold the case power button to shut it down.

The failures mainly occur during or shortly after demanding workloads, including modern games, local Qwen inference, and OCCT graphics testing. They have also happened during Chrome use or apparent idle after earlier heavy activity, but never during an older, less demanding game.

The system uses an ASUS ROG Astral RTX 5090 OC, Ryzen 9 9950X3D, ASUS ROG Crosshair X870E Hero with BIOS 2402, 64 GB of DDR5-6000 CL30 with EXPO enabled, an ASUS ROG Strix 1200W Platinum ATX 3.1 power supply, Samsung 990 Pro 4 TB storage, and Windows 11 Pro 25H2 with NVIDIA driver 616.92. Both connected displays are 4K monitors, and diagnostics report the GPU operating at PCIe Gen 5 x16.

The GPU is currently at its default settings, while EXPO remains enabled. A driver clean installation did not prevent further crashes, and removing ASUS GPU Tweak also did not resolve the issue. OCCT memory and VRAM tests, a variable-load 3D test, and a combined power test have generally passed, although one graphics stress run crashed while a later run passed. Testing each monitor separately and both together has not produced a consistent trigger.

Crash records include 0x141 VIDEO_ENGINE_TIMEOUT_DETECTED dumps, a 0x1A8 display-path failure, older 0x116 graphics recovery failures, and several 0x7F parameter-8 double faults. There were also volmgr dump-writing failures, though it is unclear whether those are related.

What should I prioritize next to separate a graphics-driver problem from a faulty GPU, power-delivery issue, RAM instability, or broader platform problem? Would testing this GPU in another suitably powered system and a known-good GPU in this one be the most useful next step, or is there a controlled test I should perform first?

4 Answers

Answered By NorthVale5 On

The store’s swap testing is probably the most decisive next step once the cable and BIOS settings are checked. Have them test your RTX 5090 in another system with a known-good, high-capacity PSU, then test a known-good GPU in your machine using your PSU and the same workload. Ideally change only one component at a time and repeat the workload that usually triggers the failure.

If your 5090 fails in the other system, the card or its power connector is strongly implicated. If another GPU fails in your system, focus on the PSU, motherboard, RAM settings, BIOS, or Windows installation. If neither swap reproduces the problem, the issue may be workload-specific or intermittent, so keep a log of driver version, temperatures, power limits, EXPO state, and exact test duration.

Answered By CopperLynx82 On

Start with the physical power path before assuming the driver is responsible. Shut the system down, reseat the GPU, and fully disconnect and reconnect the GPU power connector until it is completely latched with no visible gap. Also reseat the modular PSU-side connections and inspect the cable for damage or sharp bends near the plug. An intermittent high-current connection can produce black screens and maximum fan speed even when the PC itself is still running.

If the problem continues, temporarily undervolting or reducing the GPU power limit can be a useful diagnostic. If that makes the crashes disappear, the GPU or its power delivery becomes much more suspicious. A different known-good PSU is another worthwhile comparison, despite the existing 1200W unit being adequately rated.

Answered By QuietMarble29 On

EXPO is still a meaningful variable here. Run the machine at completely stock memory settings for a while, then test with one DIMM at a time in the recommended slot. A memory or memory-controller issue can surface as graphics-driver timeouts or seemingly unrelated double faults, so passing a short memory test does not fully clear it.

Likewise, reseating the GPU is worth doing. A card that is not perfectly seated can behave normally in lighter workloads and fail when the slot, connector, or card is under more electrical or thermal stress.

Answered By PixelHarbor6 On

Use DDU from Safe Mode to remove the NVIDIA driver completely, then install a known-stable driver with only the necessary components. The installer’s clean-install option does not remove everything that DDU does. It would also be reasonable to disable overlays and Game Bar temporarily, since they can contribute to graphics-driver timeouts in some games.

If a clean driver setup changes nothing, a clean Windows installation from USB can help rule out corrupted system software, although I would do the hardware checks first.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.