RTX 5090 black screens and Event 153 TDR crashes on AM5—could PCIe Gen 5 be the cause?

0
9
Asked By MellowPine47 On

I'm troubleshooting recurring black-screen crashes with an MSI RTX 5090 Gaming Trio in a system with a Ryzen 9 7900X3D, ASUS ROG STRIX X670E motherboard, 64 GB DDR5, Corsair HX1200 PSU, and Windows 11. The issue happens during demanding games, including Delta Force and Dota 2. The display goes black, but Discord voice chat and game audio continue for a while before the system eventually recovers or reboots. Event Viewer reports VIDEO_TDR_FAILURE (0x116) with Arg3 0xC000009A, along with NVIDIA nvlddmkm Event ID 153 messages such as "GpuRcReset TDR" and "BusReset TDR occurred." I have already performed a clean driver installation with DDU, repaired Windows updates, reseated the GPU and power connections, tested FurMark, removed overclocks, reduced GPU power, and lowered the RAM from DDR5-5600 to DDR5-3600. Lowering GPU power seemed to improve stability temporarily. The GPU is connected directly to an Alienware AW3225QF over DisplayPort, not through the integrated graphics. I'm currently using a Lian Li Strimer cable instead of the original 12V-2x6 to 3x8-pin adapter. My next test is to update the motherboard BIOS, use the PSU's original GPU power cable or the bundled adapter, and force the slot from PCIe Gen 5 to Gen 4. Has anyone experienced this combination of black screens, continuing audio, Event 153, and 0x116 TDR errors on an AM5/X670E or X870E system? Did PCIe Gen 4, a different power cable, BIOS or firmware updates, a motherboard change, or an RMA ultimately solve it?

5 Answers

Answered By CableCheckR8 On

The power connection is worth investigating, especially with a high-power RTX 5090. Temporarily remove the Strimer and use the PSU’s original cable or the adapter supplied with the card, making sure the 12V-2x6 connector is fully inserted and not under sideways tension. If that fixes the crashes, use a direct cable designed specifically for the HX1200 rather than an extension or adapter.

MellowPine47 -

The card is an MSI RTX 5090 Gaming Trio, and the connector is seated, but I’ll repeat the test with the original PSU cable or bundled adapter instead of the Strimer.

Answered By BIOSWaypoint3 On

Update the motherboard BIOS to the newest stable release before drawing conclusions about PCIe stability. After that, force the GPU slot to PCIe Gen 4 and retest with stock GPU settings and DDR5-3600. If Gen 4 stops the crashes, that points toward a Gen 5 link-training, motherboard firmware, signal-integrity, or platform compatibility issue rather than proving the GPU itself is defective.

MellowPine47 -

That is the test I’m planning next. I’ll keep the GPU at stock settings and change only the PCIe generation so the result is easier to interpret.

Answered By PixelHarbor6 On

Event ID 153 has been associated with NVIDIA driver and GPU-reset problems for quite some time. Different driver versions may change how often it appears, but it does not necessarily identify the root cause by itself. Since the crashes persist across multiple games, it is worth testing the hardware and platform variables one at a time rather than assuming the game or Windows installation is responsible.

MellowPine47 -

That makes sense. I’ve already tried a clean driver install, so I’m going to focus next on the power connection, BIOS, and PCIe generation.

Answered By LaneStabilityQ2 On

Some AM5 systems can become unstable when the CPU’s memory-controller and PCIe fabric are heavily loaded. A motherboard BIOS update is the safer first step; changing SoC voltage should only be considered carefully and within the board and CPU manufacturer’s recommended limits. It is not a guaranteed fix for a GPU TDR and should not replace checking the cable, slot, temperatures, and hardware.

MellowPine47 -

I hadn’t considered the CPU’s effect on PCIe stability, but I’ll avoid making aggressive voltage changes and test the firmware, cable, and Gen 4 settings first.

Answered By DisplayPathV5 On

Because the monitor is connected directly to the RTX 5090, the display cable and monitor are still worth ruling out, even though Event 153 and 0x116 point more strongly toward a GPU reset. Testing another certified DisplayPort cable or monitor can help exclude a display-link issue, but it would not explain every possible bus-reset error.

MellowPine47 -

The monitor is an Alienware AW3225QF over direct DisplayPort. I’ll keep the cable and display path in mind, but the power and PCIe tests seem like the higher-priority checks.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.