RTX 5090 crashes with TDR errors, but Debug Mode prevents the black screens

0
5
Asked By MellowPine47 On

I'm troubleshooting intermittent crashes on a new high-end Windows 11 PC with an ASUS ROG Astral RTX 5090 OC, Ryzen 9 9950X3D, X870E motherboard, 64 GB DDR5-6000 memory, and a 1200 W ATX 3.1 PSU. The system sometimes runs demanding games for hours, but it can also crash while browsing or shortly after closing a game. Symptoms include VIDEO_TDR_FAILURE blue screens, black screens with no HDMI signal, fans ramping to maximum, total freezes, and occasional loss of HDMI audio to my Sony 4K OLED TV.

The most useful reproducible failure occurred during an OCCT 3D Adaptive test. After about three minutes, the display went black and the system froze. Event Viewer recorded an NVIDIA nvlddmkm Event 14 reporting an uncorrectable ECC error in the GPU PCIe REORDER unit, followed by repeated GPU reset/TDR Event 153 messages and BugCheck 0x116 (VIDEO_TDR_FAILURE). The dump referenced nvlddmkm.sys, dxgkrnl.sys, watchdog.sys, and pci.sys, including GPU bus-reset activity.

CPU, system memory, VRAM, and combined CPU/GPU power tests passed. No WHEA errors were found. Safe Mode was stable, although it does not use the normal NVIDIA graphics stack. The current NVIDIA driver is 616.64, and a DDU cleanup was already performed.

The key result is that the same OCCT 3D Adaptive test completed 15 minutes without a crash when NVIDIA Debug Mode was enabled. Debug Mode temporarily removes the card's factory overclock and boost behavior, even though no manual overclock was configured. Does this point primarily to an unstable factory GPU overclock, or could an ASUS/Armoury Crate driver or RGB-control service still be responsible? What would be the best next diagnostic step without permanently reducing the card's performance?

4 Answers

Answered By SensibleCedar31 On

The Kernel-Power 41 and Event 6008 entries mainly record the forced or unexpected shutdown. The missing dump and volmgr Event 161 are consequences of the crash or dump-writing failure, not proof that an SSD is defective. Likewise, PCIe Gen 2 while idle is normal power saving, and a zero replay count after reboot does not disprove the earlier PCIe-related error.

I would focus on reproducing the failure and comparing normal mode, Debug Mode, and a clean driver-only setup. Keep notes for each run, including driver version, BIOS settings, Armoury Crate status, cable configuration, and whether the crash occurs under load or at idle.

Answered By CircuitFern2 On

Armoury Crate and motherboard utilities are worth removing from the test environment. RGB, monitoring, fan-control, and automatic driver services can interact with the GPU driver and occasionally cause instability. Uninstall Armoury Crate and Aura-related components, disable automatic hardware-driver installation where possible, reboot, and retest with only the NVIDIA driver installed. Also check whether the crashes stop after a clean boot.

That said, software alone does not fully explain why disabling the factory GPU profile makes the exact 3D test pass. Treat the Armoury Crate theory as a separate variable and change only one thing at a time so you know which test matters.

MellowPine47 -

The technician currently suspects an Armoury Crate driver or RGB service. After removing Armoury Crate, the system appeared stable for a while, but we are still checking whether a component reinstalls during boot.

Answered By ByteTrail6 On

The Event 14 ECC message and the later TDR/bus-reset entries show that the GPU stopped responding, but they do not identify whether the GPU itself, its firmware, the PCIe connection, power delivery, or a driver triggered it. Passing short CPU, RAM, VRAM, and PSU tests cannot completely rule out an intermittent fault.

For isolation, test with the latest stable driver rather than repeatedly reinstalling the same version, load BIOS defaults, temporarily disable motherboard enhancement features, reseat the GPU and its power connector, and verify that the 12V-2x6 connector is fully inserted with no sharp bend near the plug. If possible, test the card in another known-good system or test another GPU in this system. A Linux live environment can also help separate Windows-driver problems from hardware instability, although it is not a perfect substitute for a normal gaming workload.

Answered By QuietHarbor8 On

The strongest clue is the A/B result: the 3D test crashes with the card’s normal factory profile but passes in NVIDIA Debug Mode. That makes the factory overclock or boost curve the leading suspect, even though the NVIDIA App shows +0 MHz because the factory settings are part of the card’s VBIOS rather than a manual adjustment.

Run several longer tests in Debug Mode and also test real games. If everything remains stable, the card may simply be marginal at its advertised factory boost. You could then ask the manufacturer or retailer for an RMA rather than treating a permanent performance reduction as the solution. A modest power limit or clock reduction can be useful as a diagnostic, but it does not prove the underlying hardware is healthy.

MellowPine47 -

That was my concern too. I’m using Debug Mode only as a diagnostic test, since I bought the OC model specifically for its intended performance. The identical OCCT test passed for 15 minutes with Debug Mode enabled, while it failed after about three minutes normally.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.