I'm troubleshooting an intermittent graphics stability problem on a Windows 11 PC with an ASUS ROG Astral RTX 5090 OC, Ryzen 9 9950X3D, X870E motherboard, 64 GB DDR5-6000, and a 1200 W ATX 3.1 power supply. The system sometimes runs demanding games for hours, but it can also crash during light desktop use or shortly after closing a game. Symptoms include VIDEO_TDR_FAILURE blue screens, black screens with no HDMI signal, fans running at maximum, complete freezes, forced shutdowns, and occasional loss of HDMI audio to my Sony 4K OLED TV.
CPU and memory tests passed, as did OCCT power and VRAM tests. The PSU appears capable of handling simultaneous CPU and GPU loads, and there are no WHEA hardware errors. A controlled OCCT 3D Adaptive test reproduced the problem after about three minutes. Event Viewer then reported an NVIDIA nvlddmkm Event 14 stating that an uncorrectable ECC error occurred in the GPU PCIe REORDER unit, followed by repeated TDR and bus-reset events. The resulting dump showed VIDEO_TDR_FAILURE 0x116 involving nvlddmkm.sys, dxgkrnl.sys, watchdog.sys, and pci.sys.
The most important result is that the same 3D test completed 15 minutes without a crash when NVIDIA Debug Mode was enabled. Debug Mode temporarily removes the factory overclock and boost behavior from the ASUS OC card. There is no manual overclock configured, and I would prefer not to permanently reduce performance unless necessary. Could the factory overclock, GPU firmware, motherboard software, or an Armoury Crate/Aura Sync component be causing the instability? What would be the best next diagnostic step, and does this evidence justify an RMA or replacement of the graphics card?
3 Answers
The 0x116 bugcheck and nvlddmkm messages show that Windows lost communication with the graphics stack, but they do not by themselves prove that the NVIDIA driver is defective. The earlier 0x7F double-fault event is also a generic kernel failure, and the missing dump was probably a consequence of the hard freeze rather than proof of a bad SSD.
For hardware isolation, test with Armoury Crate removed, one RAM module at a time if necessary, and a different operating system or clean Windows installation. If possible, test the GPU in another compatible system or test another known-good GPU in this one. Also reseat the card and inspect the 12V power connector and cable. A successful OCCT power test makes a gross PSU failure less likely, but it cannot completely rule out an intermittent connector, transient power issue, or PCIe problem.
Do a clean software isolation test before replacing hardware. Uninstall Armoury Crate, Aura Sync, motherboard utilities, overlays, and GPU monitoring or RGB tools, then prevent those components from reinstalling automatically. Use only the motherboard chipset drivers, the current NVIDIA driver, and Windows for the test. Utilities that communicate with the GPU can trigger driver resets or interfere with power and boost control, even if they are not technically overclocking the card.
If the system becomes stable after removing Armoury Crate, the likely cause is a software or service conflict. If it still crashes at normal factory settings but remains stable in Debug Mode, the factory GPU profile or the card itself becomes much more suspicious.
The Debug Mode result is the strongest clue here. Since the same 3D workload crashed with the card’s normal factory boost behavior but passed when Debug Mode removed the factory overclock, the card may be unstable at its advertised OC settings even though you never applied a manual overclock. That can be caused by the GPU itself, its VBIOS, power delivery, or an interaction with the motherboard and driver stack.
I would first run the card at stock reference clocks as a diagnostic, update the motherboard BIOS and GPU VBIOS if an official update exists, and remove all GPU/RGB tuning utilities. If it remains stable only with the factory OC disabled, document the Event 14, 0x116 dump, and repeatable test results and contact the retailer or ASUS about an RMA. Passing short VRAM, CPU, and power tests does not rule out an intermittent GPU or PCIe-path fault.
That makes sense. I’m treating Debug Mode as a diagnostic rather than a permanent fix because the card is supposed to run at its advertised settings.

My technician suspects an Armoury Crate driver or Aura Sync service. After Armoury Crate was removed, the PC appeared stable, but we’re still testing because the problem has been intermittent.