My five-year-old ASUS TUF RX 6800 XT has gradually developed a problem where games crash after about an hour, and now they usually fail within 15 minutes. The screen goes black or freezes on the last frame, sometimes with robotic audio, but the computer does not fully restart.
My system has a Ryzen 5 5600X, 32 GB of DDR4-3600 memory, an Aorus B450-I motherboard, a 750W SFX power supply in an NZXT H1 case, and the RX 6800 XT. I have tried DDU with multiple driver installation types, fresh Windows installations, several Linux distributions, different games, and reduced memory settings. I also tested the card in another computer.
FurMark and OCCT's 3D and VRAM tests do not crash, but reducing the GPU clock to 2000 MHz and lowering the power limit to 80% only delays the failure. Temperatures normally stay below 90°C in the case, and are lower during open-air testing, although after repasting I have seen around 87°C core and 109°C hotspot. I cleaned the card and replaced one hardened thermal pad, with more replacement pads on the way.
Swapping the CPU, RAM, power supply, storage, and motherboard made no difference. An RTX 3060 works normally in my system for hours, while the RX 6800 XT crashes in both systems. I also tried the card in a secondary PCIe slot and disabled PCIe ASPM, but the problem remains. The BIOS switch makes the issue take longer in quiet mode, but does not prevent it. What should I check next, and is this likely a hardware failure in the GPU?
4 Answers
Make sure the cooler is making proper contact everywhere. The card was very dusty, and the photos suggest that some memory or VRM areas may not have had pads in the expected positions because pads can remain stuck to the heatsink. Replace the complete pad set with the correct thicknesses rather than adding random pads, clean the heatsink thoroughly, and verify that the heatsink screws are tightened evenly. A core temperature around 87°C and a 109°C hotspot are high enough to justify fixing the cooling first, even if they do not fully explain why benchmarks pass.
The fact that it only happens in games is not conclusive evidence against hardware. Different games can produce transient loads, memory access patterns, and power-state changes that a steady benchmark never reproduces. Event Viewer may only show a generic display-driver timeout or unexpected loss of the device, so the absence of a useful error there does not clear the GPU.
Also verify the power connection setup: use separate PCIe power cables from the PSU for each GPU connector rather than one daisy-chained cable, and inspect the plugs for looseness or heat damage. If the card still crashes with known-good cabling, correct thermal-pad placement, stock settings, and clean power, there is probably no practical software fix left. A repair shop with board-level diagnostic equipment could test the VRM and memory, but the cost may exceed the card's value; replacement or selling it clearly as faulty may be the more sensible option.
The component swapping strongly points to the 6800 XT itself. Since the card fails in another computer and a different GPU works in yours, this is very unlikely to be Windows, drivers, RAM, or the motherboard. A faulty VRM, memory module, solder joint, or aging power-delivery component could fail under a particular gaming workload even when synthetic tests pass. Lowering the clock and power limit reducing the frequency of crashes also fits an aging or marginal card, although it does not identify the exact part.
That is what I am starting to suspect. The reduced power limit only extends the play time, so I am going to finish replacing the thermal pads before treating it as a board-level fault.

The card is clean now, and I found one hardened pad when I opened it. I only replaced that one initially, so I have ordered a full, correctly sized set. The card usually stays below 90°C hotspot during open testing, but I want to rule out contact and VRAM or VRM cooling problems properly.