My five-year-old ASUS TUF RX 6800 XT has developed a gaming-related crash that has gradually become worse. It used to run for about an hour before failing, but now games typically crash after 15 minutes. The screen goes black or freezes on the last frame, sometimes with robotic audio, but the computer does not restart.
The rest of the system uses a Ryzen 5 5600X, 32 GB of DDR4-3600 memory, an Aorus B450-I motherboard, a 750 W SFX power supply in an NZXT H1 case, and two SSDs. I have used DDU, reinstalled drivers, tested multiple Windows and Linux installations, disabled XMP, and tried different games. FurMark and OCCT 3D/VRAM tests can run without crashing. GPU temperatures stay below 90°C in the case and below 80°C during open-air testing, although the card has recently reached about 87°C core and 109°C hotspot after being opened.
Reducing the GPU clock to 2000 MHz and setting the power limit to 80%, or switching to the quieter BIOS profile, delays the crash but does not prevent it. ReBAR is enabled, and PCIe ASPM is disabled. I tested the card in another computer and tested an RTX 3060 in mine: the RTX card works normally, while the RX 6800 XT crashes in either system. I also swapped the CPU, RAM, power supply, SSD, and motherboard without changing the result.
I cleaned the card, repasted the die, and found one hardened thermal pad. Some pads appeared to have stuck to the heatsink, so I am planning to replace the remaining pads. GPU-Z reports the correct factory BIOS. Since synthetic tests pass but real games fail, what should I check next, and is the card likely suffering from a hardware fault in its VRM, memory, or power circuitry?
3 Answers
Finish the cooler work carefully before condemning the card. Clean out the dust, make sure every memory and VRM component has the correct pad thickness and actually makes contact, and verify that the heatsink is mounted evenly. The 109°C hotspot reading is high enough to justify correcting the pads and paste, even though the card may not immediately fail a benchmark. Also try separate PCIe power cables from the PSU rather than using one daisy-chained cable, and test both performance and quiet BIOS positions.
At this point the testing strongly isolates the problem to the graphics card itself. Since it crashes in two different systems, across operating systems and games, while another GPU works normally, this is very unlikely to be a driver, motherboard, RAM, or PSU issue. A failing VRM, memory module, solder joint, or marginal core can behave differently in games than in synthetic tests. Lowering the clock and power limit reducing the frequency of the crash is also consistent with a marginal card rather than proving that temperatures are the only cause.
That is what I am worried about. I am replacing all of the thermal pads first, but if that only delays the crash again, I will probably need a board-level repair or replacement.
If a full pad replacement and correct mounting do not solve it, there is not much useful software troubleshooting left. You could run an overnight VRAM and 3D test, but passing those would not rule out a fault that appears only with a particular game workload or rapid changes in power demand. A technician with board-level diagnostic equipment may be able to test the VRM and memory, but repair costs can approach the value of the card. In practical terms, replacement or selling it clearly as faulty for parts may be more sensible.
The fact that reducing power and clock speed changes the time to failure is useful evidence, but it does not identify the exact failed component. It could still be memory, power delivery, or a connection that becomes marginal under a specific workload.

The card is now clean, the BIOS switch made no difference, and I have already tested it with separate power cables. The quiet profile only extends the runtime before the crash.