My Gigabyte RTX 5090 Gaming OC has been causing black-screen crashes followed by a hard reboot. There is no BSOD or useful dump file, although audio sometimes continues briefly in the background. The crashes began after upgrading from an RTX 2060 Super and happen most often in Unreal Engine games or during loading screens. World of Warcraft, Hitman, Kingdom Come 2, Starfield, and some other demanding games can run normally, while STALKER 2, Dispatch, and certain Crimson Desert settings fail consistently.
I have already tried power limits, core underclocks, different PCIe generations, multiple drivers, BIOS updates, several 12V-2x6 cables, a different 1500 W PSU, and a WireView monitor. I also rebuilt the entire system around the card with a new motherboard and CPU, tested Windows and Linux, and ran the RAM at JEDEC settings after memory testing. My old RTX 2060 Super works reliably in the same slot with the same PSU and cable. The 5090 was sent back under warranty, but it was returned with only a note saying the thermal paste had been replaced, and the exact same crashes remain.
Linux logging shows an NVIDIA Xid 79 error and a PCIe AER RxErr when the GPU disappears from the bus. The card's reported thermal-limit headroom can collapse from roughly 40 degrees to zero or below in a fraction of a second, even when the card is drawing only about 80–120 W and the external core temperature is in the 40–60 °C range. Limiting memory clocks allows some tests to survive, while higher core clocks eventually trigger permanent throttling or shutdown. The issue can also be reproduced with PCIe forced to Gen1.
Does this evidence point to a defective GPU package, internal hotspot, sensor problem, cooler contact issue, or something else? Should I pursue another warranty claim or replacement rather than continue changing software and platform settings?
5 Answers
There is no useful BSOD dump to retrieve here because this is a hard GPU or platform-level reset, not a normal Windows stop error. The most convincing evidence is that the 5090 fails under a controlled Linux workload while the system loses the device from the PCIe bus. Software may determine which games trigger it, but it does not explain the card disappearing across operating systems and hardware configurations. A replacement 5090 is the appropriate next step.
I would avoid spending more time on RAM slots, driver cleanup, or reinstalling the operating system. Those are reasonable early checks, but the same behavior under Linux on a rebuilt machine makes them very unlikely. Keep the card sealed, save the sensor logs, record a short video of the reproducible test, and submit a second warranty claim explicitly describing the Xid 79/AER failure and the successful 2060 Super comparison. Ask for a replacement rather than another paste application or a basic power-on test.
At this point the graphics card is the prime suspect. You have reproduced the failure across two platforms, two operating systems, different drivers, PCIe link speeds, cables, and a different power supply. The old card working in the same slot is especially useful evidence. Xid 79 combined with the PCIe error means the GPU is dropping off the bus, not producing an ordinary game crash. I would document the Linux logs and repeatable test case, then request a replacement or refund through the warranty process. I would not open or reflow the card yourself while it is still covered.
The fact that one game runs and another fails does not clear the GPU. Different engines and settings can exercise different shader, memory, voltage, boost, and initialization paths. DLAA versus DLSS can change the workload enough to expose a marginal component, and a crash during a splash screen can still be a low-level GPU failure rather than a problem with the game itself. This pattern is consistent with a card that fails under specific internal operating conditions, not necessarily with a game-specific bug.
The thermal data does look abnormal, even though the reported GPU core temperature is not high. A very fast collapse in thermal headroom while the card is at low external temperature could indicate an internal hotspot, poor contact somewhere in the package or cooler, or a faulty sensor/thermal protection path. Replacing paste would not necessarily fix a defective die, package, memory-power area, or sensor. The fact that reduced clocks and power limits change the symptoms also fits a hardware fault that appears when particular parts of the card are energized.
A custom cooler or reflow would be a last resort, not a sensible next step while warranty coverage exists. The retailer should replace the card before any irreversible repair is attempted.

That is what I am hoping to establish. I am concerned the retailer will simply test the desktop and return it again, so I am collecting the cross-platform results and the exact thermal-limit behavior as evidence.