I built a small-form-factor PC and used it for about a month before leaving it powered off for more than a year. After turning it on again, I began experiencing random blue screens, and eventually the NVMe SSD failed, so I replaced it. The current issue is that demanding games such as Crimson Desert, Half-Life: Alyx, and Dyson Sphere Program suddenly close without an error message, system freeze, or blue screen. The crashes can happen in the main menu or anywhere from a few seconds to 10–20 minutes into gameplay. Occasionally, Reliability History or Event Viewer records an error, but not consistently.
The system has been tested with Windows 11, CachyOS, and Bazzite, and the game crashes occur on all of them. CPU temperatures reach roughly 90°C in the Dan A4 v4.1 case, although the machine appears stable during benchmarks and stress tests. I have run 3DMark, Superposition, Heaven, CrystalDisk, MemTest64, Prime95, FurMark, OCCT, and Cinebench without detecting errors.
The BIOS is current and running at stock settings. I have tested with and without EXPO/XMP, changed PCIe link speeds between Gen 3, Gen 4, and Auto, adjusted the CPU SoC voltage, reseated the RAM, SSD, cables, and PCIe riser, and installed a replacement riser cable. The NVMe slots and PCIe settings have also been tested at different generations. What else should I investigate? Could the GPU or another hardware component be faulty even if synthetic benchmarks pass?
2 Answers
A fresh CPU thermal paste application and cooler remount would be worth trying, especially after the system sat unused for so long. A peak around 90°C during Prime95 does not automatically prove overheating, and the fact that it remains stable under that workload makes this less conclusive, but poor contact or uneven mounting can still cause game-specific temperature or boost behavior. Check the CPU temperature, GPU hotspot, and clock speeds immediately before the crash rather than relying only on benchmark results.
Try completely removing the graphics driver with a display-driver cleanup utility, then install a known-current driver fresh. Driver problems can cause games to exit without a blue screen, and some games exercise different GPU features than synthetic tests. Also check whether the crashes stop with a modest GPU power-limit reduction or underclock; if that changes the behavior, it points more toward GPU stability, power delivery, or the card itself.
I have reproduced the crashes on multiple operating systems with clean driver installations, so a normal driver issue seems less likely, but testing a reduced power limit could still help distinguish software from hardware.

The CPU reached about 91°C during Prime95 and stayed stable, so I was unsure whether repasting would help, but I can still remount the cooler to rule it out.