PC Hard-Crashes and Only Boots Again After Switching Off the PSU

0
7
Asked By MellowPine47 On

My PC occasionally crashes instantly and completely, then tries to reboot. It shows the ASUS TUF logo, fails to start Windows, and eventually reaches the recovery screen. The system only boots normally again after I switch the PSU off, wait briefly, and turn it back on.

The hardware is a Ryzen 5 5600X3D, Radeon RX 6800 XT, ASUS TUF Gaming B550-Plus WiFi motherboard, 32GB Corsair Vengeance Pro DDR4-3600, Corsair RM750e PSU, and a Samsung 990 Pro 1TB SSD.

This has happened roughly 15–20 times over the past year. Nearly every crash occurred while playing Fortnite, although one happened when the PC was under very light load with no game running. Temperatures and usage look normal, generally staying below 70°C for the CPU and GPU during normal use. There are no obvious power spikes, fan surges, or overheating events.

The PSU and motherboard have been replaced, known-good memory has been tested, the BIOS and drivers are current, Windows has been freshly installed twice, and the SSD reports good health. OCCT stress tests have run for hours, including CPU, GPU, memory, and combined maximum-power tests, without reproducing the crash. The CPU reached about 90°C and the GPU pulled around 270 watts without failing.

The issue began a few months after moving to a new location. The outlets, breakers, and service panel were later updated, but the crashes continued. iCUE is running whenever the crashes happen and controls several Corsair devices, including eight fans, the RAM, AIO display, mouse, keyboard, mousepad, lighting device, and headset. iCUE and the connected devices have been reinstalled or updated multiple times, but the problem still cannot be reproduced reliably with or without iCUE.

Event Viewer does not show a useful stop code. It sometimes records DCOM errors shortly before the crash, followed by Kernel-Power reporting that Windows shut down unexpectedly. The crash itself is instantaneous. What else should I test, and could the CPU's memory controller, the CPU socket, or a software/device-control utility be responsible?

2 Answers

Answered By CedarFox88 On

Check Event Viewer and Reliability Monitor for a real bugcheck or WHEA hardware error. Kernel-Power by itself usually only means Windows lost power or was reset before it could write a proper crash report, so it does not identify the failed component. If there is no dump or stop code, this may be a power or hardware reset occurring below the Windows level.

MellowPine47 -

Event Viewer only shows a few DCOM errors before the incident and then Kernel-Power after the failed reboot attempt. There is no stop code or useful entry at the exact moment of the crash.

Answered By QuietMaple52 On

The symptoms after the crash make the CPU, its memory controller, or the socket a reasonable next suspect. A bad memory-controller state can explain why the machine hangs at the motherboard logo until power is fully removed. Run a bootable test such as MemTest86 for several passes with the memory at stock settings, disable EXPO/XMP and any CPU undervolt or overclock, and reseat the CPU while checking the socket for contamination or bent pins. If it still happens with known-good RAM and default BIOS settings, testing with another compatible CPU would help isolate the problem.

SilverBirch6 -

Even if the memory modules work in another computer, the motherboard and CPU memory controller can still be involved. Testing at JEDEC defaults is important because a marginal controller may pass synthetic stress tests but fail during a particular game.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.