PC crashes with BSOD when copying large files during GPU load

0
1
Asked By MellowPebble47 On

I'm troubleshooting a repeatable Windows 11 crash. My system has an Intel i7-14700K, RTX 4080 Super, MSI PRO Z790-A MAX WIFI, 32GB DDR5, a 1TB Kingston Fury NVMe used for Windows, and a 512GB XPG S40G NVMe.

The usual sequence is GPU-heavy gaming or shader loading, followed by an SSD reaching 100% active time, a freeze or black screen, and then a BSOD or reboot. The bugchecks have included UNEXPECTED_STORE_EXCEPTION (0x154), KERNEL_DATA_INPAGE_ERROR (0x7A), CRITICAL_PROCESS_DIED (0xEF), KMODE_EXCEPTION_NOT_HANDLED (0x1E), and STATUS_IN_PAGE_ERROR (0xC0000006). One dump referenced nt!HvpGetCellPaged with an in-page error while the Registry process was active.

CPU-only stress testing is stable, GPU-only stress testing is stable, and CPU stress combined with a large file copy is also stable. However, GPU stress combined with copying large files reliably caused crashes several times. This makes me suspect an NVMe controller, PCIe slot or root-complex problem, motherboard issue, PSU instability, or an interaction between GPU and storage load.

I updated the motherboard BIOS and SSD firmware, disabled overclocks and XMP, and tried moving the drive to another M.2 slot. After moving it, I completed four 200GB file-copy tests while stressing the GPU without a crash so far. Could the original M.2 slot or PCIe path be failing, or should I still investigate the PSU, memory, and storage drive further?

4 Answers

Answered By QuietHarbor8 On

The strongest clue is that the problem stopped after moving the NVMe drive to another slot. That points more toward the original M.2 slot, its PCIe lanes, a poor connection, or a motherboard-level issue than Windows itself. Keep repeating the combined GPU and file-copy test, and check whether the two M.2 slots use different PCIe resources or share lanes with other devices. Also inspect the drive temperature and health data while testing.

MellowPebble47 -

I moved the drive and have completed four 200GB copies during GPU stress without a crash. I’m continuing to test before calling it solved.

Answered By CopperLynx29 On

Several of those bugchecks are consistent with data failing to come back from storage, so the different stop codes may all be symptoms of the same NVMe or PCIe communication problem. Confirm that the SSD is firmly seated, check its SMART and error counters, and test the drive in the other slot or another system if possible. If the issue follows the drive, replace it; if it stays with the slot or board, the motherboard path is more suspect.

Answered By PlainMaple52 On

For completeness, test with BIOS defaults, no XMP or overclocking, and verify the RAM with a proper memory test. Check PSU cabling and GPU power connections as well, especially if the PSU is borderline for the 4080 Super. Save important data before continuing, because repeated in-page errors can indicate an unstable drive or controller.

Answered By BrightCedar6 On

Firmware and BIOS updates were worth trying, since some NVMe controllers can hang under heavy I/O and cause the disk to sit at 100% active time. Since those are already current, the slot change is more informative than the original dump names. I would avoid assuming the GPU is directly damaging the SSD: the GPU load may simply expose a marginal PCIe connection, power-delivery problem, or thermal issue.

MellowPebble47 -

The Kingston firmware and motherboard BIOS were updated recently, but the crashes continued until I tried the other M.2 slot.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.