GPU Load Plus Large File Copies Cause Repeatable BSODs—Is the Kingston NVMe Failing?

0
0
Asked By MellowBirch42 On

My Windows 11 PC repeatedly crashes when the GPU is under heavy load at the same time as large file transfers. The system uses an Intel i7-14700K, RTX 4080 Super, MSI PRO Z790-A MAX WIFI, 32GB DDR5, a Kingston Fury 1TB NVMe as the Windows drive, and an XPG S40G 512GB NVMe as a secondary drive. The crashes usually follow this pattern: a game or shader load starts, an SSD reaches 100% active time, the system freezes or black-screens, and then it either BSODs or reboots. Reported bugchecks include 0x154 UNEXPECTED_STORE_EXCEPTION, 0x7A KERNEL_DATA_INPAGE_ERROR, 0xEF CRITICAL_PROCESS_DIED, 0x1E KMODE_EXCEPTION_NOT_HANDLED, and 0xC0000006 STATUS_IN_PAGE_ERROR. One dump referenced nt!HvpGetCellPaged with STATUS_IN_PAGE_ERROR, while another showed CORRUPT_MODULELIST_0xEF and failed while writing the memory dump. CPU-only stress, CPU stress during a roughly 200GB copy, and GPU-only stress are stable. Combining GPU stress with a large file copy reproduces the crash, however, and has done so several times. I have updated the motherboard BIOS and SSD firmware, disabled overclocks and XMP, and tried moving the Kingston drive to another M.2 slot. After moving it, four 200GB copy tests during GPU stress completed without a crash, although I am continuing to test. What would be the best way to distinguish between a failing SSD or controller, a motherboard or PCIe-slot problem, a PSU or power-delivery issue, and a GPU/NVMe interaction?

3 Answers

Answered By NorthstarQuill8 On

Run separate isolation tests rather than relying only on the combined workload. Put the Kingston drive in the slot normally used by the XPG drive, test the XPG drive under the same copy workload, and if possible copy between drives while monitoring both temperatures and SMART/NVMe health data. Check Event Viewer for disk, storahci, stornvme, WHEA-Logger, and PCIe errors immediately before the crash. If only the Kingston drive fails regardless of slot, replace or RMA it. If the problem follows the slot, investigate the motherboard or PCIe connection.

Answered By GlassCedar53 On

The different bugchecks do not necessarily mean several unrelated problems. If the storage device temporarily disappears or returns corrupted data, Windows can report paging errors, critical-process failures, registry faults, and even apparently random kernel exceptions. The failed dump-writing messages also fit a system that is losing access to the system drive during the crash. Moving the SSD and getting four successful combined tests is encouraging, but it is not conclusive—continue with longer tests, inspect the drive's health logs, and keep a current backup.

Answered By SolarMoth26 On

A PSU or power-delivery problem is still possible because the GPU and NVMe are being stressed simultaneously, but the fact that the issue appears tied to one drive or one M.2 location makes the storage path the stronger lead for now. Verify the GPU power cable is firmly seated, avoid adapters if possible, and check whether the PSU has enough capacity and suitable dedicated cabling. If the failure follows the Kingston drive after swapping slots and cables where applicable, replace the drive before blaming the motherboard or PSU.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.