RTX 5090 crashes with P2PREQ/REORDER ECC errors until the PSU is fully power-cycled

0
0
Asked By MellowCedar47 On

I'm troubleshooting recurring crashes with a brand-new ASUS ROG Astral RTX 5090 OC. My system uses a Ryzen 7 9800X3D, ASUS ROG Strix B850-E Gaming WiFi motherboard, Corsair RM1000x SHIFT 1000W PSU, 32 GB of DDR5-6000 memory, Windows 11, and NVIDIA driver 610.88.

The card can run games for hours, but when the problem appears, Escape from Tarkov often crashes within about two minutes of entering a raid. I get black screens, driver timeouts, reboots, or VIDEO_TDR_FAILURE (0x116) BSODs. Event Viewer repeatedly reports nvlddmkm Event ID 14 errors involving "PCIE P2PREQ" and "PCIE REORDER" uncorrectable SRAM/ECC errors, followed by Event ID 153 GPU reset messages. WinDbg points to nvlddmkm.sys and a GPU TDR timeout.

I have reinstalled Windows, installed fresh chipset and NVIDIA drivers, reseated the card, cleaned the PCIe slot, checked the power connector, moved an NVMe drive so the GPU runs at PCIe 5.0 x16, reduced the memory from DDR5-7200 to 6000, and removed all overclocking software. PCIe monitoring usually shows no receiver, replay, LCRC, TLP, or DLLP errors. Temperatures and 12V readings also look normal. PCIe Gen4 was briefly more stable, although Gen5 became stable again after reseating the card.

The strangest part is that shutting Windows down, switching the PSU off, and waiting for a complete power discharge often makes the system stable for several days. A normal reboot does not have the same effect. OCCT VRAM, GPU stress tests, and FurMark have passed, even though some games still trigger the issue.

I'm planning to try the standard Corsair 600W cable, update the official VBIOS, force PCIe Gen4, test reduced power and clock settings, and possibly try a larger PSU. Has anyone encountered these P2PREQ/REORDER uncorrectable SRAM or ECC errors, and did a VBIOS update, cable or PSU replacement, motherboard BIOS change, PCIe Gen4, or GPU replacement resolve them?

3 Answers

Answered By PixelHarbor9 On

A few useful checks would be confirming whether the card was purchased new, testing both performance-mode switches, and temporarily lowering the GPU power limit and core or memory clocks. A proper OCCT VRAM test is also worth running. If reduced power or clocks make the problem disappear, that points more toward the card or power delivery than Windows or the game. A clear photo of the connector and cable routing could also help verify that the plug is fully seated and not under excessive strain.

MellowCedar47 -

The card was bought new about two and a half months ago, and the crashes only began around two weeks ago. I have not flashed anything besides the official BIOS. OCCT VRAM, GPU stress tests, and FurMark all passed. I switched from a 90-degree Corsair cable to the standard cable that came with the PSU, and the system has been stable for two days so far, although it is too early to call that a fix.

Answered By SignalMaple61 On

The uncorrectable P2PREQ and REORDER SRAM errors are more concerning than a routine driver crash. A clean Windows installation and multiple driver versions make software less likely, especially when the failure occurs at the PCIe/GPU level. Try the latest motherboard BIOS and official GPU VBIOS, then force PCIe Gen4 as a diagnostic. If the crashes continue with a different cable, PSU, BIOS configuration, and reduced clocks, the card itself or its PCIe interface may be defective and an RMA would be appropriate. Save the Event Viewer entries and several minidumps to document the failure.

Answered By VoltTangent23 On

A 1000W unit is within the usual recommendation, but a 5090 can draw roughly 600W or more by itself, with additional transient demand. The CPU, pumps, fans, drives, and other hardware reduce the remaining headroom. Testing with a known-good higher-capacity ATX 3.x PSU and its proper 600W cable is reasonable. Since a complete PSU power-off changes the behavior while a reboot does not, power delivery or a latched GPU state is worth investigating.

MellowCedar47 -

I have ordered a 1500W PSU so I can test whether the issue is related to transient load or the current unit. I’ll also compare the original cable with the replacement cable rather than changing several variables at once.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.