My ZOTAC GeForce GTX 1080 Ti AMP 11GB recently began crashing in several games, including Valorant, CS, Black Myth: Wukong, and Hitman, despite working reliably in this same PC for about a year. Crashes range from simple exits to desktop and freezes to complete system lockups requiring a forced restart. Event Viewer repeatedly logs nvlddmkm Events 13, 14, and 153, including "Illegal Instruction Encoding," "Physical Multiple Warp Errors," and GPU errors involving GPUID 100. I have also seen LiveKernelEvent 141 and Error_DMA_PageFault reports with graphics-engine resets.
The system uses an Intel i5-10400F, Gigabyte H410M S2 V2 motherboard, 16GB DDR4, an MSI MAG A650BN 650W PSU, Windows 11, and current NVIDIA drivers. MemTest86, storage checks, fresh Windows installations, DDU driver removal, BIOS and chipset updates, and PCIe Gen2/Gen3 testing have not resolved the issue. The card operates at PCIe 3.0 x16 under load. Lowering the power limit to 80%, core clock by 150 MHz, and memory clock by 500 MHz also failed to prevent crashes. Short OCCT CPU/VRAM tests, Unigine Superposition, and Premiere Pro GPU rendering can complete successfully.
The most useful comparison is that a GeForce GT 710 works in the exact same PC with the same CPU, motherboard, RAM, storage, PSU, and Windows environment. The crashes return when the GTX 1080 Ti is reinstalled.
Do "Physical Multiple Warp Errors" and "Illegal Instruction Encoding" across multiple TPCs usually indicate failing GPU silicon, or could faulty GDDR5X VRAM produce these errors? How much does the successful GT 710 comparison implicate the GTX 1080 Ti, and are there worthwhile longer-duration VRAM, core, or diagnostic tests that could distinguish a GPU, memory, PCIe, motherboard, or driver problem before declaring the card faulty?
3 Answers
For additional confidence, run a long VRAM-focused test rather than relying on a five-minute pass, and monitor for artifacting, driver resets, or rising temperatures. Testing with the card at stock settings and then with a substantial memory underclock can be useful, but neither a pass nor a failure is definitive. The event names describe what the GPU reported after execution went wrong; they do not by themselves identify whether the original cause was a shader core, VRAM, or another part of the card. If the crashes remain after the clean software and PCIe tests you have already performed, replacement or professional board-level testing is more practical than further driver troubleshooting.
The same-system comparison is strong evidence against Windows, the motherboard, and most driver issues. Swapping only the graphics card changes the result, and the errors are being generated by the NVIDIA display driver while the 1080 Ti is under real game workloads. It is still possible for a marginal power-delivery or PCIe interaction to affect only the more powerful card, but you have already tried another PSU, forced PCIe generations, updated the BIOS, confirmed x16 operation, and tested a clean installation. At this point I would treat the 1080 Ti as likely failing hardware.
This sounds most like deteriorating VRAM, which is a fairly common failure mode on older 10-series cards. A memory problem does not always show up in short VRAM tests; particular games may access memory patterns or regions that synthetic tests never hit. The fact that reducing the memory clock by 500 MHz did not help makes the card look even more suspect, although it does not completely rule out memory as the cause. If freezes leave audio running or inputs still produce game sounds briefly, that can also be a clue that the GPU is hanging while the rest of the game continues.

My 1080 Ti behaved similarly before it started showing visible artifacts and eventually stopped working. It passed some benchmarks and still failed in games, so successful short stress tests would not convince me that the card is healthy.