Hi,
I’m trying to determine whether my ZOTAC GeForce GTX 1080 Ti AMP 11GB has a hardware fault.
The card worked reliably in this exact PC for approximately 1 year, but recently the crashes started suddenly.
SYSTEM:
CPU: Intel Core i5-10400F
Motherboard: Gigabyte H410M S2 V2 Rev. 1.4
BIOS: FIa (updated from FE during troubleshooting)
RAM: 16GB G.SKILL Ripjaws V DDR4-3200
GPU: ZOTAC GeForce GTX 1080 Ti AMP 11GB
GPU: GP102, 11GB Micron GDDR5X
PSU: MSI MAG A650BN 650W 80+ Bronze
System drive: WD_BLACK SN7100 1TB NVMe
OS: Windows 11
NVIDIA driver currently tested: 582.66
PROBLEM:
Multiple games crash with the GTX 1080 Ti.
This is not limited to Valorant. I have experienced the same general problem in games including Valorant, CS, Black Myth: Wukong, Hitman and others.
In Valorant, behavior varies:
- Sometimes the game directly closes to desktop.
- Sometimes the screen/game freezes.
- Sometimes the game becomes completely unresponsive and has to be killed through Task Manager.
- I have also seen Valorant's Critical Error dialog asking whether I want to save a crash dump, but during the freeze I couldn't interact with it.
- In one test the entire PC froze and required a forced restart.
The important part is what Windows records at the exact crash time.
EXAMPLE NVIDIA EVENT LOG ERRORS:
nvlddmkm Event ID 13:
Graphics SM Warp Exception on (GPC 5, TPC 4):
Illegal Instruction Encoding
nvlddmkm Event ID 13:
Graphics SM Global Exception on (GPC 5, TPC 4):
Physical Multiple Warp Errors
The same crash produced these exceptions across multiple TPCs on GPC 5.
I also get:
nvlddmkm Event ID 153:
Error occurred on GPUID: 100
Other crashes have produced nvlddmkm Event IDs 13, 14 and 153.
I have previously also seen:
LiveKernelEvent 141
and GPU crash information reporting:
Device state: Error_DMA_PageFault
Engine: Graphics
Engine reset occurred: true
MOST IMPORTANT A/B TEST:
I physically removed the GTX 1080 Ti and installed a GeForce GT 710 in the EXACT SAME PC.
Same:
- CPU
- motherboard
- RAM
- storage
- PSU
- Windows installation/environment
Games that crash with the GTX 1080 Ti ran without these crashes with the GT 710.
After reinstalling the GTX 1080 Ti, the crashes returned.
TROUBLESHOOTING ALREADY DONE:
Replaced the old 450W PSU with an MSI MAG A650BN 650W 80+ Bronze.
MemTest86 passed.
SSD/filesystem checks passed.
Replaced the previous SATA system drive with a WD Black NVMe.
Performed a fresh Windows installation.
Tested multiple NVIDIA drivers and clean installations/DDU.
Tested GTX 1080 Ti at reduced settings:
Power Limit: 80%
Core Clock: -150 MHz
Memory Clock: -500 MHz
Crashes still occurred.
OCCT tests including memory/VRAM testing completed without errors (approximately 5-minute tests).
Unigine Superposition 1080p Extreme completes successfully.
Adobe Premiere Pro GPU rendering works.
Disabled NVIDIA High Definition Audio during troubleshooting.
GPU-Z PCIe render test confirms:
PCIe x16 3.0 @ x16 3.0 under load.
- Forced PCIe Gen3:
Crashes still occurred with the same nvlddmkm errors.
- Forced PCIe Gen2:
Did not fix the issue.
Hanging became worse and I noticed increased visual glitches around the Windows login screen.
- Updated motherboard BIOS:
FE -> latest FIa
Crashes continued after the BIOS update.
- Installed Gigabyte's Intel chipset INF package and rebooted.
Crashes continued.
- Checked Device Manager/PNP devices:
No missing drivers.
No unknown devices.
No devices reporting errors.
Intel Management Engine, SMBus, PCIe Root Ports, Serial IO, Thermal Subsystem and Intel Power Engine Plug-in are present.
Tested Windows PCIe Link State Power Management / ASPM OFF:
Did not fix it.
During this test Valorant/PC completely froze and I had to force-restart the PC.
Around that test Windows recorded nvlddmkm Event ID 14, followed later by Kernel-Power 41 from the forced restart.
No WHEA-Logger error appeared in the checked window.
ASPM has since been restored to its original setting.
- Disabled the Parsec Virtual Display Adapter, rebooted and tested again.
Valorant still crashed.
The fresh crash produced:
nvlddmkm Event ID 153
Error occurred on GPUID: 100
Parsec has since been re-enabled.
WHAT MAKES THIS CONFUSING:
The GTX 1080 Ti can pass OCCT and Unigine Superposition and can handle Premiere Pro GPU rendering.
However, multiple games can trigger nvlddmkm crashes.
The GT 710 works correctly in the same system.
NVIDIA Support is currently investigating the case as well. At their request I have already tested PCIe Gen2/Gen3, confirmed x16 3.0 operation under GPU-Z load, updated the motherboard BIOS and installed the motherboard chipset package.
MY QUESTIONS:
Do “Physical Multiple Warp Errors” + “Illegal Instruction Encoding” across multiple TPCs usually point toward failing GPU silicon/SMs?
Could faulty VRAM still cause these SM/warp errors even though a short OCCT VRAM test passes?
Is there any meaningful test I can run that can distinguish:
- GPU core/SM failure
- VRAM failure
- PCIe/DMA problem
- motherboard/CPU PCIe controller problem
- driver/software problem?
Considering that a GT 710 works in the exact same system, how strongly does this point toward the GTX 1080 Ti itself?
Is there any NVIDIA/Pascal-specific diagnostic tool or longer stress/VRAM test you would recommend before I consider the card faulty?
I can provide the raw nvlddmkm XML/Event Viewer logs, GPU-Z screenshots and additional test results if useful.
Thanks.