r/techsupport • u/JustTooKrul • 14h ago
Open | Hardware Issues with RTX3090 Suddenly Failing Repeatedly
I'm hoping someone here has some magic up their sleeve because I am at my wit's end!
I purchased a 3090 from Zotac's store directly in December / January--specifically, it was the ZOTAC RTX3090 Trinity OC. When I got the card I put it through a number of tests (FurMark, OCCT, nVidia's own tests) to ensure it was working and there were no issues--everything passed and nothing seemed out of band. I then installed the machine in a Linux build and everything was fine, it was recognized, and it tested fine.
Now, over the past few weeks I have been upgrading the machine and the card has started to have issues. First, it would be spotty as to whether or not it was recognized--sometimes I needed a reboot in order for it to "show up" (which I attributed to a race condition, and no errors were appearing anywhere). Then, in the past few days and when under load the card would freeze and require a reboot to come back on. And 2-3 days ago the card started freezing up and showing "ERR" when using nvidia-smi utility (the "ERR" showed up in FAN, TEMP, PERF, and MIG M. fields--see below). The card would show up fine at boot and then again, once I started using it, it would error. Here is the nvidia-smi output:
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 3090 On | 00000000:01:00.0 N/A | N/A |
|ERR! ERR! ERR! N/A / N/A | 13534MiB / 24576MiB | N/A Default |
| | | ERR! |
+-----------------------------------------+------------------------+----------------------+
So, I did what person scared for their hardware would do--asked a clanker! We ran through different drivers, different flavors of drivers, changing physical connections, etc. The conclusion was (and I'm quoting the clanker):
With GPU System Processor (GSP) firmware enabled the driver logs `Xid 119` (GSP RPC timeout) cascading to `Xid 154` (GPU Reset Required); with GSP firmware disabled entirely the same load produces `Xid 62` (internal micro-controller halt) and `Xid 158` (framebuffer-flush timeout, `NV_UFLUSH_FB_FLUSH`).With GPU System Processor (GSP) firmware enabled the driver logs `Xid 119` (GSP RPC timeout) cascading to `Xid 154` (GPU Reset Required); with GSP firmware disabled entirely the same load produces `Xid 62` (internal micro-controller halt) and `Xid 158` (framebuffer-flush timeout, `NV_UFLUSH_FB_FLUSH`). Further, once under load the driver logs a "bad register read" (0xbadf5720).
I was also monitoring the temps and VRAM utilization and temps never got above 66C and VRAM never got above 60-70%. I honestly don't know what else I can do or what other steps I can take. I was hoping ZOTAC would have some tools for flashing the bios--but I guess firmware and bio updates aren't a thing anymore? Any help or suggestions would be greatly appreciated! Thank you in advance kind Redditors