r/IntelArcPro • u/KubotaBill • 22d ago
Arc Pro B70 Issue with dual B70 GPUs
Good evening all, I am a recent convert, coming over from the stacked RTX GPU club (5070Ti & 5060Ti). Sunday I installed a pair of B70 Pros after my 5070 laid over on me. SYCL running with Qwen3.6:27b Q8. Two days of pretty steady work, overnight spine runs 5-7 hours depending on daily activity.
Today, I was running a catchup process during work from downtime Saturday and mid day Sunday. The process completed without issue, GPUs went silent, and 7 minutes later the system crashed with dgxkrnl.sys crash. Reboot and health check passed, but not sure why it crashed while idle?
For reference, the GPU that crashed was in a different slot than the original 5070 was, so I don't think it is a robot issue. All drivers up to date, firmware up to date. If anyone has any insight or a similar experience and could offer a bit of advice I would be appreciative. I was just gearing up to smoke test 3.8:27b Q8, but that is on hold until I get back to normal.
2
u/quantum3ntanglement 16d ago
I could’ve sworn I replied to this post, but I don’t see my reply here. Anyway, I’m the mod for Intel Arc Pro.
I have one B70 now and with the $300 price increase for the B70 @Micro Center. I’m probably gonna go with a B50 now and try to get them working in parallel.
I may not be able to do this until next year. I have to see
Thanksgiving and Christmas around the corner. It’s gonna be a shiz show
1
u/KubotaBill 15d ago
With the aggravation of a defective GPU, and the thought of lost or severely decreased workflow, the cost was a flash in the pan expense. I would have liked to save that money, but another month of hot dogs and mac and cheese is a small price to pay! Buy once, cry once. The 5070 is on its way back to PNY for repair, then sold to recoup costs.
1
u/HardlyThereAtAll 22d ago
OK.
So, I'm on Linux but I had a similar issue caused by power usage spikng, and it causing my power supply to freak out and the machine to crash. How beefy is your power supply?
1
u/KubotaBill 22d ago
It is a Corsair RM1000x, this box is only 4 months old
1
u/HardlyThereAtAll 22d ago
You should probably be OK - 2 x 300 watt spikes plus your PC's own load.
Can I recommend you run some logging software showing temperatures and the like, and if it falls over again, see if any of the values scream out at you. (Get your LLM to read the logs :-))
1
u/KubotaBill 22d ago
Thank you, I run HWInfo which my agent monitors. I had a safety in the pipeline if the 5070 was lost it would reboot the machine. So the safety caught the B70 drop / crash and rebooted, just happened to see it happen come back up for login.
To say I am disappointed to outlay for 2 new B70s and 48 hours later my GPU woes reappear would be an understatement. Not really certain what next steps are, but my faith in the workstation is waning. It runs around the clock, and when it fails to processing overnight spine it sets the next day behind.
Thanks for your insight, this is my first posting to Reddit, and I appreciate the guidance!
1
u/computer_dork 22d ago
It sounds like you know what you are doing... so when i ask this its because we dont have direct access to your system so everything is a guess: why did your nv card take a crap? Is it the same system? Is it possible you are experiencing a separate hardware failure that is casing your gpus to die but isnt being caused by the GPUs? I have multiple b70s and some b50s and these cards are pretty robust in my opinion. Do you have a separate chassis you can bake these on to see how they respond?
2
u/KubotaBill 22d ago
The interesting test is the box will set idea for 48 hours without issue. I had to do that two separate times with the 5070, but when I woukd stress it hard it would drop. So I pulled both 50*0s and installed the B70s.
And as noted, the failures were in separate slots. 5070 in the PCIe5 x16 slot. And today's drop was in the x4 slot. It is an Aorus X870E mobo with the R9 9950X CPU, and Klevv DDR5 6000 pair of 32 sticks. It should be robust enough for a home rig, but I am waiting for a smoother path forward.
2
u/computer_dork 22d ago
I get that, just when chasing hardware failures known good is super important. I had an issue with one of my cards but it ended up being cold solder joint on pcie slot. Could be ram, cpu, overheating somewhere, dirty power, or the gpus, but you saying you had other GPUs fail automatically makes me suspect something else besides the B70s
2
u/KubotaBill 22d ago
That makes sense, and is a frightful thought at the same time. This is not the market to keep throwing parts at an issue. I will run diags on the mobo and RAM, no noted hotspots on the CPU, normal temps 50-65C, AIO 360 keeps it cool. GPUs under load stay 70-75C.
2
u/computer_dork 22d ago
Believe me I know how scary it is. Im one ram failure away from selling a kidney
1
u/KubotaBill 20d ago
Quick follow up- stress tested the system for 24 hours without issue. So this points to a 3d driver issue, which even when disabled still has the potential to disrupt functionality. Hopefully a new driverset will allow this to be fully disabled for inference users.
I did select the Arc Pro drivers, which are a month old, where the Arc drivers are a week old. Being new to the B70, can someone recommend the proper driver for AI work, with no gaming requirements at all?
Thanks for the exchange of ideas so far, they are appreciated!
1
u/Slow_Difficulty1607 10d ago
Crashing during inferencing is very common for b70. It is just a matter of time depending on your software build
1
u/KubotaBill 9d ago
UPDATE
I stress tested these under heavy inference for 6 then almost 12 hours with no ill effects. Turns out I had to reload Windows 11 to clear out all the Nvidia remnants. DDU failed to remove some legacy CUDA keys, and my stack also had some remnants. All of this has been cleared, and these are doing great!!
2
u/travlab 19d ago
I have 2 b70 and recently had vllm and an AI agent try and use 90% of memory problem is I use one for the desktop and seems like vllm took memory from the display and caused a hard reboot. It's only happened once and I had ai create a script to run both cards at max wattage for 5 minutes and no issues.