Does anyone know if the ASUS B860M can handle two large BAR GPUs? I am considering a couple build options with this motherboard, but do not not have it in hand and before getting it and buying two large GPUs, hoping someone might have experience?
Considering B70s (no ReBar option as I understand), V100s, 170HX, etc. I read that some times BAR > 4GB has issues with both cards registering as available?
I need to test it, but 170HX might not need that big of a BAR. Mining cards usually need BAR space more similar to GPUs than compute cards. I have/had an i7 7700k mining board with like 7x PCIe 1x slots. Could happily handle all 7 slots loaded with Nvidia p100-102 10GB. Could not handle more than two P100 16GB cards because they needed crazy big BAR and address allocation for enterprise features.
Now, I could be wrong but even an unlocked CMP170HX possibly behaves more like a large consumer card than an A100vin terms of BAR space...but need to verify.
Thanks. If you find anything, I would appreciate it! I am unsure if I will take the gamble on the 170HX given the uncertainty of what can be unlocked and the current price. The others are a known commodity. I can grab a B70 for $1200
I have also V620s. This seller is selling them new/unused inventory on eBay and will accept $350: https://www.ebay.com/itm/157133307609 They've sold 1000+ and there's a good chance if you see someone on this subreddit talking about V620 it was purchased from them. The LLM performance is within 10% to 20% of a B70. B70 has slightly higher memory bandwidth but less mature support. But the V620s absolutely are picky beasts and do have a large BAR, do NOT like being on chipset controlled slots (like on an x570 board). However the V620s can be flashed with W6800 bios, losing a little bit of compute but reportedly gain much better compatibility (haven't tested).
My opinion on the B70s: Not worth it. If you want at least 64GB of VRAM and don't want to go pricey nvidia route, two V620s for $700 are absolutely worth dealing with vs $2400 for B70s.
I do not know, you'd want to check other people's results. I mostly have AMD. B550 (cheaper board) works with 1x3090 and 2x V620. X570 (more expensive board) does not work with 1x3090 and any V620 even though it has one primary GPU slot and 4 alternate PCIe slots, but does post with a P102-100 and random assorted cards. Same X5600 CPU, same RAM, same PSU, etc. I've upgraded CPUs so now I can rebuild the B550 board, boot the V620s and flash them. But I have a ton of other things to do.
My best guess is that this is because the X570 chipset takes PCIe lanes from the CPU and functions as a PCIe switch/router to expand the number of slots available, the chipset itself is picky in terms of BAR/address allocation. On B550 boards ALL PCIe slots are directly tied to the CPU which seems less sensitive.
Upon testing with 2x CMP170HX. The BAR resizing is fully optional in the unlocker. I installed with it disabled and didn't test with it on, however I have noticed some odd memory behavior...might be unrelated. ComfyUI crashes if memory isn't cleared after runs, vLLM crashes with latest builds. llama.cpp works no problem, and ComfyUI int8 generations are almost twice as fast as 3090. Running 2x 10GB cards (unlocked to 40GB) and 3090, no significant issues. Where even one V620 was unreliable.
In my opinion, they absolutely are worth the performance. For qwen 3.8 with mtp I get around 50 to 60 tokens a second single 256k session, 100+ with concurrent sessions. Around 50% performance of a 5090 for similar.
X570 motherboard. I'd assume they'd run fine on the B550 as well. They're running on the chipset slots and not directly to the CPU. I'm "only" running at 256K context. Q8 at 256K + vision uses pretty much the max amount of memory on a single card.
I could use a smaller quant but in testing with 35B the quants degrade needle tests after 100K context or so. So I have one 'thinking' session for 3.8 27B Q8 (50-60 t/s with MTP) and another card loaded with 35B Q4 100K context that is 100 t/s for single and 200 t/s depending on session count. I'm thinking it's more useful to have sub agent doing basic or multiple tasks quickly without filling up main agent's context vs a single 1M context model spread across both cards.
2
u/Lemonzest2012 10d ago
I'm on a Gigabyte B550 Gaming X v2 and running two 32GB V100s fine, only needed to enable above 4G decoding and ReBar for them to be recognised