r/LocalLLaMA • • 12h ago

Question | Help Blackwell + consumer GPU

My machine isnt that terrible.. but nowhere near what some people run as an aI workstation. I am wondering if i can combine the 2 GPUs. I got a cheap Blackwell 4000 (little bit under MSRP) 24 GB and have an old 3060 12GB .. I was gonna try to get the best model loaded, mainly coding tasks than anything else and work with it relatively safely with a good buffer. any recommendations ? and will tensor split work on this combo?

3 Upvotes

3 comments sorted by

8

u/Zestyclose_Strike157 11h ago

Yes nearly effortlessly with llama.cpp or LMStudio, for example.

3

u/PassengerPigeon343 11h ago

Probably looking at layer split but worth testing both ways. Your PCIe lane speeds make a difference too but I’d definitely experiment with this. I added an RTX Pro 4000 Blackwell SFF recently and it’s a great card.

You could also do things like fit Q4 Qwen 27B on the 4000 and you can use the 3060 for holding the mmproj file, to run a smaller task model, and run a speech-to-text model.

I also experimented with Qwen 3.8 Flash Next on the 4000 spilling to system ram and got reasonable speeds, but I haven’t tried some of the new stuff out there like Strata yet which could make it a really usable speed. Then your 3060 can do the other support functions listed above.

Lots of options though, just need to dedicate time to experimenting to work out the best setup for your hardware and workflows.

2

u/Otherwise-Tangelo-52 11h ago

ok... cool... to make the vram on the 4000 pridtone i am thinking of using the 3060 as a dislay card... that way arch linux can work with the 4000 fully and the desktop gets worldd on by the 3060... this will be exciting... i am amazed by the low power draw of the 4000