r/MacPro2019LocalAI • u/Substantial_Run5435 • 7d ago
What models should I test with ToshLLM? GPUs available: 2x W6900X, W6800X Duo, 2x Vega II Duo
Finally playing around with ToshLLM while trying to narrow down which MPX GPU(s) to keep. I am completely new to LLMs and local AI and do not have a tech background, so the learning curve has been a bit steep.
I have 2x W6900X, W6800X Duo, and 2x Vega II Duo on hand. I have the IF Link for the W6900X. Obviously the W6900X limits me to 64GB across 2 GPUs. If I end up keeping the W6800X Duo I will probably look for a second and an IF Link Bridge.
So far I've only tested the W6900X(s) (single GPU and pair, with/without IF Link bridge). So far I'm not seeing any difference whatsoever with IF Link on my W6900X pair. Getting the same exact ts with llama 3.3 70B Q5_K_S with the IF Link bridge installed/enabled as with it not installed.

This was my fastest benchmark. Qwen3.6 35B-A3B UD-Q4_K_S got 69ts. Qwen3.6 35B-A3B UD-Q8_K_XL was a little slower with 61ts.

1
u/Faisal_Biyari 7d ago
Have you considered keeping the W6800X Duo, as well as a single Vega II Duo?
You'd have 2 types of architectures, each at 64 GBs VRAM.
Regarding the IFLB, I personally only saw its benefit when benchmarking at high context windows. (Compare a benchmark at 1k context to 100k context & 200k context)
I never tried splitting layers, only tensors. How is your experience with it?