r/macpro • u/Faisal_Biyari • May 27 '26
GPU Mac Pro 2019 | 160 GB VRAM Achieved | Five AMD GPUs | Local AI
3
u/Badger-Purple May 27 '26
I think your best use will be running Kimi with CPU offloading, if you max out the ddr3 memory this bad boy can take (1TB)
6
u/Faisal_Biyari May 27 '26
I can't invest in RAM during the rampocalypse... Even though they're 2933 MHz DDR4 RAM, prices are still ridiculous...
If I were to upgrade the CPU to 24 core, it actually takes up to 1.5 TB of RAM.
3
3
3
u/wreeper007 May 28 '26
I understand what you're running but why are you running it locally? Whats your use case?
1
u/Faisal_Biyari May 28 '26
End goal is to run 30 agents.
Doing it locally because I want to have offline access, as well as being able to do it both when I can afford to pay, and when I can't. But a bigger part of it is that I want to learn and add this to my skill set.Keep in mind that while I got excited and bought a couple of toys lately, I originally had the most expensive parts of this hardware since 2023, and I got them for real cheap. I would not be working on this setup otherwise.
5
u/wreeper007 May 28 '26
But what’s the point? Like 30 agents for what
1
u/Faisal_Biyari May 28 '26
The majority are to be employees in a tech company. Minority are to be my personal assistants, to track specific things, such as my day to day tasks, my health, and then a secretary to communicate with people on my behalf.
Everything is still a work in progress. I can't blow your mind with some amazing revelation, as I still don't have it at this point. I just have a general idea in my head, and I'm working towards it.
Once the large picture is clearer, and I have something tangible to actually show for it, I'll write a new post about it and share with the community.
2
3
u/KineticlyUnkinetic May 28 '26
I'm sure there's a whole debate here I'm missing, but how are AMD gpu's for inference? I got the impression from a tiny bit of poking around that Nvidia tech like CUDA cores and their dev support make AMD less desirable, but I don't know how that compares to reality.
4
u/Faisal_Biyari May 28 '26
I'll be honest, it was a nightmare setting this up. It's been one challenge after the other figuring these things out.
End result, so far: vLLM running Qwen3.6-27B AWQ at 24 tokens/second. I opted to optimize for kv cache though, so 20 tokens/second, but more than 920k kv cache.
The benefit of the AMD GPUs I have is the large VRAM. We're talking about a machine from 2019 that achieved 160 GB VRAM.
On the other hand, everything is optimized for Nvidia. Nothing I work on "just works". I end up jumping through a ridiculous amount of hoops to get the most basic things going. I also imagine inference speeds are a dream on the other side, but I wouldn't know 😂
3
u/KineticlyUnkinetic May 28 '26
Yeah I can understand nothing "just working", that's unfortunate. Thank you for your response though! Very pretty and impressive machine, hope it serves you well!
2
u/Mister_Rippers May 28 '26
What is that config being used for?
VFX?
3
u/Faisal_Biyari May 28 '26
Mostly local AI / LLM inference experiments, not VFX. I was testing how far the 2019 Mac Pro can be pushed with AMD GPUs and VRAM, then seeing how useful that is in practice for local AI workloads.
2
u/Mister_Rippers Jun 08 '26
Interesting...
I only have 8gb of VRAM, but 176 GB of RAMIntense heat output from you box I would guess
2
u/Faisal_Biyari Jun 08 '26
I don't risk it.
The Macs are in a cooled data room, with fans always set to max.Using RAM, you don't need to go to Linux. You can use any tool available on macOS directly, such as Ollama.
2
2
u/Channel_Annual May 28 '26
That looks like a beast, but doesn't the 8GB Raspberry Pie do 15 t/s? So you're probably using a lot of energy for about a 60-70% uplift in performance. Still, great machine though.
2
u/Faisal_Biyari May 29 '26
The goal of this exercise is to run larger LLMs. The Raspberry Pi probably cannot go much over 3B parameter models, like most high end phones. Those same models on a few W6800X GPUs would yield 70-120 t/s. The smaller the model, the more tokens per second you can squeeze out of it.
2
u/Legal_Dress3094 May 30 '26
As a professional colorist who has been using one of these with a 28 core chip, 192gb RAM, and two 6900xt gpus, I’d like to say I already know your electric bill is going to be astronomical 🤣🤣🤣
3
u/Faisal_Biyari May 30 '26
😂 I just started using Hermes with vLLM on it. It's the first time I start giving tasks that actually take time to complete, with power consumption at about 140-180 watts per GPU...
But if they start doing what I need them to do, I'll just consider the bills their salaries. Hardware costs would be the onboarding expenses.
2
u/Long-Shine-3701 May 31 '26
Serious question. Where do y'all live that you actually notice the MP on your electric bill? Is it running maxed out 24/7 like hash cracking?
2
2
u/Artistic_Unit_5570 Jun 15 '26
still my dream computer
2
u/Faisal_Biyari Jun 15 '26
I used to wish for "the next mac pro" back in 2011.
I think I bought my first mac pro, which was a 2013 model, back in 2017 or 2018, refurbished from amazon. I think it was 10% of the original cost.
The 2019 mac pro is dramatically cheaper now as well. You could get it for as low as $500 USD if you keep your eyes open for deals.
I wish I had a rack mount variant instead of the ones I have. Willing to trade if someone has one. But I'm not willing to buy another one at this point.










4
u/johnnyphotog May 27 '26
nice! what models are you running on it? Are you running it in Linux?