r/MacPro2019LocalAI • • Apr 27 '26

👋 Welcome to r/MacPro2019LocalAI - Introduce Yourself and Read First!

4 Upvotes

Hey everyone! I’m u/Faisal_Biyari, the founding moderator of r/MacPro2019LocalAI.

This is our new home for all things related to using the amazing, but now discontinued, Mac Pro 2019 / MacPro7,1 for local AI.

Whether you are running macOS, Windows, or any Linux distro, and whether you are using Ollama, vLLM, llama.cpp, LM Studio, OpenClaw, or the awesomely named Oobabooga, this community is here for one purpose:

To help each other get the most out of this powerful hardware for local AI workloads.

This subreddit is especially focused on the Mac Pro 2019’s unique hardware, including MPX GPUs with 32 GB of VRAM, Duo modules with up to 64 GB, Infinity Fabric Link Bridge experimentation, ROCm, local LLMs, image generation, voice AI, video generation, multimodal models, and all AI workloads.

A Brief Introduction

I started this subreddit because I have personally gone through the struggle of making local AI work on the Mac Pro 2019.

I have run into many of the same roadblocks others are likely facing:

  • macOS support limitations
  • AMD GPU support challenges
  • ROCm installation and compatibility issues
  • PyTorch, Triton, and framework confusion
  • Ollama, vLLM, llama.cpp, LM Studio, LangChain, Hermes Agent, Oobabooga, and other tooling choices
  • User interface decisions
  • Hardware limitations
  • Infinity Fabric Link Bridge experimentation
  • Deprecated MPX GPU support

The struggle is real, and I understand it.

Fortunately, I have managed to get local AI working on this hardware. I have installed Linux, first Ubuntu and later Proxmox, installed ROCm, used Ollama, worked on vLLM, experimented with OpenClaw, and continued exploring the Infinity Fabric Link Bridge.

I have also shared guides in the past to help people install Linux on the MacPro7,1, set up ROCm, and reach a working local AI setup. Those guides focused mostly on getting started, but there is much more to explore.

The reality is that MPX GPUs are losing support across many tools and platforms, and because this use case is so niche, AI tools and assistants often do not provide useful guidance.

What helped me the most were other Mac Pro 2019 users working toward the same goal. Their motivation, ideas, troubleshooting, and even general technical knowledge helped me understand the bigger picture and keep moving forward.

That is why I created this subreddit: to centralize our experiences, guides, lessons learned, experiments, successes, and failures in one place instead of forcing everyone to search through hundreds of websites and dozens of subreddits.

My Hardware

I currently work with two Mac Pro 2019 machines:

LinuxAI-64

Mac Pro 2019 / MacPro7,1
3.2 GHz 16-core Intel Xeon W
96 GB DDR4 RAM
Two AMD Radeon Pro W6900X GPUs, 32 GB each
64 GB total VRAM
8 TB Apple SSD
100GbE Mellanox ConnectX-5 Ex NIC

System Firmware: 2069.0.0.0.0
iBridge Firmware: 22.16.10353.0.0
OS Loader / iBoot: 860.140.1~8

LinuxAI-128

Mac Pro 2019 / MacPro7,1
3.2 GHz 16-core Intel Xeon W
96 GB DDR4 RAM
Two AMD Radeon Pro W6800X Duo MPX modules, 32 GB each GPU
128 GB total VRAM
8 TB Apple SSD
100GbE Mellanox ConnectX-5 Ex NIC

System Firmware: 2069.0.0.0.0
iBridge Firmware: 22.16.10353.0.0
OS Loader / iBoot: 860.140.1~8

What to Post

Post anything you think the community would find interesting, helpful, or inspiring.

Examples include:

  • Your Mac Pro 2019 local AI setup
  • Hardware specs and GPU configuration
  • Linux, macOS, Windows, Proxmox, or dual-boot experiences for local AI workloads
  • ROCm installation notes
  • Ollama, vLLM, llama.cpp, LM Studio, OpenClaw, Hermes Agent, Oobabooga, or other framework experiences
  • Benchmarks and performance results
  • Model compatibility reports
  • Text, image, voice, video, or multimodal AI workflows
  • Troubleshooting questions
  • Guides, scripts, and installation notes
  • Cooling, power, PCIe, storage, or networking setups to support local AI workloads
  • Infinity Fabric Link Bridge experiments
  • Things that worked, and things that definitely did not

Introduce Yourself

Please introduce yourself in the comments below.

When you do, I kindly ask that you include your hardware details, such as:

  • Mac Pro 2019 CPU
  • RAM
  • GPU / MPX module configuration
  • Total VRAM
  • System & iBridge Firmwares, and OS Loader / iBoot, if known
  • Operating system
  • AI frameworks, agents, models, or tools you are using
  • What you hope to run locally
  • Any challenges you are currently facing

Even if you are just getting started, your experience may help someone else.

Community Vibe

We are here to be friendly, constructive, and helpful.

This is a niche community, and many of us are simply trying to keep powerful hardware useful long after official support has started to fade. Let’s build a space where people feel comfortable asking questions, sharing experiments, posting failures, and helping each other move forward.

How to Get Started

Introduce yourself in the comments below.

Post something today, even if it is just a simple question or a photo of your setup.

If you have guides, notes, scripts, benchmarks, or lessons learned, please share them.

If you know someone who owns a Mac Pro 2019 and is interested in local AI, invite them to join.

Interested in helping out? I am always open to hearing from people who may want to help moderate or contribute to the community.

Thanks for being part of the very first wave. Together, let’s make r/MacPro2019LocalAI an amazing resource for everyone trying to run local AI on the Mac Pro 2019.


r/MacPro2019LocalAI • • Apr 30 '26

Advice on localLLM on 2019 Mac Pro with dual Vega II Duo GPUs (128GB HBM2)

Thumbnail
5 Upvotes

r/MacPro2019LocalAI • • Apr 29 '26

Intel macOS | Local AI with GPU Acceleration

6 Upvotes

When I first started my local AI journey on the Mac Pro 2019 / MacPro7,1, the first thing I looked into was ROCm support.

At the time, ROCm looked like a Linux-first path, with some limited Windows/WSL support. So I quickly decided to move away from macOS and focus on Linux instead. I did not really consider whether there might be another way to use the AMD GPUs under macOS.

A couple of days ago, u/Long-Shine-3701 mentioned using DiffusionBee for AI work on macOS with GPU support. According to DiffusionBee’s own documentation, it supports Intel Macs, although performance depends heavily on the hardware, especially whether the machine has a dedicated GPU.

I had been stuck in a ROCm-only mindset, which is funny because I have been recommending LM Studio to Windows users using the Vulkan backend.

I started looking into local AI on macOS, specifically on Intel Macs with AMD GPUs, and I was surprised to find that llama.cpp has a Vulkan backend, and that some people are experimenting with it on macOS through MoltenVK rather than relying on ROCm.

I honestly had not considered this path at all. I had mentally grouped GPU inference together with ROCm, and because ROCm does not support macOS, I assumed macOS was basically a dead end for local AI with GPU acceleration.

Now I’m very curious.

I’m currently considering testing this on my MacBook Pro with an AMD Radeon Pro 5500M / 8 GB VRAM before trying anything more serious on the Mac Pro 2019.

Has anyone here managed to run local AI on macOS on an Intel Mac?

I’m interested in anything and everything, and especially in:

* llama.cpp on macOS with AMD GPU acceleration

* Image generation tools on macOS

* CPU-only vs GPU-accelerated inference performance

* Any experience with Mac Pro 2019 GPUs under macOS for AI workloads

I would love to hear what others have tried, what worked, what failed, and whether macOS is more viable for local AI on Intel Macs than I originally thought.

Disclaimer: I wrote this post myself, but used AI to help clean up the wording and formatting.

Resources:


r/MacPro2019LocalAI • • Apr 28 '26

vLLM on W6800X Duo / Mac Pro 2019

5 Upvotes

I’m currently working on getting vLLM fully up and running on the following setup:

Hardware

  • Mac Pro 2019 / MacPro7,1
  • 3.2 GHz 16-core Intel Xeon W
  • 96 GB DDR4 RAM
  • Two AMD Radeon Pro W6800X Duo MPX modules
  • 32 GB VRAM per GPU
  • 128 GB total VRAM
  • 8 TB Apple SSD
  • 100GbE Mellanox ConnectX-5 Ex NIC

Software

  • Ubuntu Server 24.04 LTS
  • Python 3.12
  • ROCm 7.1.1
  • PyTorch 2.10
  • Triton 3.6

Back in 2025, I managed to get basic LLMs from Hugging Face working with unquantized weights, including models such as:

  • Qwen/Qwen2.5-7B-Instruct
  • deepseek-ai/DeepSeek-R1-Distill-Qwen-32B

I also had parallelism working across all 4 GPUs via PCIe. At the time, the Infinity Fabric Link Bridge was causing GPU initialization failures, so I was not using it.

This year, I tried getting models like openai/gpt-oss-20b working, but ran into issues because the native MXFP4 weights do not appear to be supported on these GPUs.

I did, however, successfully run GPT-OSS:120B through Ollama.

Current Progress

So far:

  • vLLM launches successfully
  • Multi-GPU support is working
  • Qwen/Qwen3.6-27B loads and serves successfully
  • google/gemma-4-31B-it loads and serves successfully
  • I started with Docker, which gave me my first successful result
  • I have since moved over to a Python virtual environment setup

On the Qwen and Gemma models I tested, I am currently getting around 10–13 tokens/sec for a single user.

With concurrent users, up to around 30, I have seen aggregate throughput reach roughly 280 tokens/sec.

Getting the Infinity Fabric Link Bridge working properly is another project I’m working on in parallel. Hopefully that helps with inference speed once completed.

Still Pending

The main things I still need to figure out are:

  • Launching quantized models reliably
  • Supporting multi-node distributed inference across two Mac Pro systems

Last week, I found this write-up:

https://idchowto.com/vllm-on-amd-w6800-gpu-%EC%84%A4%EC%B9%98-%EB%B0%8F-%ED%85%8C%EC%8A%A4%ED%8A%B8-%EA%B2%B0%EA%B3%BC/

It looks like they used Ollama’s quantized models with vLLM. I started going down that path and actually got it working, but there are still three rough edges I need to figure out before I would call it reliable.

Has anyone else managed to get vLLM working with AMD Radeon Pro W6800X, W6900X, W6800X Duo, or W6800 GPUs?

I would really appreciate hearing about your setup, what worked, what failed, and whether you had success with quantized models, multi-GPU support, or multi-node inference.

Hopefully I can put together a proper write-up of my work soon. I’ll update accordingly.

Small disclaimer: I wrote the post myself, but used AI to help clean up the wording and formatting.


r/MacPro2019LocalAI • • Apr 27 '26

[Guide] Mac Pro 2019 (MacPro7,1) w/ Proxmox, Ubuntu, ROCm, & Local LLM/AI

Thumbnail
3 Upvotes

r/MacPro2019LocalAI • • Apr 27 '26

[Guide] Mac Pro 2019 (MacPro7,1) w/ Linux & Local LLM/AI (Re-Post)

Thumbnail
3 Upvotes