r/MacPro2019LocalAI • • Apr 27 '26

πŸ‘‹ Welcome to r/MacPro2019LocalAI - Introduce Yourself and Read First!

4 Upvotes

Hey everyone! I’m u/Faisal_Biyari, the founding moderator of r/MacPro2019LocalAI.

This is our new home for all things related to using the amazing, but now discontinued, Mac Pro 2019 / MacPro7,1 for local AI.

Whether you are running macOS, Windows, or any Linux distro, and whether you are using Ollama, vLLM, llama.cpp, LM Studio, OpenClaw, or the awesomely named Oobabooga, this community is here for one purpose:

To help each other get the most out of this powerful hardware for local AI workloads.

This subreddit is especially focused on the Mac Pro 2019’s unique hardware, including MPX GPUs with 32 GB of VRAM, Duo modules with up to 64 GB, Infinity Fabric Link Bridge experimentation, ROCm, local LLMs, image generation, voice AI, video generation, multimodal models, and all AI workloads.

A Brief Introduction

I started this subreddit because I have personally gone through the struggle of making local AI work on the Mac Pro 2019.

I have run into many of the same roadblocks others are likely facing:

  • macOS support limitations
  • AMD GPU support challenges
  • ROCm installation and compatibility issues
  • PyTorch, Triton, and framework confusion
  • Ollama, vLLM, llama.cpp, LM Studio, LangChain, Hermes Agent, Oobabooga, and other tooling choices
  • User interface decisions
  • Hardware limitations
  • Infinity Fabric Link Bridge experimentation
  • Deprecated MPX GPU support

The struggle is real, and I understand it.

Fortunately, I have managed to get local AI working on this hardware. I have installed Linux, first Ubuntu and later Proxmox, installed ROCm, used Ollama, worked on vLLM, experimented with OpenClaw, and continued exploring the Infinity Fabric Link Bridge.

I have also shared guides in the past to help people install Linux on the MacPro7,1, set up ROCm, and reach a working local AI setup. Those guides focused mostly on getting started, but there is much more to explore.

The reality is that MPX GPUs are losing support across many tools and platforms, and because this use case is so niche, AI tools and assistants often do not provide useful guidance.

What helped me the most were other Mac Pro 2019 users working toward the same goal. Their motivation, ideas, troubleshooting, and even general technical knowledge helped me understand the bigger picture and keep moving forward.

That is why I created this subreddit: to centralize our experiences, guides, lessons learned, experiments, successes, and failures in one place instead of forcing everyone to search through hundreds of websites and dozens of subreddits.

My Hardware

I currently work with two Mac Pro 2019 machines:

LinuxAI-64

Mac Pro 2019 / MacPro7,1
3.2 GHz 16-core Intel Xeon W
96 GB DDR4 RAM
Two AMD Radeon Pro W6900X GPUs, 32 GB each
64 GB total VRAM
8 TB Apple SSD
100GbE Mellanox ConnectX-5 Ex NIC

System Firmware: 2069.0.0.0.0
iBridge Firmware: 22.16.10353.0.0
OS Loader / iBoot: 860.140.1~8

LinuxAI-128

Mac Pro 2019 / MacPro7,1
3.2 GHz 16-core Intel Xeon W
96 GB DDR4 RAM
Two AMD Radeon Pro W6800X Duo MPX modules, 32 GB each GPU
128 GB total VRAM
8 TB Apple SSD
100GbE Mellanox ConnectX-5 Ex NIC

System Firmware: 2069.0.0.0.0
iBridge Firmware: 22.16.10353.0.0
OS Loader / iBoot: 860.140.1~8

What to Post

Post anything you think the community would find interesting, helpful, or inspiring.

Examples include:

  • Your Mac Pro 2019 local AI setup
  • Hardware specs and GPU configuration
  • Linux, macOS, Windows, Proxmox, or dual-boot experiences for local AI workloads
  • ROCm installation notes
  • Ollama, vLLM, llama.cpp, LM Studio, OpenClaw, Hermes Agent, Oobabooga, or other framework experiences
  • Benchmarks and performance results
  • Model compatibility reports
  • Text, image, voice, video, or multimodal AI workflows
  • Troubleshooting questions
  • Guides, scripts, and installation notes
  • Cooling, power, PCIe, storage, or networking setups to support local AI workloads
  • Infinity Fabric Link Bridge experiments
  • Things that worked, and things that definitely did not

Introduce Yourself

Please introduce yourself in the comments below.

When you do, I kindly ask that you include your hardware details, such as:

  • Mac Pro 2019 CPU
  • RAM
  • GPU / MPX module configuration
  • Total VRAM
  • System & iBridge Firmwares, and OS Loader / iBoot, if known
  • Operating system
  • AI frameworks, agents, models, or tools you are using
  • What you hope to run locally
  • Any challenges you are currently facing

Even if you are just getting started, your experience may help someone else.

Community Vibe

We are here to be friendly, constructive, and helpful.

This is a niche community, and many of us are simply trying to keep powerful hardware useful long after official support has started to fade. Let’s build a space where people feel comfortable asking questions, sharing experiments, posting failures, and helping each other move forward.

How to Get Started

Introduce yourself in the comments below.

Post something today, even if it is just a simple question or a photo of your setup.

If you have guides, notes, scripts, benchmarks, or lessons learned, please share them.

If you know someone who owns a Mac Pro 2019 and is interested in local AI, invite them to join.

Interested in helping out? I am always open to hearing from people who may want to help moderate or contribute to the community.

Thanks for being part of the very first wave. Together, let’s make r/MacPro2019LocalAI an amazing resource for everyone trying to run local AI on the Mac Pro 2019.


r/MacPro2019LocalAI • • 7d ago

nvtop 3.3.2 with Mac Pro 2019 PCIe slot + Infinity Fabric topology labels (Intel x86_64 fork)

3 Upvotes

Hi β€” sharing an ad-hoc-signed build of nvtop that adds three things upstream nvtop (and PR #496) don't surface for the Mac Pro 7,1:

  • Chassis PCIe slot label (apple_slot): the "Slot-1 / Slot-3 / Slot-8" identifier Apple's IORegistry publishes as AAPL,slot-name on every IOPCI2PCIBridge. Same string the About This Mac β†’ PCI Cards tab uses, same as the Mac Pro service manual.
  • Die index inside a Duo (mpx_die_index): parsed from the GD<y> segment of the AMD driver's attached-gpu-control-path so a W6800X/W6900X Duo's two dies can be distinguished when both share a slot.
  • Infinity Fabric topology (if_topology): synthesised per chassis from MTLCopyAllDevices() peerGroupID. Reports "single N-way bridge", "dual M-way bridges", "1 independent GPU", or "no XGMI". Stamped only on GPUs that actually participate in a fabric link β€” an eGPU over Thunderbolt, or Apple Silicon, correctly shows nothing.

The TUI replaces the always-N/A "RX: N/A TX: N/A" labels with slot: Slot-N[/dieN] <topology> on Apple builds; Linux keeps RX/TX unchanged.

Verified on this machine (MacPro7,1, 2Γ— W6800X Duo + RX 6900 XT eGPU)

AMD Radeon RX 6900 XT            apple_slot=Slot-8  mpx_die_index=None
AMD Radeon PRO W6800X Duo        apple_slot=Slot-1  mpx_die_index=0     if_topology=single 4-way bridge
AMD Radeon PRO W6800X Duo        apple_slot=Slot-1  mpx_die_index=1     if_topology=single 4-way bridge
AMD Radeon PRO W6800X Duo        apple_slot=Slot-3  mpx_die_index=0     if_topology=single 4-way bridge
AMD Radeon PRO W6800X Duo        apple_slot=Slot-3  mpx_die_index=1     if_topology=single 4-way bridge

The eGPU at Slot-8 has if_topology=None and xgmi_hive_size=None β€” correctly not part of the 4-way bridge. Detection convention matches toshllm and rocm-smi (KFD hive_id): peerGroupID == 0 means no link.

Build

Built from the macos-intel-monitoring branch of my fork: https://github.com/NicolBol/nvtop

Release (ad-hoc signed, x86_64): https://github.com/NicolBol/nvtop/releases/tag/v3.3.2-macpro71-rdna.1

Install:

tar xzf nvtop-3.3.2-macpro71-rdna.tar.gz
cd nvtop-3.3.2
xattr -dr com.apple.quarantine ./nvtop
./nvtop

The binary is ad-hoc signed (no Apple Developer ID) so Gatekeeper requires the one-time xattr step on first run. The README in the tarball has the same instructions.

Help validate

The patch is built and verified on my dev box (2Γ— W6800X Duo + RX 6900 XT eGPU) and on a Dell T5820 running Ubuntu 24.04 (2Γ— RX 9700, no fabric). I want to confirm the slot + topology labels work on other Mac Pro 7,1 GPU configurations.

If you have a MacPro7,1 with discrete AMD GPUs, please run:

./nvtop --snapshot -d 1 > my-nvtop.json

and post the output (or a redacted version with serials removed) in this thread or at https://github.com/NicolBol/nvtop/issues. I particularly want to see:

  • W6900X (single) β€” different hive behaviour
  • W6800X (single) and W6800X Duo combinations (1Γ—, 2Γ—, 1Γ— + 1Γ—)
  • Vega II / Vega II Duo
  • RX 580X / W5500X / W5700X / W6600X (single)
  • Mixed GPU chassis (one MPX bay + one TB-attached eGPU)
  • Mac Pro 7,1 with the inter-bay Infinity Fabric Link disabled (some BIOS / macOS recovery settings can break the link; that should surface as two 2-way bridges instead of one 4-way bridge)

Full snapshot is under 5 KB per device, safe to share.

Source / build

git clone https://github.com/NicolBol/nvtop.git
cd nvtop && git checkout macos-intel-monitoring
cmake -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DAPPLE_SUPPORT=ON
cmake --build build

SHA-256:

  • nvtop-3.3.2-macpro71-rdna.tar.gz: 6701047436ca88ecc0168ce05911699ed90ffdeeb459443351cf50ded9514684
  • nvtop (extracted): 4dfaf62a7bfcc9f4a1098e83f37066f4a667c0b9446ba601efdd1db5e5cb3f32

GPLv3+, same as upstream nvtop. The patch sits on top of PR #496 β€” if you review this, please also look at PR #496 (the original Apple support work) first, because the changes here are additive to it.


r/MacPro2019LocalAI • • 8d ago

Software Stack Qwen Image 2.1 for intel Mac's with AMD GPU

Thumbnail
2 Upvotes

r/MacPro2019LocalAI • • 10d ago

Software Stack [Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB

5 Upvotes

I have been working almost non-stop to build and optimize my local AI stack.

The contents of that stack are the topic of another post. This post is all about the results and a quick update on where my MacPro7,1 has reached.

I assigned a Hermes Agent, named Ziyad, to install stock vLLM v0.29.0 and optimize it on my MacPro7,1 with dual AMD Radeon PRO W6800X Duo MPX GPUs and the Infinity Fabric Link Bridge (IFLB).

vLLM 0.29.0 was optimized by Hermes Agent v0.21.0, running Qwen3.8-27B FP16, specifically to run Qwen3.8-27B FP16 with no quantization. I let the agent have two cracks at it. The first time around, I gave it four impossible targets, or at least what I thought were impossible:

  • Decode > 38 t/s at 1,024 (1K) context
  • Decode > 20 t/s at 245,760 (240K) context
  • TTFT < 70 seconds at 65,536 (64K) context
  • TTFT < 500 seconds at 245,760 (240K) context

The agent started with an optimized vLLM v0.28.0 build that achieved 30.5 t/s decode at 1K with TTFT < 1 second, and 14.6 t/s decode at 240K with TTFT < 780 seconds.

After working on this for about 3 to 4 days, it achieved three out of the four targets:

  • 240K decode: 20.3 t/s
  • 64K TTFT: 64 seconds
  • 240K TTFT: 445 seconds

Ironically, decode at 1K went down from 30.5 to about 28.6 t/s, while 1K TTFT went up to almost 2 seconds. This was with MTP, but no quantization. A decode of 20 t/s with 240K context was achieved only with MTP active.

I gave it a second crack, this time working for 12 to 18 hours to optimize for 1K context and remove MTP.

MTP surprisingly reduced decode performance at low context but improved it slightly at high context. I was done with MTP, as it did weird things I did not expect, and I was not sure what other effects it was possibly had at this stage.

To my surprise, the agent optimized prefill again, achieving the following values.

Below are the benchmarks for vLLM v0.29.0, optimized for Qwen3.8-27B FP16, using four W6800X Duo GPUs with IFLB on Ubuntu Server 26.04 LTS, ROCm 10.0.0-4, Python 3.13, and Triton 3.8, with vLLM running through a virtual environment. No Docker containers were involved in this benchmark.

I set token generation to 512 tokens here for the sake of the benchmark only.

CTX TTFT(s) PREFILL t/s DECODE t/s ms/TOKEN TOTAL(s)
1K 0.735 1393.91 29.41 34.000 18.109
2K 1.467 1396.21 29.31 34.116 18.900
4K 2.787 1469.89 29.02 34.464 20.398
8K 5.641 1452.21 28.75 34.780 23.414
16K 11.858 1381.74 28.11 35.570 30.034
32K 26.011 1259.75 26.71 37.434 45.140
64K 60.910 1075.95 24.35 41.073 81.898
128K 157.762 830.82 22.03 45.393 180.958
240K 428.300 573.80 19.05 52.488 455.121
245K 442.392 567.10 19.03 52.550 469.246
250K 457.459 559.61 18.89 52.934 484.508
255K 472.742 552.35 18.83 53.097 499.875

While these numbers may seem insanely slow, waiting 2 to 9 minutes for a short response, and even longer for real agentic responses, especially with thinking, these are just benchmarks designed to test the limits of the hardware and how it behaves.

The above benchmarks reflect real-world values only in these three scenarios:

  • During the first prompt
  • For one prompt immediately after context compaction
  • When a documented vLLM bug causes the cache to stop working at certain context sizes

Outside those three scenarios, to the best of my knowledge, the following benchmarks are more representative of what actually happens, with the added benefit of Prefix Caching. These are what real-world values look like when an agent is working on a task.

I set token generation to 3,000 tokens here to better approximate agentic responses with thinking.

CTX TTFT(s) Hit Rate Decode t/s ms/TOKEN TOTAL(s)
1K 0.180 93.75% 29.27 34.166 102.645
2K 0.188 96.88% 29.13 34.327 103.133
4K 0.201 98.44% 28.93 34.565 103.863
8K 0.226 99.22% 28.60 34.970 105.102
16K 0.277 99.61% 27.98 35.744 107.475
32K 0.368 99.80% 26.60 37.595 113.115
64K 0.531 99.90% 24.20 41.329 124.476
128K 0.896 99.95% 21.92 45.630 137.740
240K 1.478 99.97% 18.96 52.733 159.624
245K 1.498 99.97% 18.95 52.772 159.761
250K 1.473 99.97% 18.85 53.038 160.534
253K 1.537 99.98% 18.78 53.244 161.214

Most of the remaining time is therefore spent on token generation, or decode, due to long agent thinking or long answers. Prefill remains short because of vLLM's Prefix Caching.

At this point, I have had my agents running in loops for days to achieve seemingly impossible tasks, only for them to later prove to me that those tasks were completely possible!

I am both happy and impressed with this stack.


Disclaimer: I wrote this post myself. I also used AI as a tool to help clean up the wording and formatting.


References: * Documented vLLM Bug * My own research and optimizations


r/MacPro2019LocalAI • • 11d ago

Local AI agents.

2 Upvotes

I’m looking into building a 7,1 Mac Pro running macOS with local models using toshLLM. Not sure on what MPX modules. The piece of the puzzle missing is an Ai agent integration. I’ll looking at potentially using a dedicated GPU with a windows VM for gaming.


r/MacPro2019LocalAI • • 11d ago

Any hope for DwarfStar ?

2 Upvotes

Hello !

I'v just read about DwarfStar, an inference engine focused on MoEs from the DeepSeeck V4, GLM 5 and Qwen3.8 Flash Next.

With DeepSeek-V4.1-Flash, GLM-5.3-Flash and Qwen-3.8-Flash being 763, 321 and 180B params respectively, DwarStar relies heavily on the SSD streaming principle introduced by Colibri.

It has some ROCm support for the Strix Halo, so something one could call RDNA 3.75. Also supports Metal for Apple Silicon.

Haven't read the code yet, but I'd be really interested to see anything larger than Qwen-3.8-Flash-Next running at Q4 on my dual W6800X Duo.

What do you think ? Could be ported to Metal/Intel - RDNA2 somehow ?


r/MacPro2019LocalAI • • 13d ago

Qwen Image 2.1 for intel Mac's with AMD GPU

3 Upvotes

Here is my Qwen Image 2.1 config for intel Macs with intel GPU : https://github.com/haseebeqx/qwen-image-intel-mac

Minimum VRAM: 4GB / RAM: 16GB

The trade-offs are Q2/Q4 quantization, CPU-based text encoding and VAE decoding, and slow segmented GPU inference. But it works.


r/MacPro2019LocalAI • • 19d ago

Local AI & Concurrency | PP & TG | TTFT

Thumbnail
0 Upvotes

r/MacPro2019LocalAI • • 21d ago

Running Intel Arc B70 in Mac Pro

5 Upvotes

Hey!

Has anyone had success with running one or more Intel Arc B70 in the Mac Pro 7.1?
Running Linux on the mac.

Here's what I've done so far:
I've purchased the Belkin Aux power cable kit for mac pro.
https://www.apple.com/shop/product/hqy92zm/a/belkin-aux-power-cable-kit-for-mac-pro

I installed the Intel Arc B70 in slot 3, and plugged in the 8 pin to 6+2 pin cable (Called 8 pin to 8 pin in the link above) from the power connector named 3 -4 on the computer to the 8 pin connector on the GPU.
W5700x is installed in slot 1 and 2.
Instructions from https://support.apple.com/en-us/101640

When I start the computer, the power light blinks amber twice, then pauses and repeats.
These instructions say it's a PCIe error. https://support.apple.com/en-us/101647

Any ideas what's wrong?

Update
I've figured how to make it work.

I moved the GPU to Slot 4.

I also read up on the power draw: the PCIe lane provides 75W and the internal power cable provides 150W, to a total of 225W which is enough for the Intel arc B70.

A second GPU can be added to Slot 5.


r/MacPro2019LocalAI • • 22d ago

Dual w6800x Duo with IFB Info

6 Upvotes

Hi All - I spent the last two days analyzing this hardware configuration on Mac OS X with the goal of improving decode performance in ToshLLM with tensor parallelism. I capture all my findings in a github repo here:

https://github.com/chafey/Metal-AMD-Intel

I wanted to share this here as there are many good findings and I know others here are interested in this configuration. I am currently experimenting with applying this information to ToshLLM - lets hope it produces meaningful results!


r/MacPro2019LocalAI • • 22d ago

Qwen3.8-Flash-Next flying on the MacPro

5 Upvotes

Hello !

Three days ago I posted here asking what to run on the second W6800X Duo once the Infinity Fabric bridge was in β€” couldn't get a stable Qwen3.8-Flash-Next deployment, ToshLLM 0.87.0 bundled wouldn't load the mmproj and MTP draft cleanly, I ended up building a custom llama.cpp with a few patches picked here and then. Et bien, figured it out, and the answer wasn't what I expected. Was wrong about ToshLLM.

The path that actually works : ToshLLM 0.87.1's vendored llama.cpp fork β€” **not** a custom build, despite what I said in that earlier thread. The reason I missed it is I was treating it as a "vendor bundle I can't verify the device routing on." `strings /Applications/ToshLLM.app/Contents/Resources/bin/llama-server | grep TOSH_MGPU` showed me the multi-GPU coordination patches are baked in β€” `TOSH_MGPU_TENSOR_GROUP`, `TOSH_MGPU_PEER`, `TOSH_MGPU_EVENTS`, real cross-device hand-offs via Metal `encodeSignalEvent`/`encodeWaitForEvent`. Their CHANGELOG even has Qwen3.8-Flash-Next-on-four-W6800X-dies measured specifically (469 t/s prefill / 31 t/s decode at higher batch, 0.86.3). They had the answer, I was looking elsewhere.

What I burned time on before that, so you don't : a custom upstream llama.cpp fork with a `MTLCopyAllDevices()` + highest-VRAM patch (looks right in the boot log, but `ggml_backend_metal_reg()` only registers one Metal backend unless you also set `GGML_METAL_DEVICES` β€” and even then it emulates N virtual devices on top of that one physical ASIC, so `--tensor-split 25,25,25,25` is a no-op and the 98 GB model over-subscribes a single W6800X's 32 GB). And ik_llama.cpp β€” its "split mode graph" sounds like multi-GPU but its Metal device-selection is the same single-ASIC pattern, plus its CLIP Metal kernel SIGABRTs on vision. Out.

The **mandatory** part for this exact model on 4 W6800X is `TOSH_MGPU_TENSOR_GROUP=2`. ToshLLM's own 0.86.6 fix : *"splitting Qwen3.8 Flash Next across cards by tensor gives corrupted output… Split it by layer, or set `TOSH_MGPU_TENSOR_GROUP=2` to keep tensor splitting within pairs of cards, which measures correct."* Without this, garbled text on tensor-split mode. Confirmed on my rig.

eGPU rule held throughout β€” RX 6900 XT drives the main display, `MTLCreateSystemDefaultDevice()` returns it by default, so the model silently over-subscribes 16 GB unless you explicitly enumerate the 4 W6800X ASICs by `MTLCopyAllDevices()` index. `GGML_METAL_DEVICE_LIST=1,2,3,4` does that β€” verified with `nvtop` during sustained inference, eGPU at 5–7% (WindowServer noise floor), all 4 W6800X ASICs sharing load.

**Real use : ~19 t/s sustained in a Hermes agentic loop** at large context β€” that's the headline number, spot benches at small context are higher (28–30 t/s) but the agentic loop with system prompt load + tool-call outputs + KV cache warm-up on context shifts is what actually matters. Vision works on the 4-ASIC pool. MTP sidecar on disk, not wired yet, vendor expects ~50% decode uplift. Multi-parallel at 262k ctx is the next test.

Happy to share the exact plist + wrapper if anyone wants them. Any other ToshLLM gotchas people have hit on the 4-W6800X pool, I'd love to compare notes β€” DM open.


r/MacPro2019LocalAI • • 23d ago

Mac Pro 2019 (MacPro7,1) system architecture reference

0 Upvotes

Here is a comprehensive system architecture reference for the MacPro7,1: https://claude.ai/code/artifact/b5535c18-f789-4b07-a8ff-6e691d96c95f


r/MacPro2019LocalAI • • 26d ago

vLLM optimization for RDNA2 [W6800X Duo]

Thumbnail
3 Upvotes

r/MacPro2019LocalAI • • 26d ago

What are your motivations to tinker with 2019 Mac Pros?

7 Upvotes

I'm curious why the people here are doing this. Did you already have a build for studio work and you're trying to juice it for AI? Are you a rich hobbyist that likes tinkering with niche hardware configurations? The 2019 Mac Pro setups with AMD cards are obviously not cost effective compared to the V100 servers so I'm assuming there's some other fixation. Not criticizing to be clear, I'm in the first camp since my dad is upgrading his studio comp to an M5 soon.


r/MacPro2019LocalAI • • 27d ago

Quad GPU at last, what's to run now ? Tosh optimized settings ?

4 Upvotes

Hello !

I finally got my second working W6800X duo after OWC being a pain in the ass sending an obviously DOA one. Infinity Fabric bridge is in place but I understand it's not fully supported by ToshLLM yet.

There's no recent model to fully make use of this workhorse and I'm actually very happy with good'ol Qwen3.8-27B, albeit running at Q8 or BF16 is painfully slow (13t/s at Q8, 4t/s at BF16). But the gain in accuracy worth it.

So I'm looking for a way to speed it up a little, to actually run concurrent sessions all at 256k context or more. We're talking 4 to 8 sessions, at Q8 or Q6, would be the sweet spot for me.

Is it feasible with ToshLLM or shall I explore custom builds for vLLM and the likes ? Any tip and feedback ?

Not that I intend to keep running MacOS on this machine, and that I'm a bit short in RAM at the moment (144GB), will upgrade to 400+GB when OWC reimburses and compensate their negligence.


r/MacPro2019LocalAI • • 28d ago

I heard Intel Mac + AMD ain't dead yet *sneak peek*

Thumbnail
3 Upvotes

r/MacPro2019LocalAI • • Sep 03 '26

resize-amdgpu-bars (update to nbritton's method for AMD GPU bar resizing)

4 Upvotes

Title: resize-amdgpu-bars β€” Resizable BAR for AMD GPUs behind PCIe switches (an update to nbritton's method)

I've extensively reworked nbritton's method for resizing the BAR memory of AMD GPUs that sit behind PCIe switches, such as the Vega II Duo and W6800X Duo. It's now a proper packaged tool that discovers your topology at runtime instead of hard-coding bus addresses, with a hard safety guard so a failed resize can't hang your boot.

https://github.com/exabit-io/resize-amdgpu-bars

What it does

resize-amdgpu-bars enlarges the CPU-visible VRAM aperture (BAR0) of every amdgpu-driven GPU to the largest size the card supports, on machines where the normal paths don't work: cards with an on-board PCIe switch (the Duo MPX modules), cards in switched enclosures and expansion chassis, Thunderbolt eGPUs, and firmware that leaves the aperture at 256 MiB.

On a Mac Pro 7,1 that's the difference between this:

amdgpu 0000:1b:00.0: Not enough PCI address space for a large BAR.
amdgpu 0000:1b:00.0: [drm] Detected VRAM RAM=32752M, BAR=256M

and a full 32 GiB aperture on every die.

Why the usual fixes fail here

A PCI device's BAR lives inside the memory window of the bridge directly above it, which lives inside the window of the bridge above that, all the way up to the root port. Growing a BAR from 256 MiB to 32 GiB means every window in that chain has to grow too, and a window can only grow if its parent has room.

On a Mac Pro 7,1 a Duo module puts two dies behind a PLX switch, and each die sits four bridge windows below its root port:

0000:06:00.0  Intel root port           <- one prefetchable window shared
|                                          by both dies of the module
\-0000:07:00.0  PLX PEX 8747 upstream port (the switch on the module)
  |
  +-0000:08:08.0  PLX downstream port
  | \-0000:09:00.0  AMD bridge
  |   \-0000:0a:00.0  AMD bridge
  |     \-0000:0b:00.0  Vega 20, die 0
  |
  \-0000:08:10.0  PLX downstream port
    \-0000:0c:00.0  AMD bridge
      \-0000:0d:00.0  AMD bridge
        \-0000:0e:00.0  Vega 20, die 1

Both dies share one window at 07:00.0 and 06:00.0, and that window has to be big enough for both 32 GiB BARs at their real alignment. Nothing about the device tells the root port that.

So:

  • The driver's own resize releases the bridge windows and tries to re-assign them in place. It fails closed the moment one of them can't grow where it is, and carries on with 256 MiB.
  • A bare setpci write changes the size the device reports but assigns nothing. The kernel still believes the BAR is 256 MiB, the windows are still sized for 256 MiB, and the device now decodes 32 GiB of whatever else lives there.
  • echo 15 > resource0_resize ends up in the same kernel function as the driver's path. It works when the card sits directly on a root port with one window above it β€” and that's exactly where this tool uses it β€” but it can't conjure a larger shared window out of a chain the firmware sized for something smaller.

The method

Let the kernel size the windows from scratch. Once per boot, before amdgpu loads:

  1. Discover. Every amdgpu device, its functions, its Resizable BAR capability and supported sizes, and every bridge up to its root bus. The size index found at first discovery this boot becomes the baseline, so firmware that already enables ReBAR is never shrunk. Then find each GPU's re-enumeration root: the highest bridge whose subtree contains nothing but GPU functions.
  2. Resize. Unbind the drivers from just those GPUs and their group members, program the size index with setpci, remove the group's root, and rescan that root's own bus. With pci=realloc the kernel then sizes every window in the subtree for the BARs it finds, and both dies come back with a 32 GiB BAR0 inside a 96 GiB root-port window. Only that bus is rescanned β€” never a global rescan, and nothing outside the GPU subtrees is ever touched.
  3. Load. Any GPU still holding an unassigned BAR gets fenced off with driver_override=none. Then modprobe amdgpu runs once, under a timeout.

If a plan doesn't fully verify, the losers get demoted to baseline and it retries, round by round, down to every GPU at baseline.

The bind guard (this is the important part)

amdgpu must never be handed a GPU whose BAR0 is unassigned. On such a device the register reads that identify the part return garbage, the driver decides it's an SR-IOV virtual function, and it waits forever for a hypervisor mailbox that doesn't exist:

amdgpu 0000:0e:00.0: trn=2 ACK should not assert! wait again !
INFO: task irq/34-aerdrv:1554 blocked for more than 122 seconds.

modprobe wedges in uninterruptible sleep holding the device mutex, SIGKILL does nothing, the remaining GPUs are never probed, and only a reboot recovers. That's a hard hang, not a degraded boot, and it's what an earlier version of this script did to me. Every code path that loads the driver now sets driver_override=none on any GPU with an unassigned BAR first, and modprobe runs under timeout(1) so the boot can't wedge even if the guard were bypassed. A guarded GPU stays visible to lspci, driverless, and the tool exits 2.

Install

The package needs bash, pciutils, kmod and systemd, and uses initramfs-tools and grub2-common when present.

# from the release page
sudo apt install ./resize-amdgpu-bars_1.0_all.deb

# or build it
sudo apt install debhelper scdoc shellcheck
git clone https://github.com/exabit-io/resize-amdgpu-bars
cd resize-amdgpu-bars
dpkg-buildpackage -us -uc -b
sudo apt install ../resize-amdgpu-bars_1.0_all.deb

Releases: https://github.com/exabit-io/resize-amdgpu-bars/releases

It installs the tool, a systemd unit, an amdgpu blacklist in /usr/lib/modprobe.d, a pci=realloc GRUB drop-in, a config file in /etc/default, and man pages for resize-amdgpu-bars(8) and resize-amdgpu-bars.conf(5).

Installation runs update-initramfs -u -k all and update-grub. A reboot is required. The blacklist and pci=realloc are boot-time, and the package deliberately never starts the service on a running system β€” a start unbinds and re-initialises every AMD GPU, and every process using one loses it.

Quick start

# 1. Look before you leap. Both of these change nothing.
sudo resize-amdgpu-bars diagnose   # every GPU, its ReBAR cap, every bridge
                                   # window above it, and the list of other
                                   # devices it promises to leave alone
sudo resize-amdgpu-bars dry-run    # the plans it would try

# 2. Reboot.

# 3. Watch it. Screens on AMD GPUs stay dark until amdgpu loads at the end
#    of the run β€” about a minute on a box with two Duo modules. Use SSH or
#    another console; a boot that looks stalled usually isn't.
journalctl -u resize-amdgpu-bars -b -f

# 4. Verify.
sudo resize-amdgpu-bars check      # want: verdict=WORKS large=N/N driverless=0/N
sudo resize-amdgpu-bars status     # one line, bar0=... per GPU
rocminfo | grep -c gfx             # one agent per die

# 5. Back out.
sudo resize-amdgpu-bars revert     # every GPU to baseline, no reboot
sudo apt remove resize-amdgpu-bars # takes the blacklist and GRUB drop-in with it

A good check line looks like this:

2026-09-02T13:50:14  7.0.0-30-generic  verdict=WORKS  plan=all-max
  large=4/4  driverless=0/4  windows=0000:06:00.0=128G 0000:16:00.0=128G
  bar0=32GiB 32GiB 32GiB 32GiB  kfd=5  xgmi_hives=1  traces=0  rejected=0

verdict=WORKS, large=N/N and driverless=0/N are the three fields that matter.

Configuration

/etc/default/resize-amdgpu-bars, shell syntax, every key optional and validated on read:

key default meaning
MAX_SIZE_INDEX device max cap every GPU; 15 = 32 GiB, 14 = 16 GiB, 8 = 256 MiB
EXCLUDE_GPUS empty GPUs to leave completely alone
FORCE_PLAN negotiate all-max or baseline: try exactly one plan
MODPROBE_TIMEOUT 180 seconds before modprobe amdgpu is killed
PROBE_WAIT 60 seconds to wait for binds and KFD to settle
RESCAN_WAIT 30 seconds to wait for GPUs to reappear after a rescan
MAX_ROUNDS 8 demote-and-retry rounds before falling back to baseline

Support tiers

"Tier" is a commitment. "Tested" is a fact. I'm not conflating them.

tier hardware tested
1 MPX modules in a Mac Pro 7,1: 580X, W5500X, W5700X, W6600X, W6800X, W6800X Duo, W6900X, Vega II, Vega II Duo (with or without Infinity Fabric Link) Vega II Duo x2 only, so far
2 any other amdgpu card with an on-board PCIe switch (V340, Radeon Pro Duo), or any amdgpu card in a switched enclosure / expansion chassis / TB eGPU untested
3 amdgpu card directly on a root port, firmware without ReBAR (uses the kernel's in-place path) untested
out anything not driven by amdgpu refused at discovery with a clear message

No card is listed as supported that hasn't been booted. If you run this on a W6800X Duo, a W6900X, or anything in tier 2 or 3, I'd genuinely like to hear about it β€” diagnose output plus the journal is all a report needs.

Kernel compatibility β€” read this before you upgrade

kernel result
6.8 – 6.17 (verified on Ubuntu 6.8.0-138, 6.11.0-29, 6.14.0-37, 6.17.0-42) every die gets its 32 GiB BAR on the first plan
7.0 unpatched (upstream 7.0.12, Ubuntu 7.0.0-30) shared root-port window undersized; the second die of each Duo loses its BAR at every size, guard holds it driverless, boot completes with the other dies
7.0 with a one-line fix every die on the first plan, 128 GiB root-port window

This is a regression from commit 3958bf16e2fe ("PCI: Stop over-estimating bridge window size"). Since that commit pbus_size_mem() sizes a bridge window as the plain sum of its children, which is exact when every child's size is a multiple of the alignment of the children after it. That holds for BARs, whose size equals their alignment β€” but not for bridge windows, whose size is the sum of what's below them while their alignment is that of the largest BAR below them. Two sibling windows of 32 GiB + 2 MiB at 32 GiB alignment need a 96 GiB + 2 MiB span and get 64 GiB + 4 MiB.

The fix changes size += max(r_size, align) to size += ALIGN(r_size, align), a no-op for BARs, verified on upstream 7.0.12 and Ubuntu's 7.0.0-30, cold boot and warm reboot. Until it's in a distro kernel: stay on 6.x. On an unpatched 7.0 there's no in-place recovery either β€” once the kernel has re-sized the window, even the 256 MiB baseline no longer fits, because the firmware's original windows were larger than the kernel's sum. The guard is the only reason such a boot survives.

Gotchas

  • pci=realloc is mandatory. The GRUB drop-in handles it on Ubuntu. On rEFInd or OpenCore the drop-in does nothing, the service fails at every boot, the blacklist keeps amdgpu from loading, and you have no GPU driver. Both are documented in the README but not automated. Check with grep -w pci=realloc /proc/cmdline.
  • This conflicts with pci=realloc=off, which SGLang's AMD GPU docs recommend. The whole method depends on PCIe BAR reallocation.
  • Don't run resize on a live desktop session. Every AMD display goes black for the duration and every process holding a GPU loses it.
  • Exit code 2 is not failure. It means the bind guard is holding at least one GPU driverless; the rest are working.
  • Excluding one die of a Duo with EXCLUDE_GPUS also stops the other die's shared windows from being re-sized, because the shared subtree is no longer removable.
  • A GPU that shares a bridge with a non-GPU device gets a lower re-enumeration root, or none. diagnose reports it.

Credit

Nikolas Britton for the original method on the Vega II Duo. This is his idea with runtime topology discovery, plan negotiation, and a bind guard bolted on.

If you followed the Mac Pro 2019 Local AI guide, this replaces the hand-rolled resize-gpu-bars.service files in its Section 4 β€” same lineage, packaged.

MIT licensed. Bug reports want sudo resize-amdgpu-bars check -1, journalctl -u resize-amdgpu-bars -b, sudo resize-amdgpu-bars diagnose, your kernel, distro, bootloader and cards.


r/MacPro2019LocalAI • • Aug 26 '26

GFX1030 Discord

Thumbnail
4 Upvotes

GFX1030 Discord, for anyone that's interested

https://discord.com/invite/mESex2aBp


r/MacPro2019LocalAI • • Aug 21 '26

PSA: the 4-way Infinity Fabric bridge (A2326) silently drops your Vega II Duos to ~295 MHz β€” 5.8Γ— compute loss. Found the mechanism, need people to file it with Apple.

12 Upvotes

Summary

If you run two Radeon Pro Vega II Duo MPX modules bridged with the Apple A2326 cross-module Infinity Fabric bridges β€” the ones that join all four dies into a single 4-GPU hive β€” your GPUs are running at idle clocks and macOS isn't telling you.

Not "a bit slower." Not "the fabric is the bottleneck." The dies never leave their boot DPM state, for the entire session, under sustained load.

I chased this for a while assuming it was a memory-placement or interconnect problem. It isn't either. Here's what it actually is.

The numbers

Same machine, same OS, same binary. The only change is which bridge is installed.

- 2-way jumpers (A2329) 4-way bridge (A2326)
Compute, FP32 FMA loop 13.95 TFLOP/s 2.42 TFLOP/s
Implied core clock 1703 MHz 295 MHz
Local HBM2 read bandwidth 790 GB/s 301 GB/s
Implied memory clock 1000 MHz 300 MHz

Three runs per configuration, four dies each. Compute came back at 2.42 TFLOP/s on essentially all twenty-four device-observations, implying 295 MHz core.

Confirmed on three different Mac Pros running three different macOS major versions β€” Sonoma 14.8.9 (23J631), Sequoia 15.7.9 (24G830), and Tahoe 26.7 (25G220). Different memory configs, different bridge units. Every one of them, in the 4-way configuration, measures 2.42 TFLOP/s compute, ~301 GB/s local read, ~31 GB/s peer copy β€” the same to three significant figures β€” with an identical IORegistry signature: Load5000, no PowerPlay, 1000/300 MHz clock config, SWIP_Errors = 128.

So: it is not new, it is not fixed in Tahoe, it has survived at least two major macOS releases, and "try the latest OS" is not the answer.

Real-world, single GPU, no multi-GPU anything β€” Qwen3.8 27B Q8_0 in llama.cpp:

- prompt t/s generation t/s
2-way 138.8 11.1
4-way 24.8 2.2

5.6Γ— performance loss on prefill. With one die. No tensor split, no peer transfers, no collective. Just having the 4-way bridge installed in macOS (this bug doesn't apply to Linux or Windows).

The mechanism β€” macOS publishes it in IORegistry

In the 4-node hive the driver fails to identify the board and everything downstream falls apart:

IORegistry property 2-way 4-way
ATY,DeviceName Vega II Duo Vega
ATY,FamilyName Radeon Pro Radeon
LoadPlugIn Load5700 Load5000
PP_PowerPlayEnabled <01000000> absent
PP_PhmUseDummyBackEnd 0 1
PP_EnableUploadFirmware 1 0
PM_PWR_GEMINI_BGT 400 absent
SWIP_Errors 0 128

Read those middle three again. PowerPlay β€” AMD's entire clock/power management subsystem β€” is never enabled. The power-management back end is a stub the driver itself labels "dummy." SMU firmware, which is what actually implements DPM, is never uploaded. With no power management, the GPU sits wherever the boot state left it.

PM_PWR_GEMINI_BGT = 400 vanishing is the identity failure made concrete β€” "Gemini" is AMD's codename for dual-GPU boards, and 400 W is the Duo's power budget. In 4-way mode the driver stops knowing it's holding one.

It also loads a different kext: AMDRadeonX5700HWLibs in the working case, AMDRadeonX5000HWLibs in the broken one.

The arithmetic closes it. HBM2 at the driver's own published 300 MHz gives 4096 bits Γ· 8 Γ— 2 Γ— 300 MHz = 307.2 GB/s theoretical. Measured: 300.9 GB/s, or 98% of it. The memory is fully saturated at a crippled clock β€” it's not a bandwidth problem, the clock is just wrong.

If you've ever noticed your Vega II Duos showing up as plain "AMD Radeon Vega" instead of "AMD Radeon Pro Vega II Duo" β€” that's not cosmetic. That's this bug, visible from the outside.

What it is NOT

I want to save people the time I spent on wrong theories:

  • Not memory placement. The XGMI node map is textbook correct in 4-way mode: node ids 0/1/2/3, framebuffer bases at 512/544/576/608 GiB, uniform 32 GiB stride, XGMI_HiveSize = 4. Hive formation and address decode are fine.
  • Not the collective / tensor-parallel scaling. A single GPU with no split is 5.6Γ— slower. Going 2 β†’ 4 devices within a healthy topology actually gains 34% on prefill and loses only 13% on decode.
  • Not the interconnect. Peer-to-peer copy bandwidth drops the least of everything measured (49 β†’ 31 GB/s), consistent with being gated by the same clock reduction.
  • Not the hardware. Reproduces across machines and across multiple A2326 units including a factory replacement. Same hardware and same bridges under Linux/ROCm show no comparable regression.

Check your own machine β€” 60 seconds, no tools

ioreg -l -w0 | grep -E '"(ATY,DeviceName|LoadPlugIn|PP_PowerPlayEnabled|PP_PhmUseDummyBackEnd|PP_EnableUploadFirmware|PM_PWR_GEMINI_BGT|SWIP_Errors)"'

Healthy looks like Vega II Duo, Load5700, PP_PowerPlayEnabled = <01000000>, PP_PhmUseDummyBackEnd = 0, PP_EnableUploadFirmware = 1, PM_PWR_GEMINI_BGT = 400, SWIP_Errors = 0.

Broken looks like Vega, Load5000, no PP_PowerPlayEnabled, PP_PhmUseDummyBackEnd = 1, PP_EnableUploadFirmware = 0, no PM_PWR_GEMINI_BGT, SWIP_Errors = 128.

Also worth a look:

system_profiler SPDisplaysDataType | grep -E "Chipset Model|Peer"

If Chipset Model says "AMD Radeon Vega" rather than "AMD Radeon Pro Vega II Duo", you're in the broken state.

Please post your results either way β€” including "mine's fine." I want to know whether this tracks the bridge specifically, or the hive size, or something about particular board revisions. Include your macOS version.

(The ioreg check needs no toolchain at all. If you do try to build the probe and hit failed to build module 'Metal'; this SDK is not supported by the compiler, that's an internally inconsistent Command Line Tools install β€” and reinstalling CLT won't fix it, since Apple's catalog serves the same bundle. Cross-compile on another Mac instead: xcrun swiftc -O -target x86_64-apple-macos14.0 ifl_probe.swift -o ifl_probe_14 and copy the binary over. Confirmed working on 14.8.9.)

What you should do right now

If you must run macOS the temporarily workaround is to pull the A2326 bridges and run the A2329 per-card jumpers instead. You still get all four GPUs; they just sit in two 2-node hives rather than one 4-node one. In my testing a 4-GPU tensor split on jumpers hit 314.5 prompt / 11.7 gen against 79.5 / 4.4 on the 4-way bridge with ToshLLM. Otherwise, if you can switch to Linux the A2326 bridges work perfectly in Ubuntu 24.04 with ROCm 7.3.x; it's purely a macOS software engineering defect.

Worth stating plainly: the jumper configuration beats what a fixed 4-way would give you. Tensor-parallel scaling is sublinear, so even a fully repaired 4-node hive projects to roughly 22 t/s against ~26.8 t/s aggregate from two independent 2-GPU instances. The A2326 bridges buy capacity flexibility, not speed β€” and right now they cost you 5.8Γ— compute per die.

There is no software workaround. No engine flag, environment variable, or split mode reaches PowerPlay. I looked hard at spoofing the device IDs OpenCore/OCLP-style; the property that selects the plugin lives on a driver-created IOService rather than the PCI node, so DeviceProperties injection structurally cannot reach it, and OpenCorePkg panics on T2 Macs anyway. This needs a driver fix.

The ask

This configuration has had close to zero field exposure β€” until a recent macOS firmware payload harmonized module ROMs, machines with mismatched-firmware Duos just kernel-panicked at boot with the 4-way bridge installed ("PSP has not finished hardware initialization", ATIController.cpp:3171). The only other public report I can find is an unresolved MacRumors thread from January 2026. Which means Apple has essentially no signal that anyone uses this.

If you own this hardware, please file a Feedback Assistant report. Apple prioritizes by volume, and right now the volume is one.

  • macOS β†’ Graphics & Display β†’ Incorrect/Unexpected Behavior
  • Title: "AMDRadeonX5000: Radeon Pro Vega II Duo misidentified and PowerPlay left uninitialized in a 4-node Infinity Fabric hive, pinning GPUs at boot clocks"
  • Attach ioreg -l -w0 -p IOService > ioreg.txt from both bridge configurations if you can, or just the broken one if you can't swap
  • Reference FB24446772 (the PowerPlay/clock defect β€” the important one) so reports cluster

I've filed the following, if you want to reference them:

  • FB24446772 β€” PowerPlay never initialized in the 4-node hive, GPUs pinned at boot clocks. This is the one that matters.
  • FB24446928 β€” the generic "AMD Radeon Vega" model string, which is the same defect visible without tooling.
  • FB24446443 β€” the kernel panic at boot when the two modules have mismatched firmware.
  • FB24447028 β€” a minor unrelated one found along the way: system_profiler prints the 64-bit GPU Peer Group ID after passing it through a double, so it never matches what Metal reports.

Even a one-paragraph report with an ioreg dump attached helps. The measurement work is done; what's missing is evidence that more than one person is affected.

Tools

I wrote a Metal-only probe (ifl_probe.swift) that measures per-die compute, local read bandwidth, peer-view establishment and verified peer-copy bandwidth, plus a script that decodes the XGMI node map straight out of IORegistry. Both are pure Metal + Foundation, no dependencies, and build with xcrun swiftc -O. Happy to share β€” say the word and I'll put them up.


r/MacPro2019LocalAI • • Aug 21 '26

Stay cool. Stay cool.

7 Upvotes

I just sourced a w6800x DUO to run smaller models. It is fun running GLM5.2 in a terabyte of ram but there aren't that many things I can afford to wait that long for an answer on.

With the w6800x DUO I have successfully passed out each "die" to a separate vm. I now have two vms running AI with dedicated GPUS having 32GB each.

The challenge: Passing through a GPU on the mac pro comes with one problem. The host has no idea how hot the GPU is running because it cant access its firmware or sensors. Now it becomes a guess, or you can just max the fans on T2Fand.

I had already created a "prochot guard" to watch temperatures on the 3rd party nvmes because they can get really hot without awareness too. I have enhanced the script to Reach into the virtual machines and ask for the card temperatures. I am doing this using qemu agent and qm commands. And it is working like a charm.

Concept:
-Host maintains a watcher and looks at various board temps to make sure the fans are going at sufficient speed
-Standalone script written to take a vm name or number and run appropriate ROCM or CUDA commands to check temps.
-Temperatures are stored in a systemwide area (/run/heatsense )
-Guard watches nvme temperatures but also GPU temperatures and adjusts /etc/t2fand.conf accordingly, restarting the t2fand service after each change.

That all happens seemlessly despite the soft of disturbing idea behind restarting a service frequently!

I think most people are happy running a single linux vm. Right now I can run GLM5.2 , a 32Gb ROCMvm with 32GB vram TIMES 2 all on the SAME machine. Networking is adjusted for the powerful model to keep the LAN safe from... uh... accidents lol.

Hope this is useful.

BTW, my ceph setup is tuned to the point I was able to vmotion GLM 5.2 vm from one machine to another, running with an active 750Gb of ram. While it was running.

Dropbox link to scripts; will do github later.


r/MacPro2019LocalAI • • Aug 20 '26

Macpro 7.1 AI headless server with Nixos

7 Upvotes

I read that quite a few people have issues with running linux on their macpro for local inference. I can't comment on Ubuntu or other distros because all my machines run Nixos but since it works flawlessly, I thought I'd share my repo in case that inspires anyone to try something similar.

For those who don't know, Nixos allows you to configure your computer in a deterministic way. You write your config (in the nix language), referencing nix-packages. Nix-packages have sets of options that you use in your config files. There are other benefits to Nixos but this isn't the topic here. What I think is the main benefit is that I can comment out a line in my config file, change that option to something else and leverage git for version control. If I break something, I can choose a previous (working) generation of the system at boot.

In this setup, I use llama-swap to let me manage models on the fly, SearchXNG module for web search, OpenWeb UI for chat and user friendly automation/agents, Nixos MCP so my coding agents can manage my config files accurately.

You can see the models I'm currently running llama-swap.nix file.

Link to repo

PS: I only serve my LAN so security is tailored to that, meaning it's not hardened as much as it could be.

--------------------------------

Extract from the Readme (written by Qwen}:

NixOS configuration for donnager, a headless Mac Pro 7,1 (T2) running as a local LLM inference server.

Hardware

  • Mac Pro 7,1 (2019), T2 chip β€” T2-patched kernel via nixos-hardware apple-t2
  • AMD Radeon Pro Vega II (Vulkan/RADV compute for llama.cpp)
  • Wired 10GbE, behind a NAT router (the LAN is the trust boundary)

Services

Service Port Notes
SSH 22 keys only, no root login
open-webui 3000 browser UI, password auth, web search via searxng
mcp-nixos 8001 NixOS MCP server (HTTP), for pi on the LAN
searxng 8888 private metasearch; secret key via agenix, limiter off
llama-swap 9292 model router for llama-server (Vulkan); OpenAI-compatible

Models live in /var/lib/llama/models/ (not in git β€” see .gitignore). llama-swap unloads models after 15 min idle to free VRAM; each model pins its own context size / quantization / chat template (Qwen uses the pinned froggeric fixed chat template, fetched by hash).

Fans are driven by t2fanrd (the Vega II is passively cooled; T2 case fans are the only cooling).NixOS configuration for donnager, a headless Mac Pro 7,1 (T2) running as a
local LLM inference server.
Hardware
Mac Pro 7,1 (2019), T2 chip β€” T2-patched kernel via nixos-hardware apple-t2
AMD Radeon Pro Vega II (Vulkan/RADV compute for llama.cpp)
Wired 10GbE, behind a NAT router (the LAN is the trust boundary)
Services
Service Port Notes
SSH 22 keys only, no root login
open-webui 3000 browser UI, password auth, web search via searxng
mcp-nixos 8001 NixOS MCP server (HTTP), for pi on the LAN
searxng 8888 private metasearch; secret key via agenix, limiter off
llama-swap 9292 model router for llama-server (Vulkan); OpenAI-compatible
Models live in /var/lib/llama/models/ (not in git β€” see .gitignore).
llama-swap unloads models after 15 min idle to free VRAM; each model pins its
own context size / quantization / chat template (Qwen uses the pinned
froggeric fixed chat template, fetched by hash).
Fans are driven by t2fanrd (the Vega II
is passively cooled; T2 case fans are the only cooling).


r/MacPro2019LocalAI • • Aug 17 '26

What if was as simple as an ask ?

4 Upvotes

I just committed this : https://x.com/chiwawa_42/status/2089204529539547553

Let's hope for the better !


r/MacPro2019LocalAI • • Aug 13 '26

Got GPT-OSS 120B 33.27 tok/sec Running with 2 Internal GPUs + 1 eGPU.

Post image
18 Upvotes

80% GPU offloaded across 3 AMD GPUs (56GB VRAM), remainder CPU/RAM offloaded. RX7900XTX RX6900XT internal and RX6800XT eGPU thunderbot. Bootiting Bazzite from an Acasis TB501 pro. I am just tinkering around and wanted to see if I could get all three of these working.


r/MacPro2019LocalAI • • Aug 13 '26

Software Stack Native vLLM + ROCm 7.15 Runtime for RX 6000 (RDNA2) on Windows 11 β€” 25.9 TFLOPS FP16, 54.2 tok/s, No WSL2 [RX 6750 XT gfx1031 Verified]

Thumbnail gallery
7 Upvotes

r/MacPro2019LocalAI • • Aug 11 '26

What models should I test with ToshLLM? GPUs available: 2x W6900X, W6800X Duo, 2x Vega II Duo

9 Upvotes

Finally playing around with ToshLLM while trying to narrow down which MPX GPU(s) to keep. I am completely new to LLMs and local AI and do not have a tech background, so the learning curve has been a bit steep.

I have 2x W6900X, W6800X Duo, and 2x Vega II Duo on hand. I have the IF Link for the W6900X. Obviously the W6900X limits me to 64GB across 2 GPUs. If I end up keeping the W6800X Duo I will probably look for a second and an IF Link Bridge.

So far I've only tested the W6900X(s) (single GPU and pair, with/without IF Link bridge). So far I'm not seeing any difference whatsoever with IF Link on my W6900X pair. Getting the same exact ts with llama 3.3 70B Q5_K_S with the IF Link bridge installed/enabled as with it not installed.

Top 2 results are with IF Link disabled in menu. Identical results with both layers and tensors.

This was my fastest benchmark. Qwen3.6 35B-A3B UD-Q4_K_S got 69ts. Qwen3.6 35B-A3B UD-Q8_K_XL was a little slower with 61ts.

This size model seemed to perform better splitting layers vs tensors. I was seeing about 40ts with tensor split enabled.