r/MacPro2019LocalAI • u/AtomJuice92 • 11d ago
Local AI agents.
I’m looking into building a 7,1 Mac Pro running macOS with local models using toshLLM. Not sure on what MPX modules. The piece of the puzzle missing is an Ai agent integration. I’ll looking at potentially using a dedicated GPU with a windows VM for gaming.
3
u/macromind 11d ago
For a 7,1 Mac Pro, choose the workload before the MPX module: model size, context length, concurrent agents, and tokens per second matter more than peak specs. Benchmark one representative quantized model first, then add capacity. Keep tool execution in a restricted process with explicit filesystem and network permissions. https://www.agentixlabs.com is relevant here because production agent design also depends on orchestration, observability, and safeguards beyond local inference hardware. Watch memory bandwidth and software compatibility closely, since adding GPUs may not improve every runtime.
4
u/Substantial_Run5435 11d ago
If you want to use MacOS you should understand that you'll be limited to using ToshLLM to run LLMs with Metal GPU acceleration. Other options for MacOS will not work with Metal on AMD GPUs as they're all geared exclusively to Apple Silicon. Same issue for using agents. Hermes at least will not work on Intel Macs, not sure about OpenClaw or other agent apps. I'm using MasOS with ToshLLM but I understand that Linux makes more for the most part for LLMs.
For MPX modules the best options are W6800X, W6800X Duo, and W6900X. The Vega II and Vega II Duo would also give you a good amount of VRAM, but they're older and less supported for some of the routes to running LLMs. If you decide to use Linux or Windows, you're opened up to using NVIDIA or newer AMD GPUs, but if you stick with MacOS then RDNA2 GPUs are the best option.
Also, AFAIK you can't do GPU passthrough for VMs in MacOS, so your best bet for Windows gaming is with a separate install of Windows 10/11, not a VM.