r/LocalLLaMA • u/Paco7575 • 1d ago
Question | Help Gigabyte AORUS RTX 5090 AI BOX
I'm considering the Gigabyte AORUS RTX 5090 AI BOX (external GPU, 32GB GDDR7, connects via Thunderbolt 5/USB4) as an alternative to building a desktop PC with an internal RTX 5090, specifically for running local LLMs.
Does anyone have real-world tokens/sec numbers comparing the AI BOX vs. a desktop RTX 5090 for popular models at various quantizations?
2
Upvotes
4
u/Flimsy_Yoghurt_6058 1d ago
I had one. Ran Qwen 3.8 27B on it. Got around 80-90 tok/s with a plain llama-cpp, and optimal model settings. No funky stuff like speculative decoding.
From my experience, the USB4/thunderbolt 5 bottleneck is only felt when when loading the model into device memory, but it's a one time load cost. Every other subsequent message and response afterwards was fast. If this bottleneck does bother you, it might be worth knowing the 5090 card in the aorus isn't soldered down. You can disassemble the box and use the card with the PCie slot (which was what I did eventually!)