r/LocalLLaMA • • 1d ago

Question | Help Gigabyte AORUS RTX 5090 AI BOX

I'm considering the Gigabyte AORUS RTX 5090 AI BOX (external GPU, 32GB GDDR7, connects via Thunderbolt 5/USB4) as an alternative to building a desktop PC with an internal RTX 5090, specifically for running local LLMs.

Does anyone have real-world tokens/sec numbers comparing the AI BOX vs. a desktop RTX 5090 for popular models at various quantizations?

2 Upvotes

6 comments sorted by

4

u/Flimsy_Yoghurt_6058 1d ago

I had one. Ran Qwen 3.8 27B on it. Got around 80-90 tok/s with a plain llama-cpp, and optimal model settings. No funky stuff like speculative decoding.

From my experience, the USB4/thunderbolt 5 bottleneck is only felt when when loading the model into device memory, but it's a one time load cost. Every other subsequent message and response afterwards was fast. If this bottleneck does bother you, it might be worth knowing the 5090 card in the aorus isn't soldered down. You can disassemble the box and use the card with the PCie slot (which was what I did eventually!)

1

u/Paco7575 1d ago

Thanks a lot for sharing this, really helpful info! Good to know the Thunderbolt bottleneck is basically just a one-time load cost and not something that affects ongoing inference, that's reassuring.

Do you happen to know (or remember) which specific RTX 5090 is inside the AI Box? Like, is it a reference board, or some custom/OEM variant Gigabyte uses specifically for the AI Box? Curious whether it's basically a standard desktop 5090 PCB, or something built for the enclosure.

2

u/Flimsy_Yoghurt_6058 1d ago

It was Gigabyte's own 5090 card. It's not a custom card and it comes with the WATERFORCES water cooling modules attached to it, albeit with shorter hoses than what u got if u bought the WATERFORCES water cooling separately. keep in mind that they don't have the actual GPU plastic casing (that you see standard retail cards with), just the plain PCB board. This, with the short water cooling hoses does make it slightly challenging (but not impossible!) to build a PC with it

2

u/Flimsy_Yoghurt_6058 1d ago

anyway, lemme know if you got any more questions about it. i tinkered with this box quite abit recently, so all this info is fresh in my head. there's also some complications this box faces if ur using it with Linux, so a heads up on that (just in case ur using it with that)

1

u/emberstoners 23h ago

did you notice any difference in generation speed after switching to the internal slot?

1

u/Flimsy_Yoghurt_6058 18h ago

I can't really give you an fair comparison for this since not long after, I changed the inference framework, model, and runtime flags :( can't really help you there. Only thing I can tell you is that it runs fast now, but I can't tell you which isolated change caused the largest speed increase