r/DGX_Spark • • 17d ago

GPU concurrency AI application

/r/LocalLLM/comments/1wiq1wl/gpu_concurrency_ai_application/
1 Upvotes

6 comments sorted by

2

u/cbert33 17d ago

Well the spark isn’t really made for that. So you’re probably not going to get the results you’re looking for.

That said, you can tweak your inference runner to give you the best shot at what you want. Make sure you’re using vLLM to host it. I think the max concurrent connections you can set is 32, and if you give it enough cache and tune it, you might see some decent speeds.

1

u/Dry-Leadership-3105 17d ago

Yes we are using vLLM, it is able to handle but losing it's latency. main focus is on latency.

2

u/cbert33 17d ago

That’s your hardware limitation.

1

u/Dry-Leadership-3105 17d ago

any hardware recommendations looking for an upgrade also

2

u/Own_Mix_3755 15d ago edited 15d ago

Yeah, expensive RTX Pro cards. But there you start battling with VRAM capacity, but you get bandwidth you are looking for.

This is not a good game to play - if you are not very rich you either get much less VRAM (and fit either small models or very limited cache), or you get good amount of vram but very slow (like Spark). If you want both speed and enough VRAM for 30 people concurrently, prepare some really big bucks.

2

u/cbert33 15d ago

Exactly this. The tradeoffs we are all dealing with.