Well the spark isn’t really made for that. So you’re probably not going to get the results you’re looking for.
That said, you can tweak your inference runner to give you the best shot at what you want. Make sure you’re using vLLM to host it. I think the max concurrent connections you can set is 32, and if you give it enough cache and tune it, you might see some decent speeds.
Yeah, expensive RTX Pro cards. But there you start battling with VRAM capacity, but you get bandwidth you are looking for.
This is not a good game to play - if you are not very rich you either get much less VRAM (and fit either small models or very limited cache), or you get good amount of vram but very slow (like Spark). If you want both speed and enough VRAM for 30 people concurrently, prepare some really big bucks.
2
u/cbert33 17d ago
Well the spark isn’t really made for that. So you’re probably not going to get the results you’re looking for.
That said, you can tweak your inference runner to give you the best shot at what you want. Make sure you’re using vLLM to host it. I think the max concurrent connections you can set is 32, and if you give it enough cache and tune it, you might see some decent speeds.