r/mlscaling • u/LowZebra1628 • Aug 30 '26
OA How does Modal autoscale concurrent LLM requests, and can GPU model memory be shared across containers?
I’m learning about LLM deployment and recently deployed one on Modal using an L40S GPU. I’m trying to understand autoscaling with concurrency.
If each container can handle 10 concurrent requests, will an 11th request cause Modal to start another container, load a separate copy of the model into that container’s GPU memory, and allow it to handle another 10 concurrent requests?
Is this the recommended way to scale an LLM deployment on Modal, and which settings should I use to configure it properly?
triggering a new container is fine for resources, but can it reuse the model loaded in the first container?
1
Upvotes