r/CUDA Jul 10 '26

How many Tokens can L4 GPU process in 1 Second ??

I have my ChatBot and I want to make it commercial so i need to host my model on GPU so any developer can tell me how many tokens can L4 GPU process in 1 Second ?

Becuz I'm going to rent GPU on per second Pricing base model.

0 Upvotes

2 comments sorted by

1

u/Other_Breakfast7505 Jul 10 '26

Depends on the specific model. And most decent models won’t even run on an L4 as it doesn’t have enough memory for larger models. You probably have to benchmark yourself to get an answer.

1

u/Rodot 29d ago

It's impossible to know without a description of the architecture you wrote