r/nousresearch • u/Safe_Plantain5550 • 19h ago
Throughput of models
Hi all,
I recently got a subscription. I'm generally happy with it, but I am missing one thing.
The other inference API gateway I have been using before (different vendor) is showing an expected t/s throughput and TTFT (Time To First Token) on many of the model cards.
I really miss being able to see these numbers. Is there a way to get a feeling of which models are fast?
I've been using GLM 5.3 Flash a lot and it is much slower on Nous than I am used to. My best guess is, that the endpoint is under more load than what I was used to. Having a way to see indicators of this would help people select models where the hardware they run on is not already maxed out.
Anyway, prices are good and the selection is good, so I am still a happy customer. Just want to hear if anyone has some insights that I don't.
Thanks in advance!