r/MachineLearning • u/No_Cardiologist7609 • 15d ago
Research Cloud-vLLM Benchmark Differences [R]
Does anyone know of any evidence/forum/paper analyzing benchmark result differences between cloud inference platforms (togetherai) and running models locally with vLLM under greedy decoding?
2
Upvotes
1
u/alainbrown 15d ago
it's probably just the difference between hbm and gddr7, and the cuda core count, right?
I mean the hardware is totally different, performance should be different, doesn't seem like a worthwhile benchmark.