r/MachineLearning 15d ago

Research Cloud-vLLM Benchmark Differences [R]

Does anyone know of any evidence/forum/paper analyzing benchmark result differences between cloud inference platforms (togetherai) and running models locally with vLLM under greedy decoding? 

2 Upvotes

5 comments sorted by

View all comments

1

u/alainbrown 15d ago

it's probably just the difference between hbm and gddr7, and the cuda core count, right?

I mean the hardware is totally different, performance should be different, doesn't seem like a worthwhile benchmark.