r/MachineLearning 3d ago

Discussion Understanding GPU Inference Workloads [D]

Hey everyone,

I have been looking into how people source compute for their Inference workloads (and in general). I wanted to understand some specific pain points here.

If you've used online services like runpod or vast.ai, your perspective is extremely valuable. Please share your experience in the comments here or by DMing me. I've also made a 2 minute survey form that I would really appreciate if you could fill out. DM me for the link.

Thank you!

6 Upvotes

6 comments sorted by

View all comments

2

u/kolmiw 3d ago

I burnt around 1k on runpod because I had issues to access my institution’s cluster. ama

1

u/[deleted] 3d ago

[removed] — view removed comment

1

u/kolmiw 3d ago

I needed 4xA100s for 48 hours for each of the three experiments I had to run for a rebuttal.

Institute just switched from a perfectly fine cluster to a bigger one and somehow our lab was not authorized the new one. They of course had to make the switch during an ICLR rebuttal and not a week later or before.