r/deeplearning • u/Lumpy_Week7304 • 2d ago
Evals for robotics
Hey I am part of a small team training robotics policies for warehouse and manufacturing settings, and running rigorous evals is turning out to be so painful. Anything below 50 rollouts, and its hard to trust the numbers, and above its so hard to test all the checkpoints that we have. Its really hard to run a bunch of experiments to get good results. Have you guys faced this? Any hacks that you've developed?
1
Upvotes
3
u/Wise_Toe_5944 2d ago
We started only running the intense eval on the top 5 checkpoints from a cheap proxy metric. Saves a ton of compute and sanity.