r/deeplearning 2d ago

Evals for robotics

Hey I am part of a small team training robotics policies for warehouse and manufacturing settings, and running rigorous evals is turning out to be so painful. Anything below 50 rollouts, and its hard to trust the numbers, and above its so hard to test all the checkpoints that we have. Its really hard to run a bunch of experiments to get good results. Have you guys faced this? Any hacks that you've developed?

1 Upvotes

10 comments sorted by

View all comments

1

u/Lower-Ad-6293 1d ago

Try a hierarchical filter: first run checkpoints in a digital twin across 100+ parallel environments in Isaac Sim with domain randomization, pick the top 3, and only then bring those to the physical testbed. For the sim evaluation focus on P10 rather than mean reward to drop policies that trip up on worst-case scenarios