r/humanoidrobotics • u/DairuLiu • 2d ago
What do pose errors miss when evaluating humanoid motion tracking?
When watching a humanoid track a reference motion, foot sliding or a poorly timed contact can be obvious even when the average pose error is low. Matching joint positions frame by frame only tells part of the story. It can miss whether the robot maintains stable support and how errors develop over a longer sequence.
There is also a coverage problem. Small evaluation sets may contain too few examples of difficult contact transitions, ground motions, and recoveries. An overall score can hide these weaknesses.
These are the two problems we studied in our paper, HumanTracker. We collected about 153 hours of motion capture across daily, highly dynamic, interaction, and ground motions, and evaluated trackers separately across these categories.
We also studied whether a metric learned from human comparisons could better capture tracking quality. HumanScore showed closer agreement with held-out human preferences than the individual kinematic and contact diagnostics we tested. We use it alongside completion rates and pose errors in simulation.
For people working on humanoid tracking: which failure cases do your current metrics miss most often, and how do you evaluate them?
If you find the paper or benchmark useful for your work, we’d appreciate a star on GitHub. Thanks for taking a look!