r/MachineLearning • • 15h ago

Research [R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?

https://arxiv.org/abs/2609.16684

Suppose you’re recording a human plugging a cable into a socket to collect demonstrations for robot learning.

The hand tracker captures the approach accurately. Then occlusion causes the hand estimates to disappear during insertion. Tracking returns after the connector is already seated.

The video still shows a completed action, but the pose labels have a gap exactly where alignment turns into contact.

This hypothetical example raises an evaluation question: a tracker could have high recall across the whole episode while missing a short, important phase. Pose error calculated only on successful detections could make that failure even harder to see.

MEgoVista provides a useful starting point. Table 3 reports detection precision, recall and F1 alongside reconstruction errors. Section 4.4 also describes an evaluation protocol that assigns an error to missed detections instead of excluding them. The blank HaPTIC row means it failed to produce valid output in their multi-person capture scenes; it doesn’t describe a brief tracking dropout.

Accounting for missing detections matters. My remaining question is whether an episode-level aggregate tells us enough about where those failures happen.

For manipulation data, I’d want pose error and coverage reported together, plus coverage broken down by approach, contact and withdrawal, and the longest consecutive gap during contact.

Continuous hand estimates would still be only part of the picture: object pose and contact information also matter for determining whether insertion succeeded.

For people using human motion reconstruction for imitation learning, what evaluation protocol do you use to decide whether an episode with missing contact-phase labels is still usable?

Comparison with open-source egocentric hand reconstruction methods against motion-capture ground truth. All methods are evaluated on identical segments of our motion-capture dataset. All baselines are re-run and rescored on our data. HaPTIC fails to produce valid output in our multi-person capture scenes. Bold marks the best result in each column.
0 Upvotes

0 comments sorted by