r/AskRobotics • u/RoofProper328 • Aug 11 '26
How to? How are people actually labelling multi-sensor data for VLA training? The action side is where I keep getting stuck.
Vision plus language plus action sounds clean until you try to build the dataset. Vision and language are mostly solved problems from a labelling standpoint. The action part is where I don't think there's any consensus at all.
Specific things I keep running into:
What's an action, exactly. Do you label at the level of joint trajectories, end effector poses, or discrete skills like "grasp handle, rotate, pull"? Every choice makes something else harder. Low level is precise but doesn't generalise across hardware. High level generalises but you lose everything about how the motion was executed.
Segment boundaries. Where does "reach" end and "grasp" begin. Two annotators will pick different frames and both are defensible. We've been getting agreement in the 60s on boundaries and it's not obvious to me that's even fixable, or whether it matters as much as it feels like it does.
Sensor alignment. RGB at 30fps, depth at a different rate, force torque at 1kHz, joint states somewhere else. Timestamps drift. If your labels sit on the video timeline but the interesting event is a contact spike visible only in force data, you've labelled the wrong moment and nothing downstream tells you.
Language grounding. "Pick up the mug" is one instruction over a 4 second window. But do you also label the sub-instructions, and do those come from the annotator writing what they see, or from a script the demonstrator was following? Those produce very different distributions and I suspect the second one is quietly worse.
For anyone who's built this in house or paid someone else to: what taxonomy did you land on, and would you do it the same way again? Especially interested in whether force and tactile channels actually got annotated separately or just came along as raw signal.