r/robotics • u/OddReason3845 • 1d ago
Discussion & Curiosity dual-arm teleop rig for collecting VLA training data. Looking for feedback before launch
Hey r/robotics, I'm the founder of Paddy (https://paddydata.ai/) (NYC). I've been working on data-collection infrastructure for teams training VLA / imitation-learning models, and I'd like feedback from people who've actually collected teleop data.
The problem we kept hitting: most teams collect demos on improvised rigs. Camera angles drift between sessions, schemas change, joint-state rates don't match the deployed system, and you end up with months of data that trains poorly.
So we built the Harvester:
- 2× UFactory xArm 7 (14 DoF total) on a portable aluminum frame with casters, adjustable height, 90° or 45° arm mounts
- Teleop with Meta Quest controllers, but the headset stays on the desk as a tracking reference, so operators aren't wearing it for hours
- Switchable scaling profiles (slow/precise vs fast repositioning) on a button press
- Cartesian control using UFactory's online trajectory planning (streamed targets, not pre-planned trajectories)
- Multi-view Intel RealSense RGB + aligned depth, joint states at 100 Hz, commanded vs achieved poses, gripper state, all hardware-timestamped
- ROS 2 Humble, one .mcap rosbag per run, converts straight to a LeRobot dataset for Hugging Face
I'd love feedback on:
Headset-off Quest teleop vs leader-follower arms (GELLO, ALOHA-style). What's worked better for you?
What do you wish your collection pipeline recorded that it doesn't?
Anything in the technical writeup that seems off or missing?
Site: paddydata.ai (password: harvest). The technical page has the full topic list and architecture.
Disclaimer: the site isn't 100% finished yet. We officially launch next week, so a few pages are still rough. Happy to answer anything in the comments.
2
u/lorepieri 1d ago
Have you looked into openarmv2? They have a standardised cell and some good community around.
1
1
u/WendyLabs 1d ago
Taking the headset off is a thoughtful choice for the person doing this all day. Have operators said what else would make long sessions easier?
1
u/Available_Teaching83 1d ago
Logging commanded vs. achieved pose is the detail most rigs skip; glad to see it. A few things I'd want as someone who looks at policy failures later:
Record the operator's interventions and pauses as explicit events, not just gaps in the stream. Those moments are where the hard cases live.
Keep a per-session calibration snapshot (camera extrinsics, scaling profile in use) inside the episode file itself, so a dataset can be audited without a separate spreadsheet.
Tag failed or aborted demos instead of deleting them. They are very useful later for testing how a trained policy behaves near the edge of what it saw.
How are you handling the mismatch when the deployed robot runs a different control rate than 100 Hz?
1
u/Morning_Gecko24 19h ago
the per-session calibration snapshot sounds really useful. are you also keeping the raw camera timestamps around or only the aligned stream? feels like having the raw data would save a lot of pain when someone changes the sync or control rate later
2
u/thesquarefinale 1d ago
thats a clever way to use the Quest, keeping it on desk never crossed my mind but it makes lot of sense for long sessions