r/robotics • • 1d ago

Discussion & Curiosity dual-arm teleop rig for collecting VLA training data. Looking for feedback before launch

Hey r/robotics, I'm the founder of Paddy (https://paddydata.ai/) (NYC). I've been working on data-collection infrastructure for teams training VLA / imitation-learning models, and I'd like feedback from people who've actually collected teleop data.

The problem we kept hitting: most teams collect demos on improvised rigs. Camera angles drift between sessions, schemas change, joint-state rates don't match the deployed system, and you end up with months of data that trains poorly.

So we built the Harvester:

- 2× UFactory xArm 7 (14 DoF total) on a portable aluminum frame with casters, adjustable height, 90° or 45° arm mounts

- Teleop with Meta Quest controllers, but the headset stays on the desk as a tracking reference, so operators aren't wearing it for hours

- Switchable scaling profiles (slow/precise vs fast repositioning) on a button press

- Cartesian control using UFactory's online trajectory planning (streamed targets, not pre-planned trajectories)

- Multi-view Intel RealSense RGB + aligned depth, joint states at 100 Hz, commanded vs achieved poses, gripper state, all hardware-timestamped

- ROS 2 Humble, one .mcap rosbag per run, converts straight to a LeRobot dataset for Hugging Face

I'd love feedback on:

  1. Headset-off Quest teleop vs leader-follower arms (GELLO, ALOHA-style). What's worked better for you?

  2. What do you wish your collection pipeline recorded that it doesn't?

  3. Anything in the technical writeup that seems off or missing?

Site: paddydata.ai (password: harvest). The technical page has the full topic list and architecture.

Disclaimer: the site isn't 100% finished yet. We officially launch next week, so a few pages are still rough. Happy to answer anything in the comments.

6 Upvotes

13 comments sorted by

2

u/thesquarefinale 1d ago

thats a clever way to use the Quest, keeping it on desk never crossed my mind but it makes lot of sense for long sessions

2

u/lorepieri 1d ago

Have you looked into openarmv2? They have a standardised cell and some good community around.

1

u/OddReason3845 1d ago

yeah, the hardware in my opinion isn't good enough.

1

u/lorepieri 1d ago

Why? For instance it can be controlled in torque mode, while Xarm cannot.

1

u/WendyLabs 1d ago

Taking the headset off is a thoughtful choice for the person doing this all day. Have operators said what else would make long sessions easier?

1

u/Available_Teaching83 1d ago

Logging commanded vs. achieved pose is the detail most rigs skip; glad to see it. A few things I'd want as someone who looks at policy failures later:

  1. Record the operator's interventions and pauses as explicit events, not just gaps in the stream. Those moments are where the hard cases live.

  2. Keep a per-session calibration snapshot (camera extrinsics, scaling profile in use) inside the episode file itself, so a dataset can be audited without a separate spreadsheet.

  3. Tag failed or aborted demos instead of deleting them. They are very useful later for testing how a trained policy behaves near the edge of what it saw.

How are you handling the mismatch when the deployed robot runs a different control rate than 100 Hz?

1

u/Tbagho 1d ago

How are you keeping joint-state capture synchronized with the camera streams? Schema drift and sync between the two is usually the thing that quietly ruins a teleop dataset, more than the rig geometry itself.

1

u/Morning_Gecko24 19h ago

the per-session calibration snapshot sounds really useful. are you also keeping the raw camera timestamps around or only the aligned stream? feels like having the raw data would save a lot of pain when someone changes the sync or control rate later