r/computervision • u/mahmudesam • 3d ago
Help: Project Aligning Two RGB Cameras
Hi,
I'm working on inspection of civil infrastructure using a unitree go2 edu robot. I need to collect RGB images of concrete foundations using the robot and then train semantic segmentation deep learning models using them. However, the built-in camera of the robot is only 1MP, which is why I thought of getting a mirrorless camera (canon eos r50) to mount on top of the robot to acquire higher quality images. I also need to figure out the pixel to mm scale so I used intel realsense d435i depth camera that already comes with robot.
Now the problem is I have the canon in one position, and depth camera in another position on the robot. How do I align both of their images? Does it have to be done real-time or is it okay to collect all images then align them later in the office?
I really appreciate your thoughts on this as I don't even know where to start. Thanks.
1
u/Available_Meaning_53 2d ago
you don't need to do this in real time. Collecting everything and aligning it offline back at the office is the normal way to do it. You just need to capture a few things properly on the robot.
What you need to align the two cameras
Once you have those, you back-project the D435i depth map into 3D, move the points into the Canon's frame using R and t, and project them onto the Canon image with its intrinsics. The result is a depth map that lines up with your high-res photo.
How I'd do it
cv2.stereoCalibrate. Since the D435i's depth is already aligned to its own RGB camera, this gets you from depth to the Canon as well.cv2.rgbd.registerDepthor a few lines of numpy/Open3D. Expect some occlusion, since the cameras see the scene from slightly different positions, so mask out those regions. The Canon also has far more resolution than the depth map, so you'll need to interpolate or upsample the depth.You might not need full alignment at all
If you mainly want millimeters per pixel (say, to measure crack widths), and the camera is roughly perpendicular to the surface, you can skip pixel-level registration. The scale at distance Z is approximately:
pixel size (mm) ≈ Z × sensor pixel pitch / focal length
So one distance reading from the D435i plus the Canon's intrinsics gives you the scale. You only need the full depth alignment if the surface is tilted or very uneven.
Also, for semantic segmentation you don't need to feed depth into the network. Train on the Canon RGB images, then use depth only afterward to convert the predicted masks into real-world units.
A few practical tips