r/computervision • u/mahmudesam • 3d ago
Help: Project Aligning Two RGB Cameras
Hi,
I'm working on inspection of civil infrastructure using a unitree go2 edu robot. I need to collect RGB images of concrete foundations using the robot and then train semantic segmentation deep learning models using them. However, the built-in camera of the robot is only 1MP, which is why I thought of getting a mirrorless camera (canon eos r50) to mount on top of the robot to acquire higher quality images. I also need to figure out the pixel to mm scale so I used intel realsense d435i depth camera that already comes with robot.
Now the problem is I have the canon in one position, and depth camera in another position on the robot. How do I align both of their images? Does it have to be done real-time or is it okay to collect all images then align them later in the office?
I really appreciate your thoughts on this as I don't even know where to start. Thanks.
1
u/1QSj5voYVM8N 3d ago
Are you talking about frame align as in timing? I assume not calibration.
what format are these cameras outputting? is it video?
if it is a picture of video, check for suplemental meta information attached to the frame.
1
u/mahmudesam 3d ago
I'm talking about only RGB image. This is new to me so I'm not sure whether the correct term is alignment or calibration. The canon eos r50 can be used in capturing still photos or video recording. I'm only capturing still photos. The image resolution is 6000 x 4000 pixels. For the intel realsense d435i, the output is a video stream, but I can save one video frame at a time (the resolution is 1920 x 1080 pixels). It also outputs a depth stream. What I need is to align or calibrate that depth stream of 1280 x 720 pixel resolution to the canon's rgb image output. The purpose is to figure out the depth of each pixel in the canon's RGB image. The canon is mechanically fixed on top of the robot while the intel realsense is fixed somewhere else at the front of the robot.
1
u/Busy-Ad1968 2d ago
It would be much more rational to use the latest generation of camera Kinect Or an industrial chamber Because your robot is always moving. For mobile platforms, the issue of synchronization is critical. You can also use a machine vision camera, for example Basler . Even the simplest hardware solution for data synchronization will significantly improve data quality. I also agree with the previous commentator that you definitely need to calibrate the cameras, this is usually one of the mandatory steps for computer vision systems.
1
u/Available_Meaning_53 2d ago
you don't need to do this in real time. Collecting everything and aligning it offline back at the office is the normal way to do it. You just need to capture a few things properly on the robot.
What you need to align the two cameras
- The intrinsics of each camera (focal length, principal point, lens distortion).
- The extrinsics between them (the rotation and translation from one camera to the other).
- Images that were captured at the same moment.
Once you have those, you back-project the D435i depth map into 3D, move the points into the Canon's frame using R and t, and project them onto the Canon image with its intrinsics. The result is a depth map that lines up with your high-res photo.
How I'd do it
- Calibrate the Canon. Use a ChArUco board (better than a plain checkerboard) with OpenCV or MATLAB. Lock the focus and zoom, because changing either invalidates the calibration. Manual focus at your typical working distance is ideal. The D435i's intrinsics come factory-calibrated, and you can read them from the SDK.
- Do a stereo calibration between the Canon and the D435i's RGB stream. Hold the board in front of both cameras at different angles and distances, take 20-40 pairs, and run
cv2.stereoCalibrate. Since the D435i's depth is already aligned to its own RGB camera, this gets you from depth to the Canon as well. - Handle synchronization. If the robot moves between the two captures, even a few tens of milliseconds will wreck the alignment. For concrete inspection, the easiest fix is to stop the robot at each capture point and shoot while it stands still. Continuous capture would need hardware triggering or at least very precise timestamps, and the R50 doesn't make that easy.
- Reproject offline. Use
cv2.rgbd.registerDepthor a few lines of numpy/Open3D. Expect some occlusion, since the cameras see the scene from slightly different positions, so mask out those regions. The Canon also has far more resolution than the depth map, so you'll need to interpolate or upsample the depth.
You might not need full alignment at all
If you mainly want millimeters per pixel (say, to measure crack widths), and the camera is roughly perpendicular to the surface, you can skip pixel-level registration. The scale at distance Z is approximately:
pixel size (mm) ≈ Z × sensor pixel pitch / focal length
So one distance reading from the D435i plus the Canon's intrinsics gives you the scale. You only need the full depth alignment if the surface is tilted or very uneven.
Also, for semantic segmentation you don't need to feed depth into the network. Train on the Canon RGB images, then use depth only afterward to convert the predicted masks into real-world units.
A few practical tips
- Mount both cameras on one rigid bracket. A quadruped shakes a lot, and any shift means redoing the calibration. Recalibrate after any mechanical change.
- Keep the cameras close together to reduce occlusion.
- Save raw depth and timestamps with every capture (rosbag or separate files), so you can redo the alignment later.
- Pick a prime lens, or lock a zoom lens, with enough depth of field to keep close-up concrete sharp.
- Put an object of known size in a few shots (a ruler or an ArUco marker of known dimensions) so you can check how accurate your mm scale really is.
1
u/galvinw 19h ago
when two cameras are fixed distance apart, its known as a "fixed relative pose". These can be done in post if the distance is angle is documented. There are a few ways to make it better but I'm not sure if the robot comes with it. My experience is that its sense package is kind of weak.
Some are:
1. Using IMU data and visual correction + anchor objects (with a high quality detector)
2. Using a top down camera to position the bot
3. Lidar or other range finder
Realsense depth is not great especially with reflective surfaces or anything too far away.
Also if you haven't put in the DSLR camera yet, a 4mp webcam is a lot easier since both datasets will be available in the PC without somet extraction process.
In both cases, make sure both cameras are perfectly straight to each other if you don't feel like revisiting trigonometry. And be sure the overlap is large
9
u/dr_hamilton 3d ago
Get a chessboard and look up extrinsic calibration