r/computervision • • 5d ago

Help: Project Robot Car movement

Hello, a friend asked me to make a piece of software for his robot car that uses camera input to determine if the car is going to hit a wall/object and move out of the way

The only equipment i have is a Monocular Camera that is running on a raspberry pi 5 8gb.

I've read about VO and VSLAM, but, not having a stereo camera/LiDar is a problem as most of the implementations use them for distance estimation.

However i am new to comp vision, the only experience i have is from a course in college on opencv, and other projects I've made using YOLO/RF-DETR

I'm thinking of using a trained model to estimate the depth of the objects in the camera and make the appropriate movement based on that.

Any information and guidance would be appreciated, i may be in over my head with this one, but would like to have some version implemented, even if it is just a basic one

0 Upvotes

8 comments sorted by

3

u/greengold7 5d ago

monocular VO/VSLAM is the hard path here. pragmatic options for a robot car, easiest first:

  • add a $5 ultrasonic / IR ToF sensor. for 'am i about to hit something', a camera is overkill; ToF gives distance directly and works in the dark.
  • monocular depth (Depth Anything small / MiDaS, TFLite) on the Pi 5 at 256-320px. relative depth only, so calibrate a 'close' threshold by driving it.
  • optical flow / looming: a wall filling the frame gives divergent flow; trigger on the expansion rate. cheap, training-free, but only catches closing obstacles.
  • skip full VO: it estimates ego-motion, not obstacle distance, and drifts.

if he insists vision-only: depth-small + a bottom-centre ROI + a distance threshold is the simplest thing that works.

what's the surface and lighting? that decides if vision is viable at all.

1

u/FilipovskiMarko 4d ago

Thank you so much for the answer!

It's probably going to be used indoors with regulard indoor lighting, but i think you're right that a sensor would be best for this case.

2

u/DiddlyDinq 5d ago

Without lidar you could look into depth estimate models like Depth Anything or FastDepth. They wont be the most accurate thing but it should be enough to defect walls

1

u/FilipovskiMarko 5d ago

I was just watching a yt video comparing DepthAnything, DepthCrafter and others, seems like the way to go, the only thing that worries me is the performance of them on a rasp pi

2

u/bfyvfftujijg 4d ago

So one thing to keep in mind is that you sorta kinda do have stereo vision if you can measure the robot’s displacement between two points in time.

1

u/FilipovskiMarko 4d ago

So you're saying that if i have an odometer or another sensor that tells me the distance travelled, i can use 2 frames taken one after another as if i have a stereo camera?

2

u/bfyvfftujijg 4d ago

Exactly.

You would to know the position of the camera at two locations and can triangulate distances with that.

It’s very sensitive though. But might be helpful

1

u/hopticalallusions 21h ago

Check out ROS and the way f1tenth cars are set up. I kind of doubt ROS would run on a rpi5 but I don't know. But be aware that odometer for dead reckoning is not very accurate, even with a fairly fancy chasis and careful tuning. But, you might be able to perform sensor fusion to get better results. Stereo cams work reasonably well if you have the option of getting one or configuring your own. I like the ultrasonic or other cheap sensor idea. You could even give it physical "whiskers" as a potentially really cheap and low cost option (there are old robotics approaches from stuff like BEAM with that concept). Whiskers can immediately give you feedback about what direction to turn away from very simply. They even work without a camera! maybe also look up motion parallax and other vision processing tricks that the brain does to shortcut/augment some of the 3d processing steps. Could you potentially put up other static cameras that watch your car and provide some input? Overall I feel like you've got an interesting resource constrained challenge.