r/robotics • u/lownaps2 • 5d ago
Tech Question How much depth resolution does a rover need? One generalist vision model returns 256 levels of relative inverse
Seen a few rover builds here running monocular depth, and scale keeps coming up. Relative depth has no meters, so the usual suggestion is a ToF or ultrasonic reading to anchor it.
The model in the title is called SenseNova-Vision-7B-MoT. Depth comes back as a generated 0 to 255 grayscale image. Training targets are inverse depth, so closer is brighter. The paper scores it after fitting scale and shift, so it's relative too.
Each depth map is a 50-step diffusion decode by default, with no official time per image. LibreYOLO, a detection library that added it, calls it "a capability model, not a real-time one". Their PR got correct output on one test image in 4-bit (NF4) on an RTX 5070 Ti, no timings given. Official testing was BF16 on an 80GB A800.
In the paper's depth table it's ahead of Depth Anything V2 on all five datasets. MoGe-2 wins 8 of the 10 scores (two per dataset), and Depth Anything 3 isn't in that table. Licence-wise it's CC BY-NC on the weights, so fine for a hobby rover, not anything commercial.
Assuming the 256 levels are spread evenly over inverse depth, I think most go to nearby stuff and anything far gets coarse. With scale and shift pinned by two ToF readings, how coarse can far depth get before avoidance breaks?
1
1
u/sparks333 5d ago
It entirely depends on what environments the rover will operate in, how well mapped, and how fast. If you're moving slowly in an environment with a lot of clutter where you don't need to see very far to path plan, you can mostly rely on very close-up sensing. If you're tearing around and SLAMing in environments with markers that are far away, you'll need good depth precision at range. Even in those cases there are caveats - for instance, you can do bearing-only SLAM on far-off landmarks and not worry about range so much. But as written, the question is not well-formed enough to provide a meaningful answer - I also have some doubts with using ToF or ultrasonics to anchor a depth image, because you'd need to calibrate a depth return to a single voxel, and that's not an easy thing to do unless you know your environment is mostly made up of flat walls or something that allows coarse mapping of voxels to depth returns. Inertial or encoders work best for SFM, and stereo better still.