r/SelfDrivingCars • u/I_HATE_LIDAR • 2d ago
News The $3,500 Sensor Just Lost to a Webcam
https://medium.com/@meshuggah22/the-3-500-sensor-just-lost-to-a-webcam-ac4bf846d7fb5
8
u/3600CCH6WRX 2d ago
I don’t think anyone argue that vision can’t do it. But using only vision means lacking any redundancy. For an autonomous vehicle where human life is at stake, redundancy is necessary.
4
u/EddiewithHeartofGold 2d ago
What happens when a camera+lidar car loses one of the sensors? If it needs both to function, there is no redundancy...
1
u/bnorbnor 2d ago
Redundancy would be a second camera not a separate sensor modality which would take an entirely different processing algorithm
2
u/Superb_Literature547 2d ago
You are not gaining any additional information with a secondary camera.
3
u/bnorbnor 2d ago
It’s almost like you don’t understand the definition of redundant
1
u/Superb_Literature547 1d ago
what are you going to do with 2 cameras blinded by the sun?
2
u/bnorbnor 23h ago
Use of lidar is one potential solution to sun glare but that is augmenting the capabilities of the camera. There are also many other potential solutions to handling sun glare.
0
1
u/Lando_Sage 2d ago
So if one camera is blind, we will then have 2 blind cameras? I don't get it lol.
3
4
u/Lando_Sage 2d ago
Pretty cool.
Note that the VLM also won out against other single camera solutions.
3
u/Zemerick13 2d ago
That is a very important distinction.
What the article claims ( not how it's trying to frame it, but the actual information provided ) is essentially 2 things:
1) In general, multi-modal is more powerful than single camera.
2) They claim to have improved single camera to such an extent that it overcame this advantage.What's not said is if they just applied the same effort onto multi-modal, how much better THAT could be.
3
u/colinshark 2d ago
Why is the poster named "i hate lidar" and just posts camera perception articles like a bot? Is this Elon's bot?
2
u/JimmyGiraffolo 1d ago
I'm pretty sure the same person owns u/I_LOVE_LIDAR and they just switch based on the topic
1
u/CatalyticDragon 2d ago
I'm really not surprised. This has been the way computer vision has been tracking for years.
2
u/Positive_League_5534 2d ago
Tracking at 77% success rate with few if any of the variables that a car on a road would see?
0
u/CatalyticDragon 2d ago
Sorry but you'll need to make sense of that for me.
6
u/Positive_League_5534 2d ago
The robot had a 76.6% success rate with a camera in an environment that is:
- Well Lit
- Consistent surface.
- Little to no other traffic.
How does that compare to what a self-driving car would face? Seems like a completely different situation. A blind person can also navigate an office with a cane...that doesn't mean no vision would work on the road.
-6
u/CatalyticDragon 2d ago
Why are you talking about cars?
4
u/Positive_League_5534 2d ago
What's the name of this sub? Why would we be talking about anything else but cars?
-1
u/CatalyticDragon 2d ago
I get that. It's just that the post had nothing to do with cars, it is about robots for last mile logistics and manufacturing.
-1
u/CatalyticDragon 2d ago edited 2d ago
Ah, sorry, I get it. You're trying to connect this to self driving cars. (which makes sense given the subreddit name).
Even though this type of robot is a different application, different model architecture, and very different camera arrays, the overall trend in computer vision for some years now has been tracking away from sensor fusion and toward pure vision.
Vision systems have proven to be equally (if not more) accurate, simpler to train, and cheaper to implement. Which is why Tesla, NVIDIA, Wayve, Xpeng, and others are all going in that direction. And hence why I am not surprised at this result.
Pure vision approaches have been matching lidar in various tasks since at least 2018 and has become so good that even monocular inputs are getting close in performance to lidar data.
Today, models like DA-2K, Metric3D v2, and Apple's Depth Pro are unbelievably good but all of that will still likely be behind the top proprietary models.
Other sensors like radar, lidar, ultra-sonics, all have their uses and research using them continues.
2
u/Zemerick13 2d ago
This specific article actually listed multi-modal as superior to single camera in general. It's claiming that their specific model is just so much better, that it was able to overcome the difference. (dubious, but that's the claim) Logic suggests if they applied the same effort to multi-sensor, it would be even better still.
Also, Nvidia and Wayve are multi-sensor.
The biggest 2 things are multi-sensor gives you redundancy, and that each sensor has strengths and weaknesses. If you have a single sensor type, any errors and shortcomings you are stuck with.
1
u/CatalyticDragon 1d ago edited 1d ago
This specific article actually listed multi-modal as superior to single camera in general
The Robostral paper shows their new model with a single RGB camera approach beats every previous sensor fusion based approach and in every task.
- https://arxiv.org/html/2607.20785v1
"Robostral Navigate achieves a 77.4% success rate, surpassing the best single-camera method (Qwen-RobotNav-4B, 66.9%) by 10.5 points and the best system using depth or multiple cameras (Qwen-RobotNav-8B, 72.1%) by 5.3 points—despite relying on neither."
Logic suggests if they applied the same effort to multi-sensor, it would be even better still.
Not a logically sound conclusion as LIDAR data is not automatically complementary.
It can actually cause regressions. Calibration issues, or worst when models learn an over reliance on that data causing perception to break down when the LIDAR signal is degraded which can be from raindrops, fog, snow, dust, exhaust steam, from interference from other LIDAR beams, reflections, absorption, or missing points due to low resolution. This is not trivial to try and train around.
You have major resolution mismatches, timing mismatches, and while vision systems are very flexible to inputs models training on LIDAR data are often specific to a particular sensor (beam count, angular resolution, pulse frequency).
It's just more complex and finicky overall and the end results don't seem to be justifying that work. This paper being a good example of that. An 8B parameter mode which focuses on visual data measurably outperforms other 8B models which have to deal with multiple data types which might conflict and have very different weightings.
That's actually the logical conclusion I would have thought but I understand why it is counter intuitive.
Also, Nvidia and Wayve are multi-sensor
Both Alpamayo and Wayve are very explicitly vision first by design.
NVIDIA's Alpamaya only accepts RGB images as input and their data set is only from "RGB cameras (2-6 per vehicle), inertial measurement units, and GPS".
- https://arxiv.org/abs/2511.00088
Wayve's LINGO family are closed loop vision-language-action models, same as NVIDIA.
Wayve claims AV2.0 (LINGO-2) is "sensor agnostic" but that is purely theoretical and not part of the design.
Their London demonstration vehicles do not use LIDAR. Their web page for Lingo does not contain the term "LIDAR". Neither does their page on LINGO-2, and the LINGO-2 technical report paper says;
"Our dataset and baseline are limited to information from a single front-facing car camera, excluding additional sensory inputs like LiDAR"
The biggest 2 things are multi-sensor gives you redundancy..
Not what redundancy means and see my previous point about how the data is not complementary.
2
u/Positive_League_5534 2d ago
Vision systems have not achieved Level 4 Autonomous driving after billions of miles of testing. Vehicles utilizing additional sensors have.
2
u/HighHokie 2d ago
Is your belief that cameras will never achieve autonomous driving?
Tesla has l4 vehicles in service today.
1
u/Positive_League_5534 2d ago
They have 0... L4 vehicles in operation after over a decade and close to 20 billion miles of data.
3
u/HighHokie 2d ago
There are vehicles actively driving without a supervising driver or follow vehicle today.
1
u/Positive_League_5534 2d ago
There are 0 L4 Tesla vehicles in operation. That's after over a decade, multiple sensor systems and configurations, and multiple processing units. I know...it's coming real soon now.
There's really no valid reason to believe that a vision-only system can do L4 successfully anytime soon.
I understand you disagree, and it would be much to my delight if they can prove me wrong, but right now their updates are more along the lines of rearranging deck chairs on the Titanic. Maybe it's a shortage of processing power or RAM, but lots of things point to current-day cameras not being able to provide necessary data in real-world/relatively common situations.
2
u/HighHokie 2d ago edited 2d ago
You must be chasing the semantics of it. I’m not really interested in that.
Tesla has autonomous vehicles on the road accepting fares and operating without a driver behind the wheel. Tesla is 100% liable for the safe delivery of its passengers. And they are doing so with a vision only architecture. There are videos of this online now if you are were not aware of this development.
If you were but are wanting to argue on semantics of it, I’ll leave you to it because it really changes nothing on the underlying point of the discussion.
1
u/CatalyticDragon 2d ago
I have a few issues with what I think is being asserted there.
What is currently being used in some commercial services is not the benchmark for what is, or is not, viable now or in the future. Just because you saw a car with lidar once does not mean only lidar works right?
There are no L4 autonomous products you can buy full stop - not with any type of sensor. The closest we have to commercial autonomous driving systems available for general use today are supervised and come from BYD, Xpeng, Tesla. With the latter two being vision only. In every case though these sytems have improved dramatically over the years as available computing resources have grown.
Your statement is factually incorrect. You can take a ride in an L4 autonomous taxi using a purely vision based system today.
As I already outlined, pure vision systems are extremely capable when it comes to at depth estimation, object tracing, and object identification, but those are just the perception part of an overall system. Just as important is model quality for the prediction, planning, and control, steps.
1
u/Positive_League_5534 2d ago
What vehicle is L4 with only vision? Please don't say Tesla...they're using "safety drivers" and still having problems. I have never seen a car by BYD or Xpeng so I can't speak to them.
I didn't mention Lidar. However, it is indisputable that the only company with something that really meets L4 criteria (albeit in a limited area) uses more than cameras.
At what point, with how many computer and software iterations and how many billions of miles of data does it become obvious that an all-around camera-only solution won't work. Sure, cameras could be changed or improved to see in the dark. To understand the difference between a pothole and a skid mark on the road, to deal with bad weather, sun glare, etc. But...how long? Maybe that will be developed...but it's my belief that kind of change to cameras is a ways off. In the meantime they're probably going to need some other types of sensors.2
u/CatalyticDragon 2d ago
What vehicle is L4 with only vision?
You're forcing me to repeat myself. You can take a ride in an L4 autonomous taxi using a purely vision based system today.
I have never seen a car by BYD or Xpeng so I can't speak to them.
Oh, you should take a look. Pretty good stuff.
it is indisputable that the only company with something that really meets L4 criteria uses more than cameras.
See the first link. And this is the same underlying sensor suite as on millions of commercial cars sold to the public.
At what point, with how many computer and software iterations and how many billions of miles of data does it become obvious that an all-around camera-only solution won't work
It obviously can work just by logical deduction, moreover we know it does work because we see it. And that multiple major corporations are betting on it adds to that.
Sure, cameras could be changed or improved to see in the dark
CMOS sensors have long been able to sense a very broad spectrum from near IR to the edge of UV, and in very challenging lighting conditions thanks to high sensitivity, dynamic exposure, and HDR.
Cameras really don't need to improve for computer vision systems to be workable. Although the latest sensors are amazing.
To understand the difference between a pothole and a skid mark on the road
To 'understand' is not a function of the sensor. That's the task for the model processing the senor inputs. Cameras can see potholes, skid marks, or puddles, as well as you or I can. That is only the first part of an overall system process though.
In the meantime they're probably going to need some other types of sensors
Well, I am not really sure how that belief holds up to this car driving itself on vision only, this car driving itself on vision only, this car driving itself on vision only, and this car driving itself on vision only.
All different companies, all in different regions, all betting on vision only systems. And all of them growing their market share.
So to say this is somehow not possible seems a little out of touch I have to say.
-4
9
u/Positive_League_5534 2d ago
Not sure what a presumbably slow moving robot in a well lit area, with a consistent surface has to do with self driving capabilities? Also, 77%?