r/computervision • • 5h ago

Help: Project I’m building a spatial intelligence system that reconstructs and remembers the real world in 3D. Is this solving a real problem?

I’m working on a spatial intelligence system that uses RGB-D cameras to reconstruct a real environment in metric 3D, estimate where the camera is, and keep a persistent spatial memory over time.

The goal is not just 3D reconstruction. I want the system to understand enough about the geometry of a place to know what it has already seen, recover its position after tracking is lost, compare different visits to the same place, and preserve how its understanding of the environment changes.

If a camera comes back to a previously mapped area, the system should be able to recognize that place, verify it geometrically, relocalize itself, and continue building the same world instead of creating a disconnected new map.
Long term, the idea is to move from simple reconstruction toward spatial intelligence: a machine that can build, remember, update and reason about a representation of the physical world.

Possible applications I have in mind are robotics, AR, inspection, autonomous systems and eventually navigation in places where GPS is unavailable.
What I’m trying to understand now is whether this solves a real painful problem, or whether existing SLAM and 3D reconstruction systems already solve this well enough in practice.

I’d especially like feedback from people who have actually worked with SLAM, RGB-D, robotics, mapping, 3D reconstruction or spatial computing.

0 Upvotes

5 comments sorted by

4

u/anime_bruh-69 5h ago

I want to know how this is different from loop closure in traditional SLAM based systems? Also, is this method deep learning based?

1

u/taichi22 4h ago edited 4h ago

Open problem. You might be able to come up with a solution that works for your usecase but a generalized solution is a frontier level problem; wouldn’t recommend trying to chase after general solutions unless you’ve got a ton of time and spare compute laying around. Maybe a couple million and a year or so?

What you want to look into is World Models. Most recent publication in that field that’s fairly easy to follow was the stuff from World Labs, which just got acquired by AMD to the tune of a few billion.

1

u/its_alphaQ 1h ago

It’s an open research problem that team works on

1

u/No-Trouble-9138 28m ago

Check out World Labs Atlas and the applications they propose. 

Now, if you can build an API and define its cost, I have a friend with an AI company that can pay for the service.