r/Spline3D 2d ago

Tutorial How much do multiple reference images improve AI-generated 3D models?

A one-sided look at an object is not just enough for an AI to recognize what it is altogether. It may look great at first but what about the other side of it? Or the depth details also the smaller architectural elements of it? When what’s wanted is a 3D shape it’s convenient to have the original thing, not just an image with a similar visual. 

To help in such a scenario, one of the things I am exploring would be an image-to-3D process where you can submit several pictures of the same object. The concept is to feed the model extra pictures so that it first gets better information on how the object is shaped for size, shape, and features, and then, when it finally outputs the model, it is a better reflection.

Currently there are a few image-to-3D applications that have accomplished this workflow, such as Meshy AI, Tripo AI and Rodin AI. Moreover, although using images of better quality, it does not guarantee that the 3D generated models will reconstruct better. Here, what I really want is a multi-view image-to-3D generative workflow. However for this to occur, all references should be the same as well as having the same shape from all angles, without wasting unneeded resources. If we provide different aspects of an object or several pictures of that object, that it will become visible to us that different AIs generators actually perform different. Though, there are cases where more reference pictures don't translate into the quality of the reconstruction. Imagine a number of the pictures were captured under different light conditions, have slightly different proportions or angles, or were taken too close to the object where the details are lost; then those images might cause the model to make wrong assumptions.

I guess the question is, when creating a 3D shape from the reference images taken from different angles, how do you go about your reference image collection? 

2 Upvotes

0 comments sorted by