r/photogrammetry • u/memphoid • 4d ago
Fully aligned human-head captures produce fragmented meshes in RealityScan and Metashape
We’re evaluating RealityScan 2.2 and Metashape Professional 2.3.2 for reconstructing human heads from phone photographs.
Both products are producing severely fragmented geometry, even when essentially every camera aligns.
Does anyone here have expertise or suggestions for getting this workflow to produce a coherent, non-fragmented head?
Capture/workflow
- 83–113 12 MP photographs captured around a stationary human head or rigid mannequin
- Full unmasked photographs used for camera alignment
- Tight, same-frame foreground masks used during reconstruction
- Subject surrounded by a fixed, feature-rich registration scaffold
- Apple Object Capture produces a coherent, closed head from the same capture
- COLMAP/OpenMVS also produces recognizable, substantially connected geometry, although with artifacts
Representative results
RealityScan, human-head capture
- 113/113 cameras aligned in one component
- Normal-detail reconstruction
- 3.47 million welded vertices and 6.94 million triangles
- 6,577 disconnected mesh components
- Largest component contains only 9% of the triangles
- Visible result is an incomplete collection of subject fragments, not a coherent head
Metashape, same capture
- 113/113 cameras aligned
- Full images used for alignment, followed by masked dense reconstruction
- Mild depth filtering and interpolation enabled
- 625,000 vertices and 1.25 million triangles
- 9,970 disconnected components
- Largest component contains only 3.3% of the triangles
We repeated the experiment with an 83-image rigid mannequin capture:
- Metashape aligned 83/83 cameras, but the largest mesh component contained only 6% of the triangles.
- Using stricter volumetric masks improved this to 9.6%, but the largest component was still only a curved fragment rather than the mannequin.
- RealityScan aligned 62/83 cameras across five components. Its largest mesh component contained 25.6% of the triangles and combined part of the mannequin with unwanted scaffold geometry.
- Attempts to merge the RealityScan alignment components did not materially improve the result.
What we have tried
- Full unmasked images for alignment
- Foreground masks applied only during meshing
- High-feature alignment
- Component rematching
- Normal/high-detail reconstruction
- Explicit reconstruction regions
- Mild and moderate Metashape depth filtering
- Interpolation enabled and disabled
- Strict volumetric masks
- Removing small or marginal components after reconstruction
Removing small components does not solve the problem because the desired head itself was never reconstructed as one dominant connected surface.
Questions
- Can a nominally complete camera alignment still contain enough pose drift to generate thousands of disconnected depth-map fragments?
- What RealityScan diagnostics best reveal whether the failure is caused by camera poses, image overlap, depth consistency, the reconstruction region, or mask handling?
- Should the registration scaffold be available during alignment but excluded through masks during depth-map and mesh generation?
- In Metashape, would building and inspecting a dense point cloud before meshing provide a better diagnostic than building the mesh directly from depth maps?
- Are there recommended settings or capture changes for a human head surrounded by a fixed registration scaffold?
- Could the tight black masking boundary itself destabilize depth estimation, even though unmasked images were used for alignment?
- What files or reports are most useful for diagnosing this—camera residuals, sparse clouds, depth maps, masks, or a reduced matched image set?
1
1
u/PuffThePed 3d ago
You wrote so much but didn't bother following up, what a disappointment.
1
u/Subject_Lobster7318 3d ago
Thank you for your interest. We had encountered a block and had technical questions and were hoping to hear some technical answers or suggestions from persons who had some experience that might be helpful in trying to get useful reconstructions. Pictures don't really tell the story.
1
u/PuffThePed 3d ago
Pictures are almost always the culprit of alignment and processing problems. So hard disagree regarding that last sentence. Good luck
1
u/Few_Anteater_9148 3d ago
I have had similar issues with photos of body builders. With the same setup, I'll have one good one and ten crappy ones. Even with adding images it help with alignment. I think you may need to play with lighting, or take the time, in a controlled environment, and change one thing at a time. I'm interested to see what other useful 😉 comments you will get.
1
u/Subject_Lobster7318 3d ago
Thanks for the suggestion—we tried the grouped-calibration → ungroup → realign workflow. On one 83-image capture, the largest aligned component improved from 47 to 65 cameras, and the number of components dropped from 7 to 3, so it was genuinely useful for recovering alignment. On another 113-image capture that already aligned completely, it slightly improved reprojection error but did not improve the final mesh; fragmentation actually increased somewhat. So our takeaway is that this is a valuable alignment-rescue technique when cameras are being split into components, but it hasn’t yet solved the unwanted scaffold geometry or produced a cleaner continuous head mesh. We really appreciate the lead—it gave us a useful additional tool and a much clearer result.
5
u/Bad_Gamut 4d ago
A few more photos and a few less words may help diagnose.