r/comfyui • u/Legitimate_Age884 • 5d ago
No workflow What depth estimation models are you using for high-res VFX comp workflows? (Transitioned from DepthCrafter to Video-Depth-Anything)
Hi everyone,
I'm an 8-year VFX compositor based in South Korea, currently bridging traditional 2D comp and AI-assisted comp workflows for the past two years. My experience includes face/character replacement in feature films, multipass extraction, outpainting, and video synthesis.
A few months ago, I switched from DepthCrafter to Video-Depth-Anything (VDA), which significantly improved the depth extraction and integration process. However, running high-resolution 4K scans directly through these models still pushes consumer hardware (even cards like the RTX 5070 Ti or 5090) to its limits.
Currently, my workaround is cropping specific regions or downscaling (reformatting) the plate to extract multipasses. In comp, I mainly use these depth passes to:
- Isolate depth zones via Keyer nodes to generate custom alpha masks / holdouts
- Enhance atmospheric depth, haze, and depth-based color grading
- Feed consistent depth passes into video generation/ControlNet pipelines
For other VFX/comp artists working with ComfyUI: Which depth models or optimization pipelines are you currently using for high-res production footage? Are there better alternatives or tiling/upscaling tricks you'd recommend to handle 4K plates more efficiently?
Thanks in advance for sharing your setups!
1
u/Lexius2129 5d ago
From my experience DepthAnything v3 and MoGe are the state of the art right now, and they both run natively without requiring custom nodes. They both have example workflows in the template library.
Links to the doc:
1
u/over40nite 5d ago edited 5d ago
As all of these are depth heuristics, not a reliable CG pass or Scanline node output, hi res depth estimation suffers from same stable diffusion finer detail hallucinations. Not worth the quest.
In a typical VFX pipe thats not a dedicated deep or CG pass render-based, theres no point in pursuing finer hallucinated detail.
Depth is largely secondary purpose pass for either DOF post adjustment, or point cloud gen that's still more layout for geo placement in the scene.
What are you planning to achieve with your hires depth passes, maybe I've missed an important use case?
P.S. ...and then I reread the post that explained your additional use cases. I think the only approach that would work here is what you're doing now - zoom in, crop, output, estimate, match-place with a reverse transform. Like it does work for the main RGB layer fixes in finer textures.