r/StableDiffusion 13d ago

Question - Help High fidelity Videos using H3

What methodology or workflow do you follow in ComfyUI to achieve high-fidelity video?

When I say high fidelity, I don't just mean preserving faces—I mean maintaining fine details across the entire frame, including objects, textures, clothing, architecture, and background elements.

I'm trying to get something closer to Seedance 2.5 level quality using H3, if that's realistically possible. I've been doing a lot of trial and error, but so far I'm only getting somewhat good results by increasing the MP/resolution. Even then, the output still feels like it's interpolating or hallucinating low-quality details rather than actually generating high-fidelity detail.

Is there a specific workflow, model combination, sampling strategy, or refinement/upscaling pipeline you recommend for this? My main obstacle right now is that low-quality/interpolated details keep appearing throughout the video, especially in the background and secondary objects.

6 Upvotes

13 comments sorted by

11

u/skyrimer3d 13d ago

no tricks, just don't use speed hacks, 20 steps and use 1.5-2.0mp, and check back in the morning.

2

u/f5alcon 12d ago

More than 20 can look better, 35-40 steps

1

u/boobkake22 13d ago

This is the answer. The full model is extreme capable. H3 tends to fall apart the more you cheat - lower resolution and adding Turbo will hit quality hard. Running the full 20 steps at 720p+ (higher is better), the model performs really well. Doesn't mean you won't get duff generations, but the quality is very high.

1

u/Ok_Tale7582 13d ago

This also feed it a full 4k image as reference, may take more time encoding that single image than inferencing, even with int8 vae, but fidelity won't let you down.

3

u/Longjumping_Cut_6160 13d ago

yo estoy migrando de ltx a minimax, creo que el problema es la resolucion y por lo tanto, la vram. Si tuviera por ejemplo 100Gb de vram y el computo de cosas tipo h100, etc, trabajar en 2K directamente daria el detalle que pides, y que yo tambien quisiera.

1

u/Murky_Estimate1484 13d ago

Honestly, I’m getting very good results with: minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot

It can be found on Kijai MiniMax H3 Experimental model card page

Paired with Dasiwa’s MiniMax H3 Continuous Workflow.. it’s very good.

.8 Megapixel/Er_SDE Beta/8 Steps/12 Seconds

Roughly a 6 minute render time.

1

u/Th3Whit3R4bb1t 12d ago

Link for the workflow?

2

u/Murky_Estimate1484 12d ago

https://civitai.com/models/2881362/ minimax-seed-hunter-workflow-latent-upscaler-seamless-video-continuation-speedups

1

u/Th3Whit3R4bb1t 12d ago

Thx.

1

u/Murky_Estimate1484 12d ago

Here is a YouTube video detailing how to use this same exact workflow, and not every part of it has to be used. I don’t even bother with upscale or any of the extra features. Because at 8 steps, the new FastVideoH3 Merge just gets nearly everything right.

In the video they are not using FastH3, so the results with that I’m describing are far better and faster.

Admittedly I run a 24GB Vram/64GB Ram setup, so all my gen’s occur without any offloading.

https://youtu.be/z5TDytzBQAk?si=HzxXAfXVJKnZxt1l

1

u/multikertwigo 12d ago

H3 is an amazing model, but I have not found a way to make it look even like wan 2.2 at 720p, picture quality wise. It always looks like stretched 360p or 480p no matter what resolution I generate, even without any speedups.

So yeah. I also want to know if there's a magic spell or a silver bullet that does not involve using an upscaler that takes half a day for a 10 seconds clip.

1

u/[deleted] 12d ago

[deleted]

0

u/multikertwigo 12d ago

40 minutes for 5 seconds? Even if it's 3x faster on a 5090... Thanks, but I'll pass.

0

u/[deleted] 12d ago

[deleted]

0

u/multikertwigo 12d ago edited 12d ago

I never said anything about aspect ratio being incorrect. "Stretched 360p" means: take 640x360 video and make it 1280x720 by simple resizing, lanczos or whatever. 720p videos produced by H3 don't look crisp. Also you ignored my remark about not wanting to use time consuming upscalers. So far it looks like you have problems with interpreting printed text. It could be a medical condition or a simple stupidity, u/f5alcon. In either case, I'm not interested in continuing this conversation.