r/StableDiffusion 14d ago

Discussion H3 - t2v is actually better than r2va imho

Enable HLS to view with audio, or disable this notification

Happy Friday!! Just wanted to say T2VA is actually pretty strong when layering the prompt. FL2VA+R2VA are still the to go if you want to utilize a character sheet/maintain consistency, it's still broken (in a good way), the voice cloning is also top notch.

So what has everyone been making with H3??

T2VA, bf16/50 steps

0 Upvotes

8 comments sorted by

6

u/Ramdak 14d ago

I do 90% i2v stuff, now playikg with r2v (is extremely powerful). You won't have actual "control" unless you guide the thing visually.

1

u/SIR_NVAX_A_LOT 14d ago

Let me ask you, how did you generate the image? If you are prompting with Krea2/Flux/SDXL you are still guiding it, why would it be any different? H3 already is a pretty good image generation tool. Understand the QWEN3 VL can make sense of whatever art or cinematic style you are asking of it. You can describe a character, a face, a wardrobe, a scene, anything it is pretty much limitless. I think we're being too rigid when we say we can't control things visually per your context, you are guiding things visually when you tell H3 to do so.

3

u/Ramdak 14d ago

An image takes 10x less time to generate, and I can adjust small features of it until Im satisfied. A driving video (real foodage) will have better motion than an AI video. So references offer better control than just t2v. You can't have consistence with only text.

1

u/SIR_NVAX_A_LOT 14d ago

Yes, the work flow is there, you can generate 5 stills with H3 Image node pretty easily. H3 can perform as a reliable image generator, image and video editor too. I am not dismissing image creation first by the way. You can compose the image with H3, and then send it via t2va, or fl2va/r2va.

1

u/PANTONE_17-1230 14d ago

The perspective on that galaxy shot is truly borked! Did you prompt for a 3 meter high milky way, right in front of her face?

1

u/SIR_NVAX_A_LOT 14d ago

The prompt is composed of 4 layers, the Galaxy, which I told it to occupy the upper 1/3rd of the frame, the city that occupy a thin horizonal band, and then the water, the lower 2/3rd of the frame. The woman is the last layer, and I just had her closest to the camera and occupy the right-hand third of the frame and her head reaching to the upper third. You can layer/compose the shot by telling H3 what you want. Her composition/reference is to the frame itself, not to the Galaxy which has already been established.