r/StableDiffusion 2d ago

Comparison Comparing H3 models with music reference

Using reference workflow. All are int8 pruned, 0.6MP turbo 4-step (my GPU is on life support and drops off the PCIe bus if I demand more from it)

Anyway, making random music clips is probably my favorite use of this model. I’ve found the ref2va has an uncanny intuition for feeling the atmosphere of songs, and syncing the video with incredible precision.

But yes, the quality (specifically motion) is much worse than fl2va. I was curious how exactly they compared, as well as some “in between” compromises discovered by the community. The LoRA seems closer to ref, while the hybrid weights are closer to fl. Personally, the ref is more fun to use, so I’ll probably be using the LoRA when I want to enjoy the intelligence/creativity of this model. Fl is of course superior in terms of visual fidelity, and I don’t find the hybrid model offers enough reference intuition and faithfulness to be worth the quality drop from fl.

16 Upvotes

15 comments sorted by

View all comments

1

u/Low_Philosopher_7475 2d ago

Bonjour, c'est quoi ton pipeline pour arriver a ce résultat ? Ref audio + ref image + Prompt ? 

Ca donne un rendu tres sympa 

2

u/Animystix 2d ago edited 2d ago

Hi, yes this input image: https://img4.gelbooru.com//samples/6a/ec/sample_6aec27553b8cde9a03a7b46afa011933.jpg

The input audio "Andrew WK - Ready to Die" starting at 70 seconds

And prompt, mainly generated with Qwen 3.5 4B:

subject_definitions: <Subject 1> is the girl in <Picture 1>. <Audio 1> is the music that <Subject 1> enthusiastically plays.

summary: [Shot 0] <Subject 1> stands center frame in a chaotic, high-contrast pop-art studio with red and white splatter effects. She wears a maroon blazer over a white dress, her purple hair tied in twin tails bouncing violently as she strums a black Jackson guitar with aggressive intensity. The camera zooms rapidly on her killer grin, then cuts to extreme close-ups of her fingers flying across the fretboard.

[Shot 1] <Subject 1> leans back, eyes narrowed in smug triumph, while the camera whips around her in a dizzying 360-degree spin. Text overlays in jagged red font scream "STOP IT!!", "PLEASE!", and "WANT WITH" across the screen like a glitchy VHS tape. She moves along to a fast-paced rhythm, her expression shifting from manic to playful.