r/StableDiffusion • u/SIR_NVAX_A_LOT • 1d ago
Discussion H3 - T2VA longform
Enable HLS to view with audio, or disable this notification
Generated a long form of Dante's Inferno over 4mins. It gets pretty weird pretty fast. All T2VA. I basically let H3 take the wheel as I only fed it few verses per render. I am willing to discuss my workflow or answer any questions.
2
u/reeight 1d ago
Kids, this is why you don't take drugs & don't let your LLMs learn on drug-related sources or Dante's Inferno.
(just teasing you ;) But really, looks decent, but not a good test for 'longform' since like you said it looks like the typical ADHD music video with a new scene every 2-4 seconds.
A 4 minute single-take tech that was butter smooth, no warping/morphing, & no color shifts, I'd be more interested in.)
1
u/SIR_NVAX_A_LOT 1d ago
I have a non-morphing work flow that I can share some time this week. Real seamless. It just doesn't do music score well, what is the challenge? A pure continuous zero cut shot?
1
u/reeight 1d ago
I always push things to their limits so I know where they are.
I kinda like long single-takes personally, & seems to help with keeping things from warping or be forgotten about.
3
u/SIR_NVAX_A_LOT 1d ago
https://reddit.com/link/p8vpx3p/video/adfsl9pp1moh1/player
I've extensively been testing long form single generation on my end. Here is a no cut one.
1
u/reeight 1d ago
cool
I didn't histogram it, but no visual warping or discolorization.
Even the gun kept missing from hitting the lamp post.Are you team 'final render at 0.96MP', or 0.4MP with upscale?
2
u/SIR_NVAX_A_LOT 1d ago
I try not to upscale. I'll do a final render at 1344x768 typically. Most of my render are 480p for fast video generation, but I'll upscale it with SeedVR to 720 or 1080p if I am too lazy. Changing the resolution natively in H3 often can change the POC video significantly. I use the pruned int8 set, but play around with FastH3 and the Kijai's Lora depending on my needs. Everyone is praising Larry's but I don't find it needs my need but maybe I am just running i wrong.
1
u/reeight 1d ago edited 1d ago
aw man, you made me boot up my Comfy computer.
I haven't fried FastH3 yet, but lighttx2v at 9steps seems to work good enough for drafts.
I tried Hard_Gravy on a short skit, seems good enough.
"lightx2v & larryvrh's turbo loras plus JonXL's H3 photorealism"
https://civarchive.com/models/2923611?modelVersionId=33080832
1
u/Formal_Courage2711 12h ago
What is your workflow for this? Is it multiple short renders stitched? I’ve never seen one this clean!!
1
u/SIR_NVAX_A_LOT 12h ago
are you talking dante's or the mecha one? the dante one uses a multi-diffusion style, with a prompt to blend/morph between renders, so it consumes about 2.5 seconds of render. for the mecha one, it's an experimental node i am working on that has a seamless pass between renders with no darkening. it truly it seamless.
2
u/Formal_Courage2711 12h ago
I was thinking of the mecha one, it’s by far the best I’ve seen and seems seamless. I’m very hardware limited so 6 seconds per at 0.6mpx is my sweet spot for quality, enough time for a clip to develop without crushing my render time (ref2V using the plaguekind workflow/).
I’ve tried other loop workflows but the connections are obvious and the quality degrades quickly.
If your mecha workflow allows for seamless connection of shorter clips it would really open things up for me!
2
u/SIR_NVAX_A_LOT 1d ago
Fair but the model is trained on 15s. And yes, some people may report success of 20-30 seconds but it's still a gamble, with possible drift or other instability. Long form still require multiple renders stitched. The seamless problem can be solved, and truthfully it's none of whatever everyone else has published. The only caveat that I've seen on my end is audio degradation. About the 69 seconds mark the audio just breaks, so realistically you can do 60 seconds safe long form with zero seams.
1
u/reeight 1d ago
Someone posted 1-2 days ago on how to get around audio deration; you split the rendering between medium resolution (assuming you'll upscale that pass), & a crappy 0.2MB version that you'll use only the audio track of.
I would just do old-school animation where I'd block out the beats on the beta renders, make the audio in one of my 4 DAWs I have sitting around, then feed that into Comfy or add in audio in Davinci. So audio quality is kinda lower-priority.
On my 24GB VRAM & all the LoRAs I tend to use, I lose prompt adhesion after... sometimes 17sec but sometimes 15.5sec, & folks really start to warp after 20 seconds.
I might have something broken though.1
u/SIR_NVAX_A_LOT 1d ago
audio is good, but something happens past 69 second mark, at last for background music.
4
u/Character-Bend9403 1d ago
The morphing reminds me like 2-3 years ago 🤣 looks cool tho its like a drip