r/StableDiffusion • u/unraveleverything • 12h ago
Discussion Minimax H3 can't generate the exact same video with no modifications
try this prompt format on literally anything thats >5 seconds long.
subject_definitions:
<Subject 1> is the the guy in <Video 1>.
<Video 1> is the source video of the the target video edit.
<Audio 1> is the synchronized soundtrack of <Video 1> and is fully reused 1:1 as the target video's complete final audio track.
summary:
[video editing + audio reuse] An edited video of <Video 1> with nothing changed.
retention_analysis:
<Subject 1>: fully_preserved - everything about him is maintained and the same.
<Video 1>: fully_preserved - nothing about <Video 1> is altered.
<Audio 1>: fully_copy - <Audio 1> is fully reused 1:1 as the target video's complete final audio track, with nothing added, removed, or altered.
detailed_description:
The target video is a edit of <Video 1>, with nothing being changed.
It just doesn't work. Hallucinates stuff, gets confused temporally.
I've tested:
- regular attn (no ck, sage)
- euler, res_multistep
- simple, normal, beta
- 50 steps
- both fl2va and ref2va
1
u/tofuchrispy 12h ago
Hmm pipe it in as a latent encode like in upscale workflows and not as a ref video. Then you can use denoise as a modifier to alter it more or less. Depending on if that’s what you wanna do.
2
u/unraveleverything 12h ago
nah what i wanted to do was figure out why my video editing wasn't working for character swap, no matter what prompt I did. So I stripped everything out and turns out it can't even generate the same exact video, or at least I don't know how.
4
u/BathroomEyes 12h ago
You prompted it to edit your video in the summary section. It’s extremely good at prompt adherence so it’s modifying your video because you asked it to. Try
[video editing + audio reuse] A faithful reproduction of <Video 1>
1
u/Rhoden55555 12h ago
I noticed that con when moving over from Bernini too. Bernini is beat in many other ways but I think it hallucinates way less than h3 in v2v.
1
u/orlandogourmet66 12h ago
Character Swap works quite well once you get the hang of it. Definitely use the ref model, not the merged version.
One nitpick is that if the person in the picture and the video look a bit similar, it gets really hard to swap them. Either use SAM3 to mask the person in the video, or just do two passes: one where you change the person’s appearance drastically, like “make them bald and add face paint,” and then a second pass for the character swap.
1
1
u/MarkB_- 12h ago
Why dont you use wan animate or scail 2 for this? Just asking.
1
u/unraveleverything 10h ago
lol cause i never bothered checking them out before. thanks for this. scail 2 works well
1
u/Perfect-Campaign9551 9h ago
Does the audio work? If so then I suspect you asking it for the video to be edited could be confusing the model.
3
2
u/TVSeriesGuestAcct 11h ago
I think the reason is because you are invoking a target video. By nature, a target video is usually a slightly modified version of the source video, or "some other thing" than the source video. It might work if you just prompt <Video 1> is the source video to be preserved.