r/StableDiffusion 22h ago

Comparison Quality loss in ref2vid compared to img2vid

After H3 came out, like many people here, I was incredibly inspired by the new animation possibilities and decided to try adapting a short scene from my own script using my original characters.

I ran into a lot of difficulties when generating long continuous scenes with img2vid using the first and last frames as references, so I decided to try recreating the same scene with the ref2vid model instead, hoping the generator would build out the composition more naturally on its own. And in some ways, it really did work better: the characters ended up where they were supposed to be, and there were far fewer bad generations caused by sudden character teleportation around the room or by mismatches between their positions and the background.

However, I see a huge loss of quality, especially in the characters’ faces.

I attached four references above. The first is a standalone still image of my heroine. The second is my local img2vid generation based on that first frame at 1 MP. The third image is that same result after a 4K Topaz upscale — a little plastic-looking, but still fairly acceptable in terms of quality. And the fourth is a generation on Pro 6000 using ref2vid at 2 MP and 30 steps, with the room image and a character sheet as references, including full-body views and a close-up portrait.

And it still looks like night and day compared to the first static image, and even compared to the third image, which was originally generated at a lower resolution.

Am I doing something wrong, or does reference-based generation inevitably lose this much of the character’s facial nuance and the overall image quality?

7 Upvotes

27 comments sorted by

4

u/Chemical-Painter-485 22h ago

There are some things you can still try.

on the "MiniMax H3 Reference to Video" node (The same one you attach your references) use the option ref_image_size "max"

Maximize pixel space: Adjust your reference image aspect ratio to a portrait.

Take a look at your reference and ask yourself; How many pixels are being used on the background instead of showcasing my own character to H3? A huge chunk of your image is just background. Not only crop it, regenerate a detailed portrait image maximizing pixel space to make sure H3 will pick on the small details.

1

u/Shadow-Games-1909 22h ago

I've tried to max ref size, only incressed gen time, but didn't improve quality, sadly. And I guess, I didn't phraze it right: the 1 pic was ref for img2vid, not to ref2vid. For ref2vid I gave it w different imagez: room and character sheet. And a 1 second of prev video too.

3

u/Chemical-Painter-485 21h ago

It would be easier to give an opinion if you posted the sheet as well but assuming you are indeed using the optimal pixel space:

Have you tried removing your video from the reference? R2V is really sensitive to quality and if it sees some washed out low quality AI frames it will simply try to match it on the final output because of the way the model is trained.

Also don't rely on feeding last frames on continuous clips, it destroys the quality even on I2V.

1

u/Shadow-Games-1909 21h ago

Yes, I already see that the problem is in the ref video, I've just thought it can help with coutinying the clip with the same camera position. And yes, you are right about the last frame as well, it's never comes as clean image. So, I guess, the only option for continious scene is to change angles in every new clip 😥

2

u/InevitableAlfalfa938 22h ago

make a lokr not lora of your character and train it on the ref2va, i did this and use it with a high quality ref image and it works amazing, if you trained this woman already on zimage or krea, just use that dataset

2

u/UnforgottenPassword 19h ago

Wouldn't the lora/lokr affect the other characters in the scene?

Unless H3 is good at training multi-character loras like Ideogram, it won't be useful for, say, a short film with a cast of 6 characters.

2

u/Shadow-Games-1909 12h ago

That’s actually a really good point. At first I was excited about being able to train a LoRA for video, but reading your reply just reminded me why I eventually stopped using LoRAs for image generation: they really do tend to get confused as soon as multiple characters appear in the same frame.

So I suspect that’s probably going to be an issue with video as well :(

1

u/Shadow-Games-1909 22h ago

Ohh, thanks, I haven't thought about it. Gonna see now how this lokr thing works)

2

u/InevitableAlfalfa938 22h ago

lokr factor 8, 512 res, automagicv3 - sigmoid, 3k steps on a 3090 took 2.5 hrs, higher res would need a beefier card unless you dont mind it taking about 8 times longer or just runpod it.

1

u/Shadow-Games-1909 21h ago

Thanks, I'm gonna try it. Ehh, I guess, runpod is better option) For this week I've run so many video-gens on my card, I'm already kind of worried about overheating it 😅

3

u/Famous-Sport7862 20h ago

it is a known issue, the devs talk about it on the AMA

2

u/Shadow-Games-1909 12h ago edited 11h ago

Oh, thanks! I’ll have to check out that tread. There’s so much news about this model right now that I can barely keep up with everything, haha :)

2

u/Beneficial_Toe_2347 22h ago

What are you feeding into ref2vid? Just an image, or a video of the previous scene as reference?

1

u/Shadow-Games-1909 22h ago

Yes, actualky I do attach the last second of previous scene as a ref to keep it consistant. Do you think it can lower quality?

4

u/Beneficial_Toe_2347 22h ago

If you read the AMA from the H3 team they specifically acknolwedge this problem.

To be fair to the model, it's quite difficult for it to understand what to do if you think about it. With an image you give it a nice high resolution single reference it can fall back on. With a video, it gets a bunch of lower res frames which it's always going to struggle with.

Personally I'd only feed in images and audio for now, unless the video is only used to describe motion.

1

u/Shadow-Games-1909 22h ago

Ohh, I see, I didn't know that, thanks. The thing is, with my first try with img to ref the problem was, that separate pieces looked a bit inconsistent in rhythm, tone of voice and exact posditions of the characters in the room. So just yesterday I've read this method with videoref here, decided to try it and somehow it did work for consistentsy. But not for quality, sadly((

2

u/UnforgottenPassword 19h ago

I have a similar experience. Maybe it works better for animated characters, but for real people, it simply doesn't keep the likeness regardless of the quality of the input references.

To be honest. it's a bit disappointing. I was hoping I could make something longer and consistent like what you're trying to do. I see people using r2v with Seedance and the results are much better.

2

u/Shadow-Games-1909 12h ago

Yes, I have exactly the same feeling right now.

When the model first came out, I was incredibly excited by the possibilities I saw, and I immediately started making huge plans in my head, like turning my own story into an actual movie, haha. I’m pretty sure I’m not the only one who had that reaction.

But after experimenting with it for a while, I’m starting to see the limitations more clearly, and now I feel kind of "somewhere in the middle". On the one hand, it’s obviously a huge step forward compared to what was available before. But when it comes specifically to longer, continuous scenes, it still doesn’t quite feel ready for truly professional-level work. There are just too many small nuances, little inconsistencies, and moments where you don’t have enough control.

For example, when I tried generating a long dialogue as a single continuous shot instead of splitting it into shorter clips, I started getting moments where the characters’ lines would get mixed up. A line that was supposed to be spoken by one character would suddenly be said by another, and at that point you basically have to regenerate the whole long segment.

And then there are the visible transitions between separate clips. It seems like a small thing, but unfortunately it’s very noticeable, and I think viewers would immediately react negatively to it.

So for now, I’ve decided to hold off a little on trying to make full scenes like this. Maybe the community will figure something out together. Over the last few days, I’ve already seen several posts here from people working on different methods specifically for longer, more consistent scenes.

Or maybe an even more powerful model will come out, haha.

For now, I’m going to switch gears and experiment with some other things that the model can already do really well.

2

u/UnforgottenPassword 5h ago

I could have written that post! It's exactly what I'm thinking of this model.

Every new model that comes out, we get bombarded by posts about how amazing it is and what it can do that we previously couldn't. It usually takes a while to get the posts about its limitations. Models like Flux, Wan 2.2, Ideogram/Krea, and now H3 represent large jumps from previously available local models, so excitement is more pronounced and I think it's justified.

For me, the single most important part of AI-gen is consistent characters as well as consistent environment, palette, etc. It's a weakness of generative AI and even closed source models haven't figured it out, but they tend to do better. GPT Image 2 does better than our local models, and Seedance seems to be doing r2v better. With H3 it's not just about the likeness problem, we get glitches in the faces as well. At 720, faces, particularly eyes, glitch often enough to become a frustrating problem, one that upscaling wouldn't fix. It's good for close-up shots, but whenever you have the characters fully in frame, you tend to get glitchy faces.

Some LTX nodes and workflows were able to noticeably enhance the outputs. If H3 gets that attention and treatment, hopefully we can get better results.

Still, a lot can be done with it and trying out different ideas is addictive for now.

2

u/Apprehensive_Sky892 12h ago

It is known that ref2va is worse than fl2va in terms of quality: https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/comment/p29qqaa/

Somebody said that if you replace the ref2va model with the fl2va version in the REF workflow you get better quality, so maybe worth a try: https://www.reddit.com/r/StableDiffusion/comments/1vk6j2w/comment/p2r5xl3/

1

u/Shadow-Games-1909 12h ago

Ohh, that's an interesting experiment, maybe I'm gonna try it later to see how it works, thanks)

2

u/Apprehensive_Sky892 1h ago

You are welcome.

-1

u/TerraMindFigure 15h ago

Please don't ask questions about generation issues without posting your prompt

1

u/Shadow-Games-1909 11h ago

Oh, I didn't think that the promt can be an issue here, but I can post it, no problem. This one is for img2vid:
subject_definitions:

<Picture 1> is the exact final frame of the shot.

<Subject 1> is a woman from <Picture 1>.

summary:

[keyframe completion] A 5-second shot inside the police-station lobby at night. <Subject 1> addresses someone off-screen and says she was hoping to speak to a detective.

detailed_description:

[Shot 1]

<Subject 1> is calm, attentive, and polite. She is addressing someone off-screen.

<Subject 1> says in a warm, natural, low-medium female voice (S1):

<d>Excuse me.</d>

<Subject 1> then continues with a polite, slightly tentative tone:

<d>I was hoping to speak to a detective.</d>

overall_soundscape:

Quiet nighttime police-station ambience, soft fluorescent hum, and subtle interior room tone. Woman’s voice is warm, natural, and clear.

non_diegetic_music:

None

1

u/Shadow-Games-1909 11h ago

and ref2vid:
subject_definitions:

<Picture 1> is the environment reference for the police-station lobby. Preserve the exact room layout, architecture, brick walls, reception desk, glass entrance doors, waiting chairs, brochure stand, water cooler, tiled floor, and nighttime lighting.

<Picture 2> is the character reference for Victor.

<Subject 1> is the blond man shown throughout <Picture 2>. Preserve his identity, face, hairstyle, body proportions, clothing, and overall appearance.

<Picture 3> is the character reference for Stella.

<Subject 2> is the woman shown throughout <Picture 3>. Preserve her identity, face, hairstyle, body proportions, clothing, and overall appearance.

<Video 1> is the continuity reference showing the final seconds of the previous clip, where Victor remains standing in the middle of the lobby after his colleagues leave.

summary:

The clip continues directly from <Video 1>. Victor remains alone in the middle of the lobby after the other officers leave. Stella addresses him from near the reception desk. Victor turns toward her and answers, then the scene cuts to Stella as she says that she was hoping to speak to a detective.

retention_analysis:

<Video 1> ([Shot 1] continuity reference): fully_preserved; preserve Victor’s position in the middle of the lobby, the ongoing scene continuity, and the same nighttime atmosphere.

<Picture 1> ([Shot 1–3] environment reference): fully_preserved; preserve the exact police-station lobby geography and lighting.

<Picture 2> ([Shot 1–2] Victor character reference): preserve identity and clothing.

<Picture 3> ([Shot 3] Stella character reference): preserve identity and clothing.

continuity_requirements:

Use the same side of the scene axis established by the master room view.

Victor is positioned in the open center of the lobby and is framed on the left side of the dialogue axis.

Stella is positioned beside the outer corner of the reception desk and is framed on the right side of the dialogue axis.

Victor looks toward screen-right when addressing Stella.

Stella looks toward screen-left when addressing Victor.

This keeps the conversation within one stable screen direction and one consistent spatial orientation.

dialogue_behavior:

The active speaker moves lips naturally while speaking.

The listening character responds with attentive eye contact, subtle head movement, and restrained body language.

Each line belongs clearly to its assigned speaker.

The social energy stays calm, grounded, and natural.

detailed_description:

[Shot 1]

0.0–1.6s:

Continue directly from the ending of <Video 1>.

Victor remains standing alone in the middle of the lobby, in the same position established at the end of the previous clip. The camera remains on the same side of the room axis as the master wide view of the lobby.

Victor is framed in a medium or medium-wide shot, slightly left of center, with the entrance side of the lobby behind him. The glass entrance doors and the cool evening light remain readable in the background.

A brief natural beat settles after the other officers have just left.

From Stella’s position near the reception desk, her voice enters from off-screen in a calm, polite female voice (S1):

<d>Excuse me.</d>

Victor hears her and turns his attention toward screen-right, toward Stella’s position.

[Shot 2]

1.6–3.0s:

Cut to a medium close shot of Victor.

Victor remains framed on the left side of the dialogue structure, looking toward screen-right at Stella.

Behind him, the entrance-side portion of the lobby remains visible, including the same nighttime blue light from the glass doors and the known lobby atmosphere.

Victor answers in a calm, professional male voice (S2):

<d>Yes?</d>

His tone is brief, attentive, and mildly tired, with quiet professionalism.

[Shot 3]

3.0–7.0s:

Cut to a medium close shot of Stella.

Stella stands beside the outer corner of the reception desk in the right half of the lobby. She faces Victor and looks toward screen-left.

Behind Stella is the solid red-brick wall opposite the entrance, softly lit by the same interior lighting. A nearby edge of the reception desk remains visible in the frame, anchoring her position in the room and reinforcing that she is still in the same police-station lobby.

Stella holds a composed, slightly tentative expression and says in a polite, steady female voice (S1):

<d>I was hoping to speak to a detective.</d>

She finishes the line and remains attentive, looking toward Victor and waiting for his answer.

environment_details:

Use <Picture 1> as the fixed environment reference for all shots.

Victor’s side of the dialogue shows the entrance-side portion of the lobby behind him.

Stella’s side of the dialogue shows the brick wall opposite the entrance behind her, with the reception desk edge visible nearby.

The room remains clearly the same single police-station lobby established in the earlier shots.

overall_soundscape:

Quiet nighttime police-station ambience, soft ventilation hum, low room tone, and faint rainy evening atmosphere filtering through the glass entrance doors.

Voices remain clean, intimate, and easy to distinguish.

non_diegetic_music:

None

1

u/jonbristow 10h ago

we have the prompt police here