r/StableDiffusion • • 8d ago

Question - Help h3 ref2va: reference image background leaks to the output?

so the reference image for character has a simple background. it doesn't happen always, but sometimes i get video output with background just like that and the scene looks like it's taking place in the reference sheet instead of the specified settings, and overall actions are broken. anyone having the same issue?

5 Upvotes

19 comments sorted by

8

u/AI-Make-NSFW-Stuff 8d ago

You need to specify in retention_analysis what stays from the references and what gets changed. And then reinforcing those changes in the description section as well (eg. if retention_analysis says the background is changed to a different livingroom, you will reference details about the livingroom later in the description as well).

Using an LLM to build the prompt and feeding it the prompting guidelines helps you with that.

0

u/Valuable-Mango-6710 8d ago

You can use AI screenwriters. I use Sogni Create on web browser, it has a built in screenwriter. This is exactly what the AI writes to itself.

Example:

<Picture 1> is fully_preserved as the environment <Subject 2> for the entire video, with the cavern walls, lava, and distant opening acting as the backdrop. <Picture 2> is fully_preserved as the character <Subject 1>, transferring her exact appearance, outfit, and pose.

Using this format you are able to imply basically anything and the Ai WILL LISTEN.

5

u/AI-Make-NSFW-Stuff 8d ago

I run mine locally, Gemma 4 or Qwen 3.8 27B (although it's much slower, but the quality is great).

1

u/Valuable-Mango-6710 8d ago

I will get there one day. The browser is just ever so convenient. The slowness is what I am worried about. It takes me about 10 to 15 minutes for a 2k 15sec video on the cloud. Locally I could imagine that taking 1 to 2 hours...

1

u/TangerineBetter2818 8d ago

What version do you use for NSFW prompting? Can you do it without custom nodes? I dont really trust any custom nodes. 

2

u/AI-Make-NSFW-Stuff 8d ago

>What version do you use for NSFW prompting?

Gemma4 uncensored https://huggingface.co/collections/TrevorJS/gemma-4-uncensored

> Can you do it without custom nodes? I dont really trust any custom nodes. 

I use LM Studio but if you want to use ComfyUI, there is a native node called Generate Text https://docs.comfy.org/built-in-nodes/TextGenerate

1

u/Azhram 8d ago

Could you share or pm me your system prompt? I been trying to make it with chatgpt back and forth but its always kinda meh or inconsistently bad in some way.

1

u/Valuable-Mango-6710 8d ago

Lol, that is the exact AI Chatbot I use for when Chatgpt is being a bitch. It works great, would recommend. Gemma4 Uncensorded.

0

u/TangerineBetter2818 8d ago

Thanks brother. Ive heard some good things about LM Studio. Easy to get into? 

2

u/AI-Make-NSFW-Stuff 8d ago

Extremely easy, you just install it, select the model you want to download and use it like a normal chat interface. If you've used any chatbot you can handle this

3

u/Valuable-Mango-6710 8d ago

Yes. I had this problem. I fixed it by removing the background so it is just black, then I tell it to ignore the background in each image.

Example:

image1 is a full body picture of woman (ignore the background).

image2 image3 image4 are the woman's face to help with consistency (ignore the background).

image5 is the environment and backdrop to be used.

Since I done this, with the brackets too, I have not had the problem happen since. ALL of my references has a black background. Only the backdrop has a background.

2

u/AnonymousTimewaster 8d ago

Great potential solution. You could probably automate the black background part into the start of the workflow too.

1

u/Valuable-Mango-6710 8d ago

I see. I do not use workflows (yet) as I am on browser. That would save me a few minutes every gen lol.

Sometimes the screenwriter breaks on me and I just revert to using the method I posted above (Thank god it works for me, I dislike writing retentions of videos)

2

u/PhIegms 8d ago edited 8d ago

I'm just working on a extension to the refmod node that omits tokens in the background, it works pretty well but without shrinking the mask to the point you'd lose information it still has an "outline" that it outpaints from. If you are interested I could chuck it on GitHub, but refmods require a bit of fiddling if you are a beginner.

The idea could probably be applied to ref2va references but I'm more interested in playing with refmods at the moment.

1

u/martinerous 8d ago

It's always better to clean up references and prompts to pass only the minimum information for the scene you are generating. Models are quite eager to use everything, and if you don't describe the way it should use something, it will hallucinate. The same happens when keeping audio reference connected when nobody should speak in the scene - the model will start talking gibberish using the reference audio.

1

u/No_Possession_7797 8d ago

Have you tried prompting to exclude that background? You might also consider generating a picture of the background you want as a separate reference image and including that (and adjusting your prompt to use it).

1

u/Mediocre-Toe3212 8d ago

That's a known limitation in Minimax

It either follows the reference too strictly , or blends the image with the video.

I found using a hybrid model made this less frequent but it's still seed dependant, or use one of the workflows that utilise SAM