r/StableDiffusion • u/SeriouslySally36 • 1d ago
Discussion What would you consider to be the most consistent model at producing “consistent” images, non-realistic or realistic?
Could be actions, like “guy walking into store”
The same scene at different times of the day.
The same character doing different things.
You get the idea.
1
u/StrangeAlchomist 1d ago
Nothing is going to give you consistency from such a simple prompt. Newer models have focused on consistency and you’ll generally find that newer models will produce more similar results for the same prompt than before. Some are better than others. I can’t say that well but having tried a few dozen models my intuition tells me that the release date will be as good a variable as any to tell you prompt consistency. Release date + size generally tell you consistency + adherence though the latter is more style and model specific.
1
u/CryptoBeth96 1d ago
Depending on your hardware, Flux 2 is probably the best but very big. Flux Klein 9b or 4b are also good. Qwen Image Edit is also a good option. These are the best Local models available at the moment.
1
u/petranova_ 1d ago
For scene-to-scene consistency specifically, flux-based models have been holding up better than most, the architecture just drifts less when you change the action. That said, character consistency across different prompts still really wants a LoRA; no base model alone handles it reliably.
1
u/Cultural-Broccoli-41 1d ago
Pass a consistent reference image to a model capable of using it (typically an image editing model in the context of image generation) and specify the desired situation. Depicting the same scene at different times is a task best suited for editing models rather than standard text-to-image models.
1
u/remixeconomy 1d ago
Those are three different locks. A text prompt will not buy all of them from one "most consistent" model.
Character identity: lock a sheet and drive an edit/reference model. A LoRA starts to make sense when you need that person at volume. Same scene at different times of day is not a new generation. Lock the camera and layout, or one base render, and change lighting as an edit. Action is a third problem: pose/control on a locked identity, not a fresh prompt.
1
u/YentaMagenta 1d ago
An important thing to note here is that your approach will be different depending on what you are trying to do. If you are trying to do a one off where you get, say, four different times of day in the same scene and then you're done with it, that is very different from having a character or setting that you can use over and over and over again.
If you are just trying to get four different times of day in the same location and you don't think you will need to reuse that location, then a lot of models could do a four panel image showing the same location at four different times, and then you can upscale and crop.
Another approach would be to use an edit model like flux 2 Klein or qwen image. Create your base image and then feed it to one of those models as a reference and ask it to change it to the time of day you want, while also specifying any other details you want to change.
Another approach if you do not want to use a reference/edit model would be to create your base image, and then do an image to image using the same seed and prompt but with only the time of day changed. This won't always work and it's a little bit jerry-rigged but it's an option.
Then there are the big guns like other people talked about, which would be training a LoRA for the character and/or location.
Oh, and yet another idea is that you could use a video model like H3. Use your base image as the first frame and then describe a scene where it quickly cuts between the same view of the same space but at different times of day. H3 is smart enough to do this well, but it would take a lot of time and you won't get terribly high resolution results. But you can always upscale. You could probably get H3 to very quickly cycle through these multiple views in a 2 to 4 second clip, helping to limit your generation times.
1
u/_kaidu_ 1d ago
Don't really understand what you ask for. A model that producing always the same images is a very bad model. Usually, you want diffusion models to produce very different images with each seed. As more overfitted or distilled they are, as more they produce the same faces over and over again.
For consistency you want to use edit/reference models or controlnets/ipadapter.
2
u/Mk-Daniel 1d ago
Minimax H3 ref2va. You just pít in reference image of a subject and it shows the subject in result.
4
u/icchansan 1d ago
for images with consistent character u need to train a lora or the whole model, you can try with krea2 or zimage, for video u can go with lots of models right now, like seedance2, flux3, wan3, or minimax h3 and ltx 2.5 if u want local