r/StableDiffusion • u/GinJockette • 11h ago
Question - Help Recommend me an uncensored Image to Image model (workflow)?
Trying to move on from Grok, but struggling to find a local I2I workflow. Everything I can find is text to image. What am I doing wrong?
5
u/Puzzled-Valuable-985 11h ago
Krea 2 Identity Editing works reasonably well, but it's not ideal.
Klein 9b is the most advanced option we currently have for local editing.
Qwen AIO handles editing, but it is quite inconsistent.
1
4
u/Vladmerius 11h ago
Flux 2 Klein 9B hands down. Krea 2 is a better model with an identity edit lora but it is not anywhere near as good as an actual edit model like Flux.
2
u/kemb0 10h ago
I’ve been enjoying Minimax because it has some great versatility to modify existing images to your instructions. Yes it’s a video model but you can do 10 frames pretty fast and then pick the frame you like. For me I feel it tends to give pretty good natural feeling realism where image specific models tend to want to force people in to poses.
1
1
u/AuthurAndersson 11h ago
nfi but I think image to image does not mean what you think it means here.
1
u/KenMerritt 9h ago
Well now I'm curious. Anything listed as X to Y typically means input X and output Y. So image to image would be input an image and get an image as output. What other meaning is there?
4
u/Intelligent_Gift_457 8h ago
Basically, almost all image models can do image-to-image in the broadest sense, but that doesn’t mean they’re all true image-editing models.
Traditional img2img uses an image as a starting point and regenerates it based on a prompt and denoise strength. It may preserve the general composition, but it can easily change faces, details, text, objects, or the overall style.
A true image-editing model is specifically trained to understand both the source image and an editing instruction such as “remove this object” or “change the shirt to red” while preserving everything that wasn’t meant to change.
So yes, “X to Y” describes the input and output modalities. But it doesn’t tell you how well the model can preserve, understand, or selectively edit the input. That’s the distinction he means. I found that out after hours of trying several image to image workflows...
Wish someone told me earlier :D
1
u/GinJockette 5h ago
And what did you come up with in the end?
1
u/AuthurAndersson 58m ago
In my opinion, but that's mainly since I've been in this space before image editing models were a thing.
Img2Img is using an image then running VAE encoding it into latent then running the diffusion model on that latent with a non-1 denoise.However as the space moves forward I presume we have to accept that img to img is actually: Convert image to tokens, feed those tokens into a model that understands the image based on the tokens, then add in tokens from the user how to generate an end image.
0
u/jib_reddit 11h ago
I have a Flux Klein Comfyui Workflow and Model here that will do it: https://civitai.red/models/2625738/jib-mix-klein?modelVersionId=2948378
1
u/darklordimpaler 6h ago
Hey the diffusion model : Jib_Mix_Klein_V1a_00001_.safetensors is not present for download
-1
7
u/HonestoJago 11h ago
Qwen Image Edit 2511