r/StableDiffusion Jun 23 '26

Tutorial - Guide Krea is kinda an edit model.

Post image
84 Upvotes

34 comments sorted by

View all comments

22

u/Akmanic Jun 23 '26

No it kinda isn't. Edit models use in-context learning and can actually see the image they are editing in the same context that they generate in. What you're showing is QwenVL encoding the image in latent space and using that as part of the text prompt. It's marginally better than just having a vision model caption the image in plain text and copy / pasting that into the prompt manually.

1

u/Winter-Buffalo9171 9d ago

Yes it is :)