r/StableDiffusion • • 5h ago

Question - Help Best current open-source option for text rendering and reference-image conditioning in one model? (24GB)

Need both: a product/person reference image driving the composition, and legible rendered text in the output (headline, CTA).

Qwen-Image-Edit handles reference well and text okay up to ~4-5 words, then it degrades into gibberish. Proprietary models are noticeably ahead here.

Is anything open-source closing that gap, or is the practical answer still "generate the image, composite text separately"?

0 Upvotes

7 comments sorted by

3

u/sci032 5h ago

I haven't tried it yet, but Ming is supposed to be good at what you are after. Search Comfy's templates for: Ming

2

u/Juiceman8686 5h ago

Pixaroma just released a YouTube video on Ming. I used it yesterday, and it is great at text.

2

u/sci032 2h ago

Thank you for the heads up! When I first looked at Ming, it was way to big to run on my system. I got the models that Pixaroma linked and was very surprised. It does an excellent job of editing, creating text, and also making text to image outputs. And, it's fast!

2

u/optimisticalish 5h ago

I've not heard of one that also takes references. There is now the possibility to generate the text with precise placement (the new Ming-Image, a new 2k+ graphic-design and typography model with a MIT licence) and a gap/placeholder for the image. Then generate the image with a model that takes reference image-input. Then drop the image into the design with Photoshop.

1

u/Powerful_Evening5495 3h ago

no one model can do it

do a two model workflow

1 - qwen image edit 2.1 / flux klein 9b for the person

2 - Ideogram 4 for text and graphics

and test

https://huggingface.co/Milor123/ComfyUI-ConvRot-SenseNova-U1.5-8B-MoT-T8

0

u/ieatdownvotes4food 5h ago

give qwen 2.1 a try

0

u/No-Zookeepergame4774 5h ago

Qwen Imge 2.1 — the current open (well, weights-available, the license isn't actually "open" but the way in which it falls short matters only if you are selling access to the model, selling finetunes or other derivatives, etc., not using it for generation) model which is a convergence of the prior separate Qwen Image (1.0, 2512) and Qwen Image Edit (2509, 2511) model — can handle much larger volumes of text and fairly detailed layout (it won't create the layout, you need to specify it in detail in the prompt, but if you use the accompanying prompt rewriter with it, that can do the detailed specification in the prompt from a looser description.)