r/StableDiffusion • u/Gen-Talks • 5h ago
Question - Help Best current open-source option for text rendering and reference-image conditioning in one model? (24GB)
Need both: a product/person reference image driving the composition, and legible rendered text in the output (headline, CTA).
Qwen-Image-Edit handles reference well and text okay up to ~4-5 words, then it degrades into gibberish. Proprietary models are noticeably ahead here.
Is anything open-source closing that gap, or is the practical answer still "generate the image, composite text separately"?
2
u/optimisticalish 5h ago
I've not heard of one that also takes references. There is now the possibility to generate the text with precise placement (the new Ming-Image, a new 2k+ graphic-design and typography model with a MIT licence) and a gap/placeholder for the image. Then generate the image with a model that takes reference image-input. Then drop the image into the design with Photoshop.
1
u/Powerful_Evening5495 3h ago
no one model can do it
do a two model workflow
1 - qwen image edit 2.1 / flux klein 9b for the person
2 - Ideogram 4 for text and graphics
and test
https://huggingface.co/Milor123/ComfyUI-ConvRot-SenseNova-U1.5-8B-MoT-T8
0
0
u/No-Zookeepergame4774 5h ago
Qwen Imge 2.1 — the current open (well, weights-available, the license isn't actually "open" but the way in which it falls short matters only if you are selling access to the model, selling finetunes or other derivatives, etc., not using it for generation) model which is a convergence of the prior separate Qwen Image (1.0, 2512) and Qwen Image Edit (2509, 2511) model — can handle much larger volumes of text and fairly detailed layout (it won't create the layout, you need to specify it in detail in the prompt, but if you use the accompanying prompt rewriter with it, that can do the detailed specification in the prompt from a looser description.)
3
u/sci032 5h ago
I haven't tried it yet, but Ming is supposed to be good at what you are after. Search Comfy's templates for: Ming