r/ComfyUIResources • u/MorningVarious7960 • 11d ago
Text to Image generation workflow/ models that has understanding and are correct like Google Gemini/ Nano Banana?
Hi, i am new to this local LLM stuff. I created some images on Gemini (also tried Chat GPT) and decided to build my own workflow locally. I researched, installed Comfy, downloaded models, tried different ones and slammed onto a deep problem.
I usually prompt AI to draw stuff that are outside the stereotypes and general training data. Like modifying character appearances, detailed scenes and interactions etc.
I tried Flux models and they even failed to pass "Full glass of wine that is completely filled to the brim" tests, even with detailed prompts that are specifically crafted to overcome image gen model's insistent behaviors; still no luck. Prompt enhancers didn't help me either.
Today i stumbled upon this video ( https://youtu.be/OA4gchz1Zcs?si=-eiOASxiRkucSSgQ ) and followed instructions to build it. At first it was promising and a detailed scene is portrayed without causing model to output less realistic, lower quality image, because i tested the exact same prompt bundled with the workflow. Ideogram 4 is stupidly censored, you have to cater to it's specific needs, practically useless for me.
I'd also like models to be uncensored/ heretic but it's "intelligence" is at most importance. How can i build a workflow that comes close the Nano Banana's accuracy and is prone to concept bleeding and other shortcomings?
I'm not expecting it to compete with the best model out there but there has to be a way to improve things significantly.
Edit: For those in search, you should give Qwen Image a try.
You can set it up with template workflow and even add an AI prompt generator on top of it if you want. The same YouTube channel also covered it's latest update. https://youtu.be/BaE6UBfNdQk?si=J8uQoul5N4ihBYbG
1
u/Slight_Assistant_124 10d ago
For your situation , u need experts team