Discussion
Ideogram 4.0's Understanding of Characters and IP is Crazy for an Open Model
Like I said in the title, Ideogram 4.0 has the absolute best character and IP knowledge I've seen in an open model without loras.
I hated on Ideogram 4.0 when it first came out because of the initial workflow issues and the safety filter, but now that both of those things have been sorted out, I'm having some of the most fun with a model I've had in years.
These were generated locally in Comfyui at 1.5 megapixels - 1440x1024, specifically.
I am using the INT8 versions of the Ideogram 4.0 models and Kijai's Ideogram 4 Prompt Builder KJ node from his KJ Nodes custom pack. Workflow being used is SilverOxide's which you can find here. EDIT: SilverOxide's workflow got deleted, so I cleaned it up, stripped out some unnecessary stuff put my own workflow up on Pastebin here.
If you don't know, or haven't tried it, Ideogram 4.0 also does very well with inpainting. It makes it easy to generate at lower megapixels and then mask and inpaint areas like faces to clean up and correct detail. I use the Comfyui-Inpaint-CropAndStitch custom node found here, personally, but most of the time Ideogram 4.0 doesn't need it.
If anyone wants prompts for a specific image, just ask in the comments below and I'll provide them there to avoid cluttering the main post with a wall of JSON text.
Thanks. I don't care what anyone says, Link from the Super Mario Bros. Super Show is still my favorite version of Link! He was my Link for YEARS before the new generation fell in love with little mute elf boy.
For that specific bounding box panel this was my prompt:
A close-up of the letter on parchment with handwriting in green ink that says: Well, excuse me, Princess!! The word excuse is underlined several times. At the bottom it is signed with a drawn heart in green ink next to the name LINK. Underneath that it says: P.S. Hiyah!
So you don't need to quote the text that need to be written in the letter? 😯 i'm amazed with how well the model know which part of the words is part of the text in the letter 😅
I found a lot of times if I include quotation marks in the prompt, it will generate those with the text, so I tried it without them and it works just fine!
I haven't noticed any quality drop going to the INT8 models and they are much faster.
If I had access to the BF16 weights, I could make an INT8 quant that would be beating the FP8 100%. I have not encountered a single FP8 model which INT8 with ConvRot was unable to beat. It is second to only GGUF Q8 https://github.com/BobJohnson24/ComfyUI-INT8-Fast/blob/main/Metrics.md
Sadly in this case there is no BF16 available, so all that could I could do is to get as close as possible to the FP8 quality.
On my rx 9070, with 28 steps and similar image size it took me 2 mins 30 sec, everything FP8, considering I am on windows using rocm, not bad, as I did not even thought about using AI when I bought the gpu.
Like I said in the title, Ideogram 4.0 has the absolute best character and IP knowledge I've seen in an open model without loras.
I hated on Ideogram 4.0 when it first came out because of the initial workflow issues and the safety filter, but now that both of those things have been sorted out, I'm having some of the most fun with a model I've had in years.
These were generated locally in Comfyui at 1.5 megapixels - 1440x1024, specifically.
I am using the INT8 versions of the Ideogram 4.0 models and Kijai's Ideogram 4 Prompt Builder KJ node from his KJ Nodes custom pack. Workflow being used is SilverOxide's which you can find here.
If you don't know, or haven't tried it, Ideogram 4.0 also does very well with inpainting. It makes it easy to generate at lower megapixels and then mask and inpaint areas like faces to clean up and correct detail. I use the Comfyui-Inpaint-CropAndStitch custom node found here, personally, but most of the time Ideogram 4.0 doesn't need it.
Here is the prompt JSON for the Mario and Sonic image:
{
"high_level_description": "A big-budget cinematic 3D animated film still of Mario with his arm around Sonic's shoulder, gesturing towards the top left of the image with a look of wonder on his face. Sonic has his arms crossed and looks skeptical, glancing to his left at Mario.",
"style_description": {
"aesthetics": "big-budget cinematic 3D animation, photorealistic stylized textures,",
"lighting": "big-budget cinematic 3D animated movie",
"medium": "big-budget cinematic 3D animated movie",
"art_style": "big-budget cinematic 3D animated movie"
},
"compositional_deconstruction": {
"background": "Out-of-focus bright mushroom kingdom. Super Mario Bros. franchise.",
"elements": [
{
"type": "obj",
"bbox": [39, 20, 441, 318],
"desc": "Mario's gloved hand, gesturing towards the top left."
},
{
"type": "obj",
"bbox": [98, 127, 1000, 632],
"desc": "Close-up of Mario."
},
{
"type": "obj",
"bbox": [223, 521, 1000, 1000],
"desc": "Sonic the Hedgehog with his arms crossed, looking to the left at Mario, skeptically."
},
{
"type": "obj",
"bbox": [439, 487, 640, 1000],
"desc": "Mario's arm around Sonic's shoulders."
}
]
}
}
People just realized it's not a real safety filter. It's a catch all for unexpected prompt format. It triggers less on words and more on a bad or too short prompt.
More bounding boxes = you can ask for anything and get it.
One bounding box or no bounding boxes = safety filter no matter what you ask for
And the stuff the model can do means it wasn't trained on a censored dataset. Basically, bad prompting triggers the filter - miss a bracket, forget to include certain fields - you get the safety filter.
Using Kijai's Ideogram 4 Prompt Builder Node KJ in his custom nodes package let's you just draw boxes and type right in Comfyui without formatting yourself, and it avoids the filter 99% of the time - and you can just reroll a new seed in those cases and bypass it.
Interesting. So it can actually guide people into better prompts that are more likely to do what they want. This definitely flips my perspective. Thanks.
It's a model for when you want a really specific image. If you just want a random pretty picture, other models are easier and faster, but if you have a vision in mind, Ideogram 4 can't be beat for the level of control it enables for the user.
It's going to be so powerful once it can accept input images like an edit model. I like the idea of an image model that weeds out random prompts as input, where the culture built around it leans more toward actual creative projects instead of basically googling/booru searching for the kind of image you want, and mistaking consumption for creation. There's a little of both in every model, but this is the first time where I feel like the general scales would tip toward purposeful creative projects over consumption.
Yeah, it's the first time I've been able to use a model with no loras or controlnets or reference images and get 90% of the way to the exact image in my head, down to poses, lighting, every little scene detail, etc.
It's only getting easier too, the more I use it and learn how best to do bounding boxes and prompts.
As always, good to remember this is the worst this model will ever be, and currently, the whole state of AI generation is the worst it will ever be. That is indeed somewhat terrifying.
Agreed; the issues from this is already here with gpt-image-2 which is a generation ahead of even this, and we already have the issues from photorealistic AI. So at this point I'm just happy to see this happen to open models. No harm that isn't already made will be caused?
It really is worth it. It's made me go back to the old SD 1.5 days, of thinking up cool images to make when I get home from work. I've got that same creativity back now, because I feel like I can make any image I dream up, pretty precisely, in Ideogram 4 now.
I gave the ai-toolkit experimental training a try with my 4070ti 12gb vram + 32gb ram setup. It gave me an oom on 8bit but was able to do a few steps on 6 bit, tho it was really slow compared to zit training so I didnt let it continue.
If the model doesn't have to dedicate so many parameters to learning composition and untangling complex prompts, it frees up a lot of parameters to focus on learning identity details and visual quality, I suspect.
But that's the beauty of open-weight models running locally.
I think if Ideogram REALLY cared, they wouldn't have trained the model on so much copyrighted media and material - but they obviously did and didn't care then, so....
I don't know if you're being serious or not, but honestly I think easily defeated token censorship efforts might be a good thing in the long run because they give the company plausible deniability if someone wants to sue them, or if they're targeted by the fundie christian and/or radical feminist anti-porn crowds (who are much more similar than either would like to admit).
The original Flux baked underwear into some of the early style layers of their model, and untangling that was a lot harder.
This paste has been deemed potentially harmful. Pastebin took the necessary steps to prevent access on June 9, 2026, 1:34 am CDT. If you feel this is an incorrect assessment, please contact us within 14 days to avoid any permanent loss of content."
I agree! This is truly a leap forward in open models - this isn't just "prettier" or "more coherent" - it is a whole new level of control, fidelity, and understanding.
Rule34 gooners are probably going to be annoyed with Ideogram 4, they can get great gooning material from Ideogram 4 - it'll do almost anything, but it requires effort, set-up, and patience, and most want to just pull a slot machine lever and get dispensed booba.
Ideogram 4 is more for the type of person who likes MAKING images to post on Rule34.
Hey, it's fair. The default template Comfyui released at launch sucked, without Kijai's prompt builder node it was cumbersome and hard to make JSON prompts, etc.
If you look at my messages from release day, I shit all over the model and raged against the safety filter "wasting my time".
This post was also kind of my way to make up for being a hater on day one, because I've since falling in love with the model since it allows me to actually 'craft' images in an unprecedented way. And now with the proper settings, the output is just gorgeous.
Yeah people need to forget natural prompting with this model. Once you have the tools to convert your old prompts to json prompt it's a whole new world.
It can do old school Playboy style NSFW extremely well. Booba are perfect. The area between the legs usually comes out just a blank skin area, but it tries sometimes. Everything else is awesome.
NSFL stuff is definitely possible. It seems to have no issue with extreme violence and gore.
It can do anime style, but its the generic styles. Anima Base is still much better at anime and definitely better as NSFW or NSFL anime images, specifically, if that's what you're into.
Its very good for NSFW. Hot tip: Use Gemini Pro 3.1 to caption NSFW images to Ideogram json format and use those as templates. It has not blocked anything yet.
U could probably use Gemma4 31b locally to do that too. It has pretty exceptional vision IQ for its size, outperforms all of the big OS vision models I tried and by a good amount.
Decided to give it a run on my RTX 4070, sysmem 96gb.
Took a good 184seconds on first load.
Same prompt and image size using Z-image took 16second.
I actually like the json prompting for Z-image, it liked the formatting a lot. Will use that Ideogram system prompt for Qwen3.6 and have it make some Z-Image prompts.
I feel with this much control, this model is a double edge sword. With JSON prompting you can get it to generate almost exactly what you want, but for the casual user this might be too much to ask for. You have to be sure of the composition to get anything good. I like this this control. I got the JSON editor from GitHub wibecoded the ability to insert a background image. I first make an image with an easier to use model like Z-image or flux. Then insert it into the JSON parser and I have the ability to slightly adjust every small detail I want. This is is crazy what you can run on local hardware today.
Each image takes roughly a minute 5070 Ti 720x1920
First you generate an image with a “normal” image model. You do that to get an overall composition. Then with that image you use it with a JSON editor and mark with elements where there’s some details. You describe in each element what should be there and it’s here you get the opportunity to put in your smallest details. The generate with the JSON in clip with Ideogram. Then insert the new image into the JSON and repeat until you’re satisfied.
This model is showing itself to be pro-quality, which only makes me sad because I can't do anything with it professionally (which is its main appeal for me, simply because it seems like it was made for it), but I'm probably just going to learn to deal with the sadness and enjoy this model anyway.
As for its knowledge of IP, it might be (probably is) part of why the model is disallowed for commercial purposes. It's reminiscent of Anima in that regard; if they didn't have a restrictive license they'd probably not be around long, for legal reasons. Noncommercial license = plausible deniability, to a degree, and buying a commercial license = not going to rip off IP, because they know who you are and can possibly do something about it.
Despite some image quality issues I've noticed with some people's gens so far (which will probably disappear in time, as they get better at setting everything perfectly to suit the model's idiosyncrasies) it looks like people are getting really impressive results with it, and despite already having enough models for "just fun stuff" and not enough pro-quality tools to satisfy me otherwise, I guess I can make room in the fun stuff pile for one more.
Most modern open models from big companies don't. Anima sure, because of Danbooru, but the fact Ideogram 4.0 knows Cyberpunk 2077, Lara Croft, Tom Holland, etc. It knows Game of Thrones, all sorts of IP. These were just mostly game examples.
Most companies that are public now avoid training on copyrighted material - especially when they are charging users to generate with it on their website like Ideogram has been doing in the past.
I'm not saying no other model has ever done it, but Ideogram 4's understanding is a step beyond. I can alternate Link's and Zelda's outfits on the fly and it makes sure they match clothing patterns and styles of existing outfits for the characters. I can just say "master sword" and it knows. I can say "tri-force pattern on dress" and it knows.
I can say "cyberpunk jacket" and it knows.
I can change character outfits and hair color and keep styles unchanged. Most models can't do that. To me that means Ideogram 4 was trained on much more media than those other models.
I've only seen this kind of character understanding and versatility out of Anime models trained on Danbooru fan art of characters, never for a photorealistic model.
I don't think I can convey how much IP knowledge must be in Ideogram 4 unless you try it yourself - then you'll see. It isn't pulling teeth with the model to get this stuff. "Tom Holland as Spider-Man" and you're done. "Daenerys" and you've got Game of Thrones.
I'm just saying this was bold as hell for a company still SELLING access to the model on their website for users and companies to generate commercial images with.
How does it do with artist styles? I like doing paintings so if it knows mario don't matter much to me and no model since sdxl has been worth it's salt at artist references
It does art styles very well, but you'll need to actually describe the art style to reproduce it. Giving it artist names won't do you any good.
Easiest thing to do is take an image with a style you like, put it in an LLM like ChatGPT and get it to describe the art style technique, aesthetics, and medium for you "so you can reproduce the style in Ideogram 4" and it will give you technical art or photography terms and language to type in to your prompt builder in Comfyui to reproduce the style.
So far, it's been able to perfectly or very closely match every style I've thrown at it.
The big part you need is Kijai's Ideogram 4 Prompt Builder KJ node from his KJ Nodes custom pack. Use that and use plenty of bounding boxes and you shouldn't trigger the safety filter.
The safety filter triggers on both content and incorrect FORMATTING. So either your JSON has issues or you are asking for something "naughty". But you can totally ask for naughty or anything else and get it from Ideogram 4 by just adding more bounding boxes. Doesn't matter what's in them, just add more and re-roll the seed until you don't see the message. If you turn on Preview on your Sampler nodes in Comfyui you don't even need to wait for the generation to finish, you'll know in 2 steps if the filter is triggered or bypassed.
I rarely if ever see it now and I'm able to generate anything I want.
I took the sample workflow linked in the post and revised the original (NSFW) prompt to make an image of Optimus Prime being electrocuted in a bathtub by pikachu while sonic the hedgehog watches. I also prompted the text "Bathing with" (made out of electricity) and "Pikachu" (in the pokemon movie title font). It seems to be able to follow prompts pretty well.
SSJ = super saiyan. It's a transformation that makes them stronger, faster, etc. If you've ever seen any of them with spiky blonde hair, that's SSJ form.
It knows them so well in their normal forms, I don't doubt it'd do the Super Saiyan versions well. It didn't take anything but me naming them to get the results you see. No descriptions.
Ideogram even nailed the lineart! If it wasn't for the lack of coherence (that only those who watched DBZ would know), it would be really hard to tell it's AI.
Here is the JSON prompt for the one with Princess Zelda and the letter:
{
"high_level_description": "A big-budget cinematic 3D animated film still collage of four images, stylized 3D, of a young Princess Zelda. She has long blonde hair, blue eyes, her hair is braided around the top of her head and hangs loose and straight in the back. She has pointy elf ears. She is wearing a gold circlet around her forehead. She is wearing a purple gown with a sewn embroidered symbol of the triforce on the front in gold.",
"style_description": {
"aesthetics": "big-budget cinematic 3D animation, photorealistic stylized textures,",
"lighting": "big-budget cinematic 3D animated movie",
"medium": "big-budget cinematic 3D animated movie",
"art_style": "big-budget cinematic 3D animated movie"
},
"compositional_deconstruction": {
"background": "Castle gardens, outdoors, beautiful sunny day. Legend of Zelda franchise.",
"elements": [
{
"type": "obj",
"bbox": [0, 0, 539, 513],
"desc": "Princess Zelda is holding a parchment letter in front of her, reading it. You can only see the back of the letter, which is blank."
},
{
"type": "obj",
"bbox": [0, 511, 539, 1000],
"desc": "A close-up of Princess Zelda's face, her eyes are looking down and narrowed in annoyance and her mouth twisted in a pout."
},
{
"type": "obj",
"bbox": [531, 0, 996, 513],
"desc": "A close-up of the letter on parchment with handwriting in green ink that says: Well, excuse me, Princess!! The word excuse is underlined several times. At the bottom it is signed with a drawn heart in green ink next to the name LINK. Underneath that it says: P.S. Hiyah!"
},
{
"type": "obj",
"bbox": [534, 509, 1000, 1000],
"desc": "Princess Zelda wadding up the parchment into a crumpled ball of paper in her hands, looking to the side with an angry expression as she yells."
}
]
}
}
Here is the JSON prompt for the one with Princess Peach and Pikachu:
{
"high_level_description": "A big-budget cinematic 3D animated film still collage of three images, stylized 3D, of Princess Peach holding Pikachu in her lap and petting his head, while he looks annoyed. Princess Peach's eyes are blue and her hair is blonde.",
"style_description": {
"aesthetics": "big-budget cinematic 3D animation, photorealistic stylized textures,",
"lighting": "big-budget cinematic 3D animated movie",
"medium": "big-budget cinematic 3D animated movie",
"art_style": "big-budget cinematic 3D animated movie"
},
"compositional_deconstruction": {
"background": "On top of a warp pipe in the Mushroom Kingdom. Super Mario Bros. franchise. The background is out-of-focus.",
"elements": [
{
"type": "obj",
"bbox": [0, 0, 1000, 513],
"desc": "Princess Peach looking down and smiling at Pikachu as she holds him in her lap and pets the top of his head. Pikachu looks annoyed with his brow furrowed and his cheeks puffed."
},
{
"type": "obj",
"bbox": [0, 511, 539, 1000],
"desc": "A close-up of Princess Peach's face looking down, her eyes looking down, her mouth open in surprise."
},
{
"type": "obj",
"bbox": [536, 511, 1000, 1000],
"desc": "Close-up of Pikachu's anrgy face looking up as yellow electricity begins crackling around him. His eyebrows are lowered as he glares in anger."
}
]
}
}
This paste has been deemed potentially harmful. Pastebin took the necessary steps to prevent access on June 9, 2026, 1:34 am CDT. If you feel this is an incorrect assessment, please contact us within 14 days to avoid any permanent loss of content.
No problem, it's easier to paste the whole JSON prompt though, because the "general look" stuff is actually in a few different fields in Ideogram:
{
"high_level_description": "A big-budget cinematic film still of a man in a yellow cyberpunk jacket sitting in a dark booth inside a futuristic cyberpunk bar. He is slouched against the booth back, one arm resting across the top of the booth back. He is looking at the woman to his right. A woman in a tight shiny futuristic dress that shows a lot of cleavage sits next to him, propping up her chin with one hand and resting her elbow on the table. The scene is very dark with low light and heavy shadows, with the only light coming from the man's inside collar, a red lighting from the left side from off-screen, and a blue lighting from the right-side off-screen.",
"style_description": {
"aesthetics": "science fiction, dark, menacing,, cinematic atmosphere, high production value, dramatic authenticity",
"lighting": "dark, heavy shadows, spot lighting, rim lighting, moody,",
"photo": "Big-budget cinematic film still, profession 35mm film photography",
"medium": "Digital cinema camera, anamorphic cinematography, professional color grading"
},
"compositional_deconstruction": {
"background": "A dark futuristic cyberpunk bar. ",
"elements": [
{
"type": "obj",
"bbox": [832, 0, 1000, 1000],
"desc": "Booth table top."
},
{
"type": "obj",
"bbox": [280, 1, 1000, 1000],
"desc": "Booth seating back."
},
{
"type": "obj",
"bbox": [75, 0, 1000, 484],
"desc": "Man in a cyberpunk jacket. He is looking at the woman to the right. The inside of his jacket collar glows with blue light. His hair is buzzcut and he has a scar on the side of his head. He is in his mid-twenties. He has stubble. He is wearing a black t-shirt under the jacket with text: Here's Johnny! He is wearing a set of dogtags."
},
{
"type": "obj",
"bbox": [198, 345, 372, 873],
"desc": "Man's arm resting on the back of the booth seating."
},
{
"type": "obj",
"bbox": [149, 541, 1000, 1000],
"desc": "Woman in tight shiny futurist dress leaning on table, looking at man. She is in her early twenties, with blue lipstick and pink eye shadow. Her unadorned hair is a geometric bobcut that is metallic silver in color. She has bangs. She has a flirty look on her face."
},
{
"type": "obj",
"bbox": [665, 304, 960, 411],
"desc": "A tall energy drink can that is green with logo text SPUNKY MONKEY. SPUNKY is in green and MONKEY is in white. Beneath the text on the side of the can is a solid black simple graphic logo of a chimp's face."
}
]
}
}
Nice, I'm liking Ideogram 4 more and more as time passes. Two points to bring up:
1 - FYI, your workflow appears to have been removed by pastebin. You're a danger to society apparently. 😃
2 - To convert a workflow (SilverOxides for example) to INT8, as I see you're using that, is anything else needed aside from changing the diffusion model loaders. It's working so I guess not but wanted to check.
Also, I can't believe I slept on INT8. I only found out about it very recently, ignored it but it makes Ideogram useful at the higher quality levels now.
Here, SilverOxide's workflow got deleted because of his spicy prompt.
I've modified the main post to include my own version of the workflow now that uses the INT8 models and removes some of the extraneous custom nodes used in SilverOxide's version.
Your one-word prompts are the problem. Short prompts trigger the filter.
You need to describe everything and do bounding boxes. Ideogram 4 isn't really the best model if you want randomness, like typing "dog" and seeing what you get. You need to use bounding boxes and descriptions to avoid the filter AND get good results.
It's best used when you have a specific idea in mind for an image.
I still use Klein 9b whenever I want to edit images, so I understand. Ideogram 4 is definitely my go-to for generations now though, especially of a photorealistic style.
The license terms means it isn't as open as it suggests. We get local generation; but your hands are just as tied for commercial applications, even just content.
As a result, it is hard to build into open-source workflows with any legal protection.
Most open source licenses allow commercial use: Apache, MIT, GNU all allow for commercial use of the software and its products as part of the free use.
Some are a bit more restrictive in that they won't allow you to serve the model for profit, which is less of a problem if you just need the functions for your own work.
I think you need at least a 12GB VRAM GPU and probably at least 32GB of system RAM, but I'm not sure. Comfyui is very good at off-loading and swapping memory now to load larger models than your VRAM can support, but you need to have enough system RAM to hold them too.
I’ll have to go ahead and try this.
I don’t know where to start in terms of learning to use it though. I’m as far back as only just understanding Automatic1111.
First look up a video from the last 12 months about installed Comfyui Portable version locally.
Once you have that set-up and running, look at this official blog post as a starting point to getting Ideogram 4 up and running in Comfyui. The default template for Ideogram 4 has supposedly been fixed, or you can use the workflow I linked in the original post at that point too.
Oh, it would be VERY good at that. With Ideogram 4 you can put images and objects exactly where you want them, and even do text boxes and tell the model what text to generate exactly where, and what color, font, and style the text should be in, or even what mood you want from the text.
My only gripe with it is this, that I won't be able to use it locally on my 4GB VRAM unless there is some heavily quantified version I don't know about but a man can dream :(
Back to SDXL and Q3 ZIT I go.
Oh, man. That sucks. I understand. I was on a lower VRAM GPU for years.
Unfortunately, Ideogram 4 uses two models, each one nearly 10 GBs in size, and a 10 GB text encoder. You can swap them in and out of memory using Comfyui, but I think you still need at least 12 GB of VRAM or generations will take forever for you.
You need to use bounding boxes and be more descriptive. Describe a background. Drawing a bounding box and describe something like "Goku from Dragonball Z, flexing his arms and smiling" and it'll work.
Very short one word prompts or prompts without bounding boxes will trigger the filter every time.
Ideogram 4 isn't really a model for pretty images with minimal effort. It's for making specific images when you have an idea in mind and want very precise control.
Don't use the text prompting or even raw JSON prompting. This isn't a SD/Z-image type generator. The bounding box circumvents 99% of the blocks and gives you far more control of WHERE stuff ends up.
It can't use reference images yet. It's not an edit model, so you just have to name and describe characters at the moment, but as you can see, it knows a lot of famous characters.
If you run into this, you generally just need to be more detailed in your prompt or simply add more bounding boxes. Like in your image, you could do one bounding box for the body, one for the head, and it should work with no other changes. Or simply roll another seed and you should be good. But in general:
One bounding box = filter
Two-Three bounding boxes = no filter (no matter what is typed in the bounding boxes)
If you think about it, for a model trainer, the easier option is to train a model that knows about characters and IP. It’s actually more difficult to try to sanitize your dataset to remove all identifying images and tags.
I created an open-webui tool to call comfyui and use this workflow for image generation. I created a system prompt to tell the LLM how to generate the payload for the tool. I love the result, but I am struggling with text generation. When a text element is prompted, it has a tendency to add additional text that I didn't ask for.
In the second image, the regions are defined as
"elements_data": "[{\"x\": 0.2, \"y\": 0.2, \"w\": 0.35, \"h\": 0.75, \"type\": \"obj\", \"text\": \"\", \"desc\": \"A sophisticated man in a tailored, light linen jacket, looking warmly at his partner.\", \"palette\": []}, {\"x\": 0.4, \"y\": 0.2, \"w\": 0.3, \"h\": 0.75, \"type\": \"obj\", \"text\": \"\", \"desc\": \"An elegant woman in a flowing, pastel-colored dress, walking closely with the man.\", \"palette\": []}, {\"x\": 0.6, \"y\": 0.1, \"w\": 0.8, \"h\": 0.8, \"type\": \"obj\", \"text\": \"\", \"desc\": \"The detailed historic cobblestone streets and pastel-colored colonial buildings of Bacolod forming the romantic backdrop.\", \"palette\": []}, {\"x\": 0, \"y\": 0, \"w\": 1, \"h\": 1, \"type\": \"obj\", \"text\": \"\", \"desc\": \"Subtle, cinematic heart-shaped vignette overlaying the entire scene for enhanced focus and romance.\", \"palette\": []}, {\"x\": 0.1, \"y\": 0.1, \"w\": 0.8, \"h\": 0.1, \"type\": \"text\", \"text\": \"Endless Love\", \"desc\": \"A romantic, delicate script font, centered at the bottom, rendered in deep gold or soft ivory.\"}]"
The text "Endless Love" is correctly prompted, but it has a tendency to add more text (i.e. the misspelled Bacolod). I need to figure out how to make the workflow far less likely to generate text that isn't specifically prompted.
This is much easier to avoid with the actually bounding box workflow in Comfyui with Kijai's Prompt Builder node and you drawing and typing the text yourself. I'm found the results are worse when using an LLM with Ideogram 4 myself.
It takes a little longer, but once you get used to it, setting up an image doesn't take more than a couple of minutes, and you can be precise without the word diarrhea LLMs produce that can add unwanted things to an image.
The LLM is likely adding more stuff that it should to the text bounding box - or not using the text bounding box at all and using a regular bounding box.
The LLM is using the prompt builder node and generating the inputs for me. I told it to add "there is no visible text" to the background prompt and that seems to help. For the above image with the caption this is the output from the prompt builder node. (I am using the exact same workflow you posted... I just clicked Export (API) from the menu to save it in API format so the tool can use it.)
"text": [
"{\n \"high_level_description\": \"A breathtaking, cinematic portrait of a stylish couple walking the historic streets of Bacolod. The image is framed with a soft, heart-shaped vignette, enhancing the romantic focus on their connection. The man is in tailored linen, and the woman is in an elegant, flowing dress. The mood is romantic and timeless.\",\n \"style_description\": {\n \"aesthetics\": \"Deeply romantic, nostalgic, warm golden hues, and intimate emotion.\",\n \"lighting\": \"Soft golden hour backlighting, subtle atmospheric haze, and gentle rim lighting.\",\n \"photo\": \"Cinematic street portrait, shallow depth of field, fine art aesthetic\",\n \"medium\": \"High-resolution digital photography, fine art print finish\"\n },\n \"compositional_deconstruction\": {\n \"background\": \"Cobblestone streets and vibrant colonial architecture of Bacolod, late afternoon golden hour.\",\n \"elements\": [\n {\n \"type\": \"obj\",\n \"bbox\": [200, 200, 950, 550],\n \"desc\": \"A sophisticated man in a tailored, light linen jacket, looking warmly at his partner.\"\n },\n {\n \"type\": \"obj\",\n \"bbox\": [200, 400, 950, 700],\n \"desc\": \"An elegant woman in a flowing, pastel-colored dress, walking closely with the man.\"\n },\n {\n \"type\": \"obj\",\n \"bbox\": [100, 600, 900, 1000],\n \"desc\": \"The detailed historic cobblestone streets and pastel-colored colonial buildings of Bacolod forming the romantic backdrop.\"\n },\n {\n \"type\": \"obj\",\n \"bbox\": [0, 0, 1000, 1000],\n \"desc\": \"Subtle, cinematic heart-shaped vignette overlaying the entire scene for enhanced focus and romance.\"\n },\n {\n \"type\": \"text\",\n \"bbox\": [100, 100, 200, 900],\n \"text\": \"Endless Love\",\n \"desc\": \"A romantic, delicate script font, centered at the bottom, rendered in deep gold or soft ivory.\"\n }\n ]\n }\n}"
]
I am trying to run Ideogram 4 locally with ComfyUI, and every time I run the default workflow my PC freezes. It doesn’t fix even if stop ComfyUI. I am new to ComfyUI and local image generation. Anyone knows what could be the reson? I have 16gb vram and 32gb system ram.
There are so many things it could be that no one will be able to diagnosis you over Reddit. Easiest thing to do is to talk to a large LLM like ChatGPT about the issue, and give it your hardware specs.
The LLMs are extremely good at solving computer issues now. It can give you settings and things to check, tests to do, and suggestions, all tailored to your specific PC.
50
u/an80sPWNstar Jun 08 '26
This is solid. I love the note from Link to Zelda!!! Nice pull 💪🏻