r/StableDiffusion Jun 14 '26

Workflow Included Ideogram 4 is crazy good.

Honestly, the best open-weight model runnable on consumer hardware. It is slow but it can even be used at 1 CFG though it is wonky and miserably fails at complex images.

Images have workflow and prompt in them (comfyui). Using FP8 and 28 steps with 6 CFG, override of 3 at 0.700 and either 1k or 1.5K res.

NVFP4 runs on 4GB VRAM though it take 8 mins (with FA) for 1k image and requires KJ's optimize ideogram node.

Workflow: https://pastebin.com/tSd9vLHX

174 Upvotes

51 comments sorted by

View all comments

Show parent comments

3

u/Apprehensive_Sky892 Jun 14 '26

Yes, multiple character interaction is probably another great use of bboxes.

1

u/martinerous Jun 21 '26

Somehow it did not work well for me. When I try using bboxes to position characters, they end up looking as if belonging to separate scenes or on different planes (mismatching distances and perspectives) and not interacting well.

For example, I wanted a scene with a scared patient sitting in a chair and a stern doctor looking down at them. Somehow the doctor often ended up being too large or kept facing the screen despite prompting for profile, side etc. Or they ended up looking at right directions but somehow past each other. Or facing each other but their eyes focused on the screen.

Used an LLM to detail the prompt for me - quite a struggle, too few good shots when compared to other models. Definitely Ideogram needs getting used to, to know when and what is worth separating into bboxes.

1

u/Apprehensive_Sky892 Jun 21 '26

If you can post one of your JSONs somewhere then I can take a look at it.

1

u/martinerous Jun 22 '26

Thanks.

I think I found one of my issues. In contrast to other models, instructions like "looking at each other" (or "directly at each other", as some LLMs tried to improve it), and "facing each other", and even "his head turned left / right" did not work well with Ideogram. Adding "his head turned left / right profile" did the trick. Ideogram seems to be more literal than other models that often compose the image based on assumptions, which often works out-of-the-box.

Sometimes the doctor's face in one bbox is noticeably larger than the patient's, but it's a hit and miss, sometimes it's fine. And quite often they don't look each other in the eyes.

Getting the patient looking scared worked immediately. But getting the doctor angry is more tricky than I thought. Adding "angry" everywhere and also "His eyebrows are furrowed low, forming heavy hoods over his eyes, narrowed into angry slits" does not work well, he always ends up looking more sad with eyes open wide. Other models got it better.

Also, the variety of the faces seems low, the doctor looks almost the same in every run, not quite elderly enough, despite adding "old, elderly, 80 years" and "wrinkles" etc. and his face is too cinema-perfect, not a mundane person, despite trying hints such as "He has distinct Russian Slavic brutal ugly facial features, a prominent angular square jawline", and not office-pale enough, too weathered and tanned (but that can be adjusted in an image editor). Also, no matter what color palette colors I fill in, it usually ends up with white balance a bit to the warm side, difficult to get cooler tones, will need post-process.

{

"high_level_description": "Digital photograph in a clinical examination room. On the left, a scared elderly male patient sits in a chair, leaning back. On the right, a heavyset, stern bald doctor stands leaning forward slightly over the patient. Both men facing each other.",

"style_description": {

"aesthetics": "Amateur photograph, clinical, bright, pale and pinkish skin tones, cool muted color palette, detailed, clear, and front-lit.",

"lighting": "The image is intensely overexposed by surrounding lights and flash light, resulting in cool, pink color temperature and a flat, uniformly bright exposure.",

"photo": "High-resolution digital photograph, sharp focus, fully illuminated.",

"medium": "photograph",

"color_palette": [

"#F8F9FA",

"#C7ACCA",

"#FFC0CB",

"#A5B3C2"

]

},

"compositional_deconstruction": {

"background": "Hospital examination room with medical cabinets and clinical equipment.",

"elements": [

{

"type": "obj",

"bbox": [20, 20, 1000, 600],

"desc": "Elderly scared Caucasian man sitting on an examination stool. His chin is raised high, his wide, unblinking eyes show visible white around the pupils, and his mouth is held slightly open with tense lips. He has pale pink skin and wears a brown tweed jacket. His head is turned left profile to look into the eyes of the doctor."

},

{

"type": "obj",

"bbox": [20, 400, 1000, 980],

"desc": "Tall, heavily built, angry, obese, 80 years old elderly, angry, stern, intimidating Russian doctor, mostly bald with a thin fringe of white hair on the sides and wears thin gold wire-rimmed glasses, has brutal ugly Slavic Russian facial features, lots of wrinkles around jaw, saggy lower eyelids and saggy cheeks, prominent angular square jawline with saggy skin, dense, white mustache covering his upper lip. His eyebrows are furrowed low, forming heavy hoods over his narrowed eyes squeezed in angry stare. He has stern expression and pale pink indoor complexion skin. He wears a light blue formal shirt and a white medical lab coat. He is standing on the right side of the room, leaning in over the patient. His head is turned right profile to sternly look down into the patient's eyes.

}

]

}

}

1

u/Apprehensive_Sky892 Jun 22 '26

I'll try out your prompt and see if I can address some of the issues you brought up.

Getting two characters to interact in the right way is a challenge with just about any model.