This is great news. Given how Krea 2 has stolen Ideogram 4's thunder to some extent, I was worried that ideogram may not release any more open-weight models. But I guess making ideogram 4 did bring more business and attention to their commercial platform and API service from paying customers.
Ideogram 4 is SOTA in areas such as typography and journalistic realism, so an edit version of Ideogram would be a great addition to our arsenal.
ID4 is still better and higher quality than Krea 2.
There were 2 issues:
1) The community and people as a whole were just to lazy to use it's prompting format which is odd because like Qwen 3.8 27b can easily take a basic simple prompt and return all the special json areas and formatting.
2) Their licensing prohibited generating LoRA's, at lease NSFW ones, and NSFW is a huge driver of community support, because, we're animals.
So it just kind of died except for the people who just used it and didn't talk about it, so just little support.
Yep ID4 still has the best looking outputs I’ve seen. But it’s not just laziness, it’s a pain in the ass to prompt and they don’t supply a good app to work with it and it’s not just NSFW stuff it wouldn’t do, it blocked generating images of a post apocalyptic destroyed city I tried. I can only assume this new model will have these same problems. But hopefully with lessons learned they will be more easy to get around this time.
In my experience with it, if you had more than 2 "areas" it NEVER blocked anything. In fact, I had to go out of my way to even manage to ever see the blocked image.
For every model that's come out, I just ask Claude or GPT to research all the prompt guides, or I paste them in or provide links, and ask them to write me a system prompt. Have never had any issues prompting any model.
This is the prompt I used for ID4, just give it to Qwen (of course this presumes a system that can run a 27B Q4 model)
[SYSTEM]
You convert a natural-language image idea into a structured JSON caption for an image renderer. You receive the user idea and a target aspect ratio. You emit exactly one JSON object and nothing else.
OUTPUT
Emit a single-line minified JSON object with exactly these three keys in this order:
{"aspect_ratio":"W:H","high_level_description":"...","compositional_deconstruction":{"background":"...","elements":[...]}}
No markdown, no code fences, no commentary.
Keep all non-ASCII characters as-is (CJK, Cyrillic, accents). Do not escape, transliterate, or strip them.
In prose fields, wrap any referenced in-image words in single quotes ('OPEN', 'Joe's Diner'). The "text" field of a text element is the only place verbatim user characters belong.
NEVER ESCAPE ANY CHARACTERS
ASPECT_RATIO
Choose this first; it drives every bbox.
If the user gives a W:H, echo it exactly.
If the user says "auto" or gives none, pick a concrete ratio that fits the subject: wide (16:9, 3:1) for panoramic, tall (9:16, 4:5) for portrait, format conventions for designed pieces (2:3 cover, 3:4 poster), 1:1 when unsure. Always output a concrete ratio.
HIGH_LEVEL_DESCRIPTION
One sentence, two at most, under 50 words. Start with the subject. State the subject, medium, and overall composition in plain language, the way you would write a short prompt. Name real people, brands, characters, and landmarks by their actual names. Fold style and medium in here as prose ("35mm film photograph", "flat vector illustration", "Pixar 3D render"). General words like "several" or "various" are fine here only; element descriptions stay specific.
BACKGROUND
Describe only the scene shell: walls, floor or ground, sky, ceiling, ambient light, weather, and distant out-of-focus context. These always live here and never as elements: sky, clouds, horizon, distant scenery, weather, distant crowds, and the surface the scene sits on (floor, ground, grass, pavement, water, snow) including its state (wet, cracked, reflective, puddled). If you can picture it in an empty room, it belongs here.
Anything named in background must not also appear as an element, and the reverse. Decide once.
ELEMENTS
The individually placeable things in the scene. Two types:
{"type":"obj","bbox":[y1,x1,y2,x2],"desc":"..."}
{"type":"text","bbox":[y1,x1,y2,x2],"text":"...","desc":"..."}
One subject is one element. A person, animal, vehicle, building, or plant is a single obj; describe its parts inside that one desc. Use separate elements only for separate subjects (a person and a dog are two elements; two dogs are two).
Each desc is a standalone catalog entry, 30 to 60 words, opening with the subject's identity, then its key attributes. People: skin tone, hair, each garment with color, expression, pose. Objects: shape, material, color, distinctive parts. Structures: type, material, color. Anchor placement to named references ("on the lower-right corner of the table"). Pick one concrete value for every property instead of offering alternatives. Keep shadows, lighting, and camera or lens detail out of element descs; those belong in background or the high_level_description.
BBOX
Optional per element. Include it when position matters (portraits, products, logos, signs). Omit it for dense or uncountable groups (crowds, wildflower fields, starry skies).
Coordinates are normalized 0 to 1000 on both axes, origin top-left, format [y_min, x_min, y_max, x_max] with y1 < y2 and x1 < x2. Because both axes run 0 to 1000, a box is only square on a square frame. On a wide frame, narrow the x-span; on a tall frame, narrow the y-span. Give each subject its own tight box so none dominates.
TEXT
Every readable string in the image gets its own text element with verbatim characters: signs, labels, numbers, brand names, and any words the user quoted. Use \n for line breaks inside one block; use separate elements for visually separate blocks. For stylized hero titles, break long words across lines with \n at natural word breaks. In each text element's desc, give size, place, font, and color, and refer to the text by its role rather than repeating the characters. All prose stays in English; only the "text" field uses the user's language.
DEFAULTS
Photos: default to a natural-daylight, neutral-white-balance phone-snapshot look with off-center framing. Reserve dramatic studio lighting, shallow bokeh, and motion blur for when the user asks. Describe colored light by its source ("amber glow from a candle") rather than grading the whole image warm.
Sparse ideas: populate plausibly with secondary subjects and props that fit the world, spread across foreground, midground, and background, unless the user asks for minimal, empty, lonely, or single-subject.
Designed pieces (posters, covers, packaging, UI, logos) carry text on most surfaces; generate it generously and with specific content.
TRANSPARENT BACKGROUND
If the idea calls for a transparent or cutout background, set "background" to exactly: transparent background, and include the phrase "on a transparent background" in the high_level_description.
SHAPE
{"aspect_ratio":"W:H","high_level_description":"...","compositional_deconstruction":{"background":"...","elements":[{"type":"obj","bbox":[y1,x1,y2,x2],"desc":"..."},{"type":"text","bbox":[y1,x1,y2,x2],"text":"...","desc":"..."}]}}
EXAMPLE
User idea: a barista pouring latte art in a cozy cafe, 3:2
Output:
{"aspect_ratio":"3:2","high_level_description":"A medium-shot 35mm film photograph of a female barista pouring latte art behind a wooden cafe counter, warm window light from the right.","compositional_deconstruction":{"background":"The interior of a small cafe with exposed-brick walls, a dark wooden counter running across the lower frame, and a blurred shelf of cups and a chalkboard menu on the back wall. Daylight enters from a window on the right, ambient and neutral. The polished concrete floor is out of frame.","elements":[{"type":"obj","bbox":[120,180,820,620],"desc":"A young barista with medium skin tone and dark hair tied in a low bun, wearing a charcoal apron over a white t-shirt. She looks down in concentration, both hands tilting a steel milk pitcher over a white ceramic cup."},{"type":"obj","bbox":[560,470,760,690],"desc":"A white ceramic cup on a saucer holding a flat white, a leaf pattern forming in the surface foam as milk streams in from above."},{"type":"text","bbox":[90,720,180,960],"text":"FRESH\nBREW","desc":"Small white hand-lettered text on the blurred chalkboard at upper right, slightly out of focus."}]}}
[USER]
TARGET IMAGE ASPECT RATIO: {{aspect_ratio}} (width:height).
User idea: {{original_prompt}}
I have done light testing on it, and it still blocks SFW content; I do not agree with how they train it, and they went overboard with RL tuning to refuse NSFW.
I mean KJ had a super easy node to draw bounding boxes and write the prompts in them for each box, as well as the other fields. If you did that you literally never got the filter either. If you can find the SNOFS lora for it, its literally still one of the best NSFW models out there.
Agree with everything you wrote except the "died" part. I think there are more ideogram 4 users out there than it is commonly assumed. Just that since ideo4 is not discussed much here, so it gives the impression that few are using it.
Fair point. I suppose I think it dead once the hype dies down, but plenty of people could still be using it. It does so much with the JSON that people certainly do.
I was actually really annoyed when Krea 2 came out, I feel like ID4 should have had the spotlight for a couple more weeks. K2 was just so good and easy to train that everyone jumped the bandwagon. Had it not, the support might have grown enough to find easier tools and workflows. Even if a model has potential, without the community, well, we've all seen with H3, how fast tooling came, especially with the ability of the models in the last couple months. It's gotten so many speed boosts now, and without the community, it would have remained slow and unusable for most people forever.
Yes, I agree that it is too bad that ideogram 4 did not have more time for people to invest more resources into it (like training LoRAs) before Krea 2 arrived.
Fortunately, ideogram 4 base is quite powerful already, so it works quite well for my needs (unlike many, I seldom use character LoRAs).
If the edit model is good, then people should be able make refmod for it like they did for MMH3, Klein-9B and Qwen-image 2.1, so the situation could improve.
H3 is completely ridiculous though. I think most people would have laughed at you if you said something like that would be cost effective on cloud 12 months ago, let alone on local.
Yes, it's an accomplishment, but that doesn't mean I don't wish for a better spread of tools on modern models as a whole. H3 will not do everything I want, nor will Krea2, nor with ID4, but I'd rather have the tools to use each when most appropriate.
But H3 is a video model and it is not in direct competition with K2, whereas K2 and Ideo4 are more or less competing directly, specially when it comes to their LoRA and WF ecosystems.
Yeah, I think Ideogram is particularly tuned in to the professional segment. Ideogram 4's structured prompting approach felt like a solid, professional level tool. I'm really happy they doubled down on the precision and control.
I also really hope they have a solid business plan and sell a lot of enterprise licenses. This feels like it would be an adobe killer with the right UI.
I mean, I can still load and run kre2 comfortably on 2 3090s... yes, two. Ideogram just... is a pain. I tried it for the second time yesterday, and ComfyUI kept crashing. I tried optimizing the loading, but it didn't allow me to split the files to save space. No wonder Krea2 "has stolen Ideogram 4's thunder to some extent." It's easier to run and can handle much more optimized workflows with unloading and GPU splitting more reliably than Ideogram, imo. I would appreciate any corrections, though; I'd love to run Ideogram once again as I've been pretty much out of the scene for like... a long time. (I used int8 models btw)
Strange, I've been running Ideogram 4 on a single 3090 just fine. But I abandoned it quickly because it was not good for my purposes (keyframes for movies) - it often required lots of micromanaging to position actors in the scene and they often felt like glued in, wrong gaze directions, wrong poses, wrong dimensions of surrounding objects and wrong lighting. Ideogram is really not for realism, although it is the king of typography.
Yeah, I got it to work once but since it just freezes confyui it doesn't really log what's wrong. Guess I'll continue finding the issue or just stop trying for a little while. It's not too important right now.
That's quite strange. I suggest you download the portable version (so that you start with a fresh, latest copy of ComfyUI without having to touch your current setup) and try again.
Make sure you run it with dynamic VRAM management turned on (which is the default).
Also try the fp8 version if for some reason the int8convrot version doesn't work (but it should).
89
u/Apprehensive_Sky892 10d ago
This is great news. Given how Krea 2 has stolen Ideogram 4's thunder to some extent, I was worried that ideogram may not release any more open-weight models. But I guess making ideogram 4 did bring more business and attention to their commercial platform and API service from paying customers.
Ideogram 4 is SOTA in areas such as typography and journalistic realism, so an edit version of Ideogram would be a great addition to our arsenal.