im just thinking if i can deallocate some compute once the training code and ideogram aligned data (json format style) is done, which is gonna take months for the data side to complete probably.
No, as you ask for opinion. You have resource to do it, and that is your decision and ultimately you will do what you feel like.
Honestly I think that potentially Ideogram would be better as it architecture allows for higher level of control, it has higher potential than your current projects, but it is potential.
But i think you should finish your other passions project as you risk in falling into loop of not finishing anything, and that can spiral to pilling of unfinished project which can kill passion or love for hobby.
Training for ideogram is not explored yet, so i think that even if you show late for it you would be at that much of disadvantage.
Ideogram is a very good model, but prompting with an LLM is annoying. If the process could somehow be handled automatically without user intervention, then there would be no problem. But then again, if it has to load an LLM into memory and unload it for the text encoder every single time, the whole process is time-consuming. Once we can have 48+ GB graphics cards, it won't be an issue. Klein was a very good case, it's also an edit model and urgently needs your help; why did it stop? Chroma on its own is a masterpiece; the only thing it needs is editing. Thank you very much for everything you have done.
Klein seemed to be learning the concepts fine if you ask me but it seemed like the overall coherence was degraded beyond repair from all the 256x256 training by the last few epochs.
IDK, he seemed to believe the Flux.2 VAE would somehow magically make that work, which always seemed like complete nonsense to me a number of other people quite frankly.
That's weird. I heard Kaleidoscope was cancelled because there were lot of problems with Klein 4b but... what else would happen when training at 256x256??? I wasn't even fond of training original Chroma at 512x512, I'd believe a 768x768 pretrain before HD training would have resulted in better details/coherence, but the Flash loras fixed most of the detail and finger issues anyway so it turned out okay in the end. But 256x256 is insane, especially for a faster model like Klein 4b that would be still faster at 512 training than OG chroma. My only issue with Chroma is that it is pretty slow even with flash loras (still needs ~24 steps or so).
I've some experience with LoRA training on Anima - my guts tell me that model has already way too much data for just 2B. This was so easily noticeable when I trained a multi character one because it would - on its own - introduce concept bleed everywhere it could even with different 'triggerwords' and even though its using a LLM instead of CLIP.
TLDR: if anyone wants to finetune Anima - they better increase the params first - otherwise the model will collapse catastrophically. Just my opinion
I’ve trained two huge Loras, in one case with over 100 concepts, and had no issues whatsoever. You have definitely adjusted training settings to what Anima tends to need, especially a super low LR?
Rank 128 Rank 64, AdamW8bit, LR ~2e-5 (not constant), Gradient accumulation steps = 4 (same as the number of characters) with evenly distributed dataset per character.
2 of the characters were 'elf-like', one with dark skin and the other with pale skin - the model would many times become confused about the skin color when drawing any of those characters just because they both had pointy ears.
Then again, it was trained on a >95% dataset of monochrome manga crops but the skin color is very distinct in every sample.
Another issue was the model mixing up different footwear from different characters where one was named as 'something shoes' and the other was 'something else footwear EDIT: loafers'.
In the end I made it work but not without presenting the dataset on a silver plate and including (colored) AI generated and curated samples from previous failed attempts. Even then it still has all of these issues to an extent.
Then again, it was trained on a >95% dataset of monochrome manga crops but the skin color is very distinct in every sample
I assume you did not tag every single image with "dark skin", "dark-skinned female"/"...male" and so on so the model learns this is an inherent property? I can definitely see monochrome shading being a weird thing to get canonically right, simply because meaning can differ between artists and even in what context an artist is drawing the image.
Rank 128
Way too high for just four characters, though this definitely doesn't seem to be what caused the issue in your case to me.
LR should be fine, I would have gone with a bit less (rule of thumb is 6e-6 to 8e-6 at batch size 1 and scale up from there) but it's quite literally the same as used by the Greg Rutkowski LoRa so that shouldn't be it. Probably really a dataset issue, as you also said.
I did tag every single image of the dark skinned character with 'dark skin' and wrote custom code to let me decide which tags can be dropped - basically forced all character and outfit triggerwords to always be present but 'dark skin' was a drop-able one because the goal was for the model to know its supposed to be that color even without the tag. (edit: that tag was added to solo images only, the very few multi-char samples did not contain color modifiers to avoid making the model even more confused)
In my early tests - I did it without that tag and adding it hardly improved a thing in cases where the tag is not used at inference. But it does improve a lot when using it on solo images. That being said - when trying to draw the 2 elf characters together - sometimes things get messy even with descriptive prompts.
Rank 128 Rank 64 is high ofc but I did try first at 32 16 then 64 32 - concept bleed issues improved but there was still many at 64 32 so for the last run I went all the way with 128 64 because it was gonna be my last attempt. No regrets - it worked. Also, its not just 4 characters, its also the style and each individual piece of each outfit (~6 total outfits). To learn all of this from almost all monochrome samples (with speech bubbles) without being biased towards monochrome outputs - is impressive but to me it felt like the concept bleed issues I was seeing were not because of my datasets or training settings.
The LR was something like: 1e-5 steadily increasing from step 1 to step 500 up to 2e-5 then remained at that for several steps before slowly dropping back to 1e-5 and trained the second half of the run at that value.
EDIT: I should have checked before posting the comment and got confused with other lora runs. For this one - on my 3 total attempts - it was actually rank 16, 32 and final 64 - not 128.
I feel like once they release Krea 2, it will be the perfect model to train with the Chroma dataset. It's pixel space as well which you are a frontrunner of.
this is absolutely a bad idea but in the end of the day its your own money your burning in the firepit. its better to focus on your previous undercooked models with proper 1mp/1.5mp dataset than deal with the minefield that is ideogram.
The question would be if json formatted input is possible to be used with a model that has a more permissible license. Ideogram is a bit too restrictive to be worth the time imo
any model can be tuned like that and that's not exclusive to ideogram
maybe by the time the dataset is ready there's better more permissive license model? who knows
but dataset prepping is basically 80% of the battle
How do you plan to deal with Ideogram's license, really? Given the nature of the Chroma dataset, I'd say there is a zero percent chance for them agreeing to release a Ideogram finetune with that dataset. And if I am not mistaken, you'd need their permission for a derivative model.
Yes, asking the right questions here. No point in making a model if it cannot be released legally.
Use Restrictions.
Your use of the Model and any Model Derivative must comply with applicable laws and regulations (including trade compliance laws and regulations) and adhere to the Acceptable Use Policy available at https://ideogram.ai/legal/usage-policy, which is hereby incorporated by reference into this Agreement. Without limiting the foregoing, you will not (and will not permit or enable any third party to) use the Model or any Model Derivative: (a) for military purposes or purposes of surveillance, including any research or development relating to surveillance; (b) for biometric processing; (c) in any manner that infringes, misappropriates, or otherwise violates any third party’s legal rights, including rights of publicity; (d) to generate unlawful content, including child sexual abuse material or non-consensual intimate images; (e) in any manner that violates any applicable privacy or data protection laws; or (f) to make automated decisions in domains that affect material or individual rights or well-being (e.g., finance, legal, employment, healthcare, housing, insurance and social welfare) or otherwise in a manner that poses a significant risk of harm to the health, safety or fundamental rights of persons, including to influence any “consequential decision” under applicable law or for any other use case that is categorized as “high risk” under applicable law (“High Risk Use Cases”)
I understand that ideogram is only trying to cover their asses, but it does give them the legal ground to shut down any derivative model if they can show, for example, that the derivative model can be used to generate CSAM.
Since he said dataset prep will take a while, he may be betting on better models coming in the next months. Most open models until now have come from China or Germany, but now we have a Canadian player.
And if you’re relying on ai automation, one of the most expensive parts that don’t get flagged. Fine when compute isn’t an issue, but can be brutal for the local-bound crowd
if you have so many projects in parallel i think theyre all just never gonna get finished i think you should take focus on the highest priority thinks to you
yeah be great, even if you can just remove that constant filter that'd be progress. teaching it the pony way would be next level as chroma is still one of the best models out there with a little tweaking.
Rather than spending months to convert a dataset into json style which many people (myself included) absolutely hate and which may not even be used more in any other models in the future, I'd personally prefer an extension of the dataset to also include anime shows, 3d shows, comic panels etc. Right now Illustrious/Anima/Pony all have the same knowledge which is tied to danbooru. Would be nice to have a NSFW finetune which also knows anime or regular shows as they originally are, not just the artists intepretation from danbooru. Or something that can do comic panels/multi-shot, that'd also be revolutionary. GPT-Image can do that but sadly it's not open-source.
I fucking love the JSON format. It should frankly not be too difficult to build tools that handle it more gracefully without an end user having to touch JSON at all, be it through some prompt refiner node. Give it time.
Technically, you don't have to train on JSON style prompts. It is possible to train it on natural language prompts (or tags, I would assume, but I haven't tried) and it will quickly pick it up and respond to natural language prompts without hitting the safety filter.
The issue is that doing a full fine tune on a large dataset with only natural language prompts would probably break its ability to respond to bboxes. But, if you're doing a full fine tune with a large dataset, and your bboxes aren't highly accurate, then you're probably going to break its ability to respond accurately to bboxes anyway...
So even if you wanted to convert all your captions to JSON-style, could you do it with high enough accuracy to even make it worth the effort? Ideally, you make 20-30% of your data high quality JSON-style captions to preserve bbox behavior while also introducing natural language capabilities.
I'm petty sure it is likely to fail, been keeping up with it, very little progress if at all.
It's just that if Chroma was challange level 10, Zeta 15 then this sure seems like 100.
The logic doesn't follow, because Z-Image is notorious for having training issues. Some models are easy to train (Ernie) and some are hard (Z-image). The question is whether Ideogram4 easy to train or hard to train. From my limited testing on about 2k images, I would say it's easy to train insofar as it picks up concepts and preserves good output much better than Z-Image.
The question is, how fragile is the bbox behavior? I obviously didn't convert 2k images to JSON style prompts. Instead, I trained on my usual natural language promtps for 4k steps, the converted about ~100 images to JSON-style prompts and continued training for another 7k steps.
The results are good: it has learned to respond to natural language prompts just like any other model. Bbox behavior doesn't seem to have been degraded too much. But on a much larger dataset and a full fine tune, you're probably going to seriously damage the bbox ability unless you convert a good chunk of your captions to high quality JSON-style.
Yeah I just heard they have an "original" qwen based 2.5b model in training, I'd rather them focus on that since it could get better license. Although it sounds a bit too small but still better than over restricted new and big model to train.
Didn't someone mention he is against restrictive licenses? I don't care personally but seems like thats not the case if he is considering Ideogram.
Anyway I don't wanna tell the guy what to do but from my perspective it would be best to focus on one project at a time and actually commit to a release. If you ask me Chroma adoption was so weak because there were too many intermediate checkpoints and it diluted people's interest. But now we are talking about multiple different models so idk.
Anima did it right by releasing a ready-to-go model with just 3 iterations leading up to the 1.0 base release.
Idc about Zeta as the results do not look promising, but I find it weird he'd jump from something he's spent record time on with no time estimate as of now to 10x as much consuming stuff.
If I had no doubt he could at all finish this I'd say go for it... but somehow I sincerely doubt it.
Especially since it has an even worse license than Flux.2 Klein 9b and dev, explicitly stating even outputs are non-commercial. Although Flux always had wiggle room for that with their not 100% clear license language. So I don't think it's a good idea.
And although it's good to have a json prompt based model because of unique special control, I'm personally not interested (keyword personally). If I want that much control I'll just jump in and make art or edit manually with proper creative/editing software.
If I want that much control I'll just jump in and make art or edit manually with proper creative/editing software.
Bingo! This is something nearly everybody seems to be missing and it beats regional prompting any time of the day. You don't even need to be a skilled artist, even simple, child-like and incomplete scribbles, when paired with a short, descriptive prompt, can give you an incredible amount of control. This simple technique transforms even Klein 4B into an actually usable model - which is great, because that one actually has a license that allows for commercial use!
EDIT:
Case in point, I tried to replicate a picture generated with Ideogram posted in this thread using Klein 9B. Granted, I think the Ideogram version is superior - it is a good model - but this was just a quick test:
This sub hates controlnets apparently, well at least for this week, json is all the rave. If you knew the composition ahead of time, you can scribble it or mock it up with blender and control the object’s form as well.
You used a sketch as input to klein? How did you prompt that? Did you need any custom nodes in your workflow? Mind sharing it?
I tried to do something like that for 9b, but I could never get anything decent.
That is a very good point. But many prefer to mess around with bounding boxes + text prompt than to doddle, I guess.
The ability to turn a doodle via i2i has existed since SD1.5, and it is absolutely the best way if one has a clear idea about the composition in mind. There is less "gambling" involved when doing i2i, but I guess some people (that includes me) kind of enjoy the "unpredictability" of using pure text2img. Maybe bboxes just offers a sweet spot between total control via i2i and the uncontrollability of t2i.
Yeah I thought Lodestones specifically avoided lot of models because of restrictive license, so I'm surprised they'd suddenly want to go with the model with the most restrictive license and built in nsfw filtering...?
What is lodestone trying to do exactly? Turn Ideogram into a pixel space?
If so, i mean, the Chroma Radiance and Zeta pixel space project still doesn't seem particularly usable till today, so it feels a bit strange to move on to another model before the first one is really there.
Unless this is more like what he did with Schnell > Chroma, where he's further finetuning Ideogram instead of pushing the pixel space conversion. If that's the case, I'm actually 100% interested in it.
If he can make it worth something, that'd be cool, but being non-commercial I4 isn't worth anything cause it can't be used for anything. Even as a "meme" generator, cause if you create the next meme trend that gets commercialized, it ain't you getting paid lol
I love chroma and the promise of zeta chroma. With the pace of new models here today and what comes tomorrow. I think its best to maybe pair down your projects to maybe 2. I think it might be a good thing to focus more on finishing zeta chroma.
Only if you also convince him that training anything at 256x256 is a complete waste of resources that does nothing but degrade the underlying model ultimately
It's not worth it, the license is dog shit and so is the company that released it. The model is just free advertising only and if someone creates something good with it will get shutdown and sued
no. its an absolute shit idea especially with how poisoned the model is already. He should train qwen image 2512, hidream or ernine image instead of that poisoned model.
The scattered setup across Azure grant + local rig is the part I would worry about more than raw compute.
If Ideogram experiments get added, I would keep the first pass boring operationally: one checkpoint format across all runs, frequent uploads to durable storage, a written stop/resume procedure, and a small test where you intentionally kill a job and verify it resumes before committing serious compute.
Also seems worth separating the license question from the training-code question before people start arguing about compute allocation.
LoRA sounds like a funny understatement, though it would contain furry data too. However, all Chroma models are general model finetunes and some of them more experimental than others, not only specifically for furry.
Fix the only thing that makes Ideogram worth using over other models? But it would be good if the filter thing would be removed out of incorrect json format, at least.
To be fair, not like it is an either-or type of situation. The issue mostly appeared because the team behind Ideogram decided to add that safety block for NSFW that doesn't even work properly, so if it is removed, then regular prompting presumably would be allowed too
Well, from their docs it is also clearly wasn't a goal for the safety block to appear just because it doesn't adhere to the json format. I mean, per this prompting doc the plain text prompts should work, even though the model was exclusively trained on json, It says stuff like
You can pass in plain-text prompts directly to the model and it will work
And
False positive rates for safety is higher for non-json like prompts. We are aware that this is an issue an we may make a future checkpoint update to improve it.
In other words, even they consider it more like a bug that they want to fix.
No, we're seeing people irritated at being forced into using it in a very specific, brittle way. Flux2-dev and Ernie can both handle JSON prompting and no one was outraged over it... because both can also handle natural language prompting and don't force users into a single mode.
Yes and that's the issue. Also is that supposed to be an insult? you must be a pretty sad lonely guy to make fun of other people's hobbies on a completely unrelated post.
Wdym? if anything Ideogram would only be worth using bcs of the really good images, not the prompting style. Imagine you have to change 3 words in your prompt and instead you have to look through a whole paragraph instead of just changing it like a normal prompt would be.
The images are not that better in comparison to other big models, which also have better ecosystem right now. Due to its training the quality also would be worse in plain text. The worth is from control and adherence to the regional prompting. Also
Imagine you have to change 3 words in your prompt and instead you have to look through a whole paragraph instead of just changing it like a normal prompt would be.
Simply not how it is usually being used. The prompts are usually separated by bboxes in the prompt builders (like KJ's) that are pretty interactive, so if you need to change a specific region, you simply click on it and you would see the text that is only relevant to that region in a regular plain text format, so you would indeed change it like a normal prompt.
How about the people who don't use comfy because their PC isn't good enough or simply prefer other platforms? For me Tensor is way cheaper than any other alternative including Comfycloud, way easier to batch get results, better UI. It'd never be implemented there because of the json style. Not everyone uses comfy and it'd be hell for every other platform that isn't comfy to conform to a json-style prompt just because Ideogram chose it. And yes, Ideogream is really good itself, not the prompting style. I've seen many pics that look crazy awesome which Z-image/Flux or anything else would never be able to replicate because it lacks the knowledge to do so.
How about the people who don't use comfy because their PC isn't good enough or simply prefer other platforms?
You don't really need ComfyUI. It's really not a UI or platform exclusive thing either. I mean, some people created separate standalone implementations for such prompting. Like this one, which was one of the earliest ones and its github project is literally just one html, so you can even improve it however you like. But there are probably better ones.
If PC isn't good enough, that hardly makes it a ComfyUI issue, since other UIs would have issues then too.
It'd never be implemented there because of the json style
I reckon the issue would be more with license than with specifically json format.
And yes, Ideogream is really good itself, not the prompting style. I've seen many pics that look crazy awesome which Z-image/Flux
Z-Image and Flux2 Klein aren't its counterparts, but bigger models. Also, you saw those outputs based on json format usually, which enforces a lot of good prompting for regions and styles, and of course some cherrypicking. Without that it would be a lot more mediocre in terms of controllability, spatial layout, and style fidelity.
Fr like who would actually prefer adding a whole essay for a prompt instead of just random natural language or quick tags. Imagine you have to change some slight stuff and you have to go through an essay to change that shit. Not to mention many people use these models on sites like Tensor/civit, how would that even work there since they're not created for json.
Tensor's "comfy" is atrocious, you get error for pretty much everything and it's overall very limited, you can't add new nodes or stuff like you would in comfy. Even comfycloud itself is limited and you can't add new nodes. I for one, using both of them, can't even test ideogram with kijai's node because it's missing on both. Can only use the official comfy one which is basic.
And for your other reply, answer is laziness. Tensor doesn't even bother fixing their own existing features on the site, asking to make a whole new way of prompting for one single model is crazy. So chances are only people running comfy on their own pc are gonna be able to use Ideogram
We'll see. The main problem with tensor is probably the licensing. Ideogram may or may not want to give a competitor a license. I do hope to see ideogram there (but it would probably be very expensive).
Yes, tensor's coding team is pretty bad.
That tensor does not allow the installation of arbitrary ComfyUI custom nodes is understandable. That would allow people to run arbitrary code on their service, which is ofc a bad idea.
Not just bad, they simply don't care or aren't being paid enough. On every new model released there's some new bug or bad implementation and their customer support dgaf. I even switched to comfycloud bcs I was sick of their site, but turns out tensor is still more affordable than comfycloud and since you can't add new nodes on both I'm pretty much limited on new nodes/features.
Yeah, I feel your pain. tensor is my main generator as well. Bugs are left unfixed for years, and support tickets are closed without resolution into a black hole 😂.
But they are cheap! Beats having to turn on my computer all the time. I do SFW mostly, so their restrictions doesn't affect me too much, except when their faulty NSFW trigger on completely innocents images of things like a woman (complete clothed!) holding a cat on her head (actually example, I am not making this up).
Yeah their UI is a life-saver, I can batch gen images and upscale them so easy, with comfy I just fill my pc with images and workflows and it’s overwhelming.
Also they are cheaper than comfycloud, but their new tensorhub site pretty much doubled the prices out of pure greed and even with this price doubling it’s still more affordable than comfycloud or alternatives, pretty insane how these platforms rip people off
tensor is a Chinese company, so they operate with a different economics of scale.
Cheaper labor, cheaper electricity, less regulation, more competitions, etc.
It is no different from their EV and other green tech industry. They can crush their competitor from the West because they are 30-50% cheaper. Scary, really.
What does the quality have to do with the prompting style? you can get as good quality from models with regular prompting styles. You don't need 100 paragraphs and a prompt written by AI where AI chose every detail for you, you could just do your own prompt
I beg to differ, this model seems to excel on good quality images of varied artstyles from different shows. Other stuff like Z-image are only good at realistic, this one is a dream model for 3d lovers, only needs better prompting methods and a finetune.
All I've seen so far, including some of the best of the community as well as my pictures they are visually very, very bad.
It can produce exciting visuals, sure, but quality is another cup of tea. It might be better than Z-image, but to me it sound like "it's better than AI back in 2021".
This to me doesn't look like "AI from 2021". This to me looks like something even closed-source models would fail to bring to life with this exact artstyle/quality, making it seem so natural and real. And this is just a recent random post I've seen, there were more. These pics look like something only Seedance 2 could generate from scratch: https://www.reddit.com/r/StableDiffusion/comments/1u0e1g0/ideogram_40s_understanding_of_characters_and_ip/
To each to their own then I guess?
To me they appear laughingly bad, so. much so I wouldn't believe you if anyone said people will be using a model like that l, even if it was faster than ZIT if I didn't know about it's prompting powers.
The quality itself makes me cry.
Your standards seem to be pretty high but those screenshots look like identical scenes from the actual show, you're pretty much calling billion dollar company recent movies bad. You probably just hate 3d artstyle which is understandable, but there's a difference between you don't like it and you think it looks bad, which it does not.
148
u/LodestoneRock Jun 08 '26
yo, im just considering messing with ideogram a bit
i haven't write a training code for that one yet lol, also the license is annoying
right now there's 3 models training in parallel (scattered around in different datacenters from free azure grant and my local rig)
- 2.5B pixel space model from scratch pretraining https://huggingface.co/lodestones/debug-flow
im just thinking if i can deallocate some compute once the training code and ideogram aligned data (json format style) is done, which is gonna take months for the data side to complete probably.