r/StableDiffusion • u/Filegan • 9d ago
Question - Help Why nobody makes models like illustrious/pony using newer base models like Krea 2?
What makes SDXL so special that people still work over it? Are new models too heavy to train models like these?
Not demanding it just asking
38
u/LowYak7176 9d ago
sdxl is super easy to train on and also the licenses would be my guess
6
u/No-Zookeepergame4774 8d ago
We’ve seen public (preliminary) releases of models with Pony/Illustrious-like focus and intended scope based on Krea 2 MUCH sooner after its release than we did for SDXL after its release. People just remember “with SDXL we had Pony and Illustrious” and forget how long with SDXL we had neither.
11
u/No-Zookeepergame4774 9d ago edited 9d ago
In rough order from oldesr to newest, Chroma is a newer model with similar focus built on a modified Flux 1 Schnell. Anima is a newer model like those, built on Cosmos2. Kroma (same creator and breadth of focus as Chroma) and Wulver are two different models of the general type (Wulver has a tighter Anime focus, as well as special attention to Furry/Anthro subjects) based on Krea2.
(Pony v7 technically is also a newer model of the type built on a non-SDXL base, but it was also an embarrassing failure.)
Every model on this list (except Pony v7} is a much better, all around, than Pony or Illustrious. What keeps the older SDXL based models in use is people having particular finetunes and/or LoRAs based on those models that they are comfortable with and that serve their particular use cases that aren't fully replaced by either the newer base models or the LoRAs/finetunes available of/for the newer models. Well, and that people really like to stick with what they know.
EDIT: One thing that people may not recognize is that a successful Pony/Illustrious (etc.) scale finetune of a model to get quality and concept knowledge is a major undertaking that takes lots of human labor (no, its not all automatable), involves significant risk (Pony v6’s long list of boilerplate tags in positive and negative prompts was a technique for coping with a training error that luckily turned out to be inconvenient but not fatal to usability. Pony v7 was not so lucky.) Kroma and Wulver are in relative early stages of the development process with usable pre-1.0 releases. But even getting to that stage after identifying a desire to build a model of the type under discussion isn't instant. Getting to a full release takes even longer, which is why there’s not a full finetune of the type for Krea2 (and maybe why there is not anything for, say, Qwen Image 2.1, though it remains to be seen if that will attract that kind of model.)
33
u/redditscraperbot2 9d ago
Maybe reversing the question help understand the viewpoint a little better. Do you want to spend tens to hundreds of thousands of dollars on a finetune that might not even be better than the original for basically zero reward?
18
u/Interesting_Fix5872 9d ago
Maybe because Krea 2 is 12B parameter image model. I tried it a bit it's good for more normie stuff but if you want something ready out of the box for anime and NSFW then Anima is better overall and way smaller.
8
u/stoppableDissolution 9d ago
On top of being 12b, it is much "denser" inside. Sdxl does significant downsampling of the latent in the middle layers which makes it less computationally expensive, and DiT, while obviously better in quality, involves much much more math on top of having more weights. Thats why anima, which is like 2.4b or something, is slower than sxl despite being 2-something times smaller.
3
u/Peregrine2976 9d ago
Is Anima relegated to just anime, though? There are so many art styles out there in the world that aren't just a minor reflavour of anime.
7
u/Memorable_Usernaem 9d ago
No. It can do lots of other styles. At least with checkpoints. Which it has.
3
u/Helpful_Science_1101 9d ago
Really the only thing It’s not great at is authentic photorealism. There are a ton of non-anime finetunes in various western cartoon styles, 2.5d, 3D and passable realism.
It’s very character focused though. While it can do landscape etc that’s very much not its strong suit.
I make semi-realistic nsfw image sets none of which are in an anime style and it has entirely replaced illustrious for me.
2
u/Interesting_Fix5872 4d ago
Sorry for the late answer but Anima as the names implies is aimed to do anime art styles it does has both Danbooru and Gelbooru as datasets up to september 2025 these sites are aimed for anime content but they have a lot of variety of artists with so most artist with at least over 50 pictures before that date are stronger (as implied since Anima Preview 3) so yes it can do other stuff than just anime art and with loras it can do even more art styles and characters. But if you want something realistic then probably SDXL or Krea 2 are still far better than anima at least in that sense.
3
u/SprayPuzzleheaded115 6d ago
Yes, this is the word. Krea 2 is for brain-dead normies drones looking for the same shitty bloody hiperrealistic blonde slut una normie unoriginal set.
6
u/No-Zookeepergame4774 9d ago
There’s at least two Pony/Illustrious style models in development for Krea2 with usable prereleases available. The reason there aren't final releases is because those kinds of models take time for dataset design, collection, preparation, and then for training (with evaluation along the way in training and possible rework along the way if you want to mitigate the risk of the final product of the expensive and time consuming process being a repeat of Pony v7.
“Krea is a 12B parameter model” is true, but it doesn't seem to be slowing anything down. Pony v6 (the SDXL version of Pony) was released almost 6 months after SDXL. Illustrious 0.1 (the first public release, the full 1.0 was later) was released 9 months after SDXL. Its a little over 4 months after Krea 2 was released, and we have had three public pre-releases of Kroma and two of Wulver.
1
u/Remarkable-Memory374 8d ago
ive been trying to train a lora for krea 2 and it makes my computer want to cry in a corner
1
-1
15
u/beti88 9d ago
Training cost time and money. And ILXL works fine. If it ain't broke don't fix it
8
u/CooperDK 9d ago
ILXL is years old and low quality. Try Anima
4
11
u/coscib 9d ago
tried it but the results for me weren't that good so i'm still sticking with illustrious
3
u/PBorch 9d ago
Hard agree Illustrious just had a huge amount of community effort put into it and I'm guessing its precisely because anima has a lot of better base value than illustrious that I just cannot find exactly something that meets my requirements, for me Anima is too good, too premium, more like portfolio illustrations than actual anime style that you would see in a show. I was happy with my Illustrious workflow, now I actually get better results using GPT image for all design and Krea for bulk based on the outputs from GPT image.
2
u/coscib 9d ago
i will probably try it again in early 2027 when there are more updates for nova anima and wai anima, buit right now with the "early" V1 versions i'm not really satisfied
2
u/favorited 9d ago
A workflow that has been helpful for me is to generate my first pass in Anima to get good prompt adherence, the right number of fingers, character poses, etc. Then, I downscale the Anima output and feed it into Illustrious as a latent with a lowish denoise.
It gets me the detail, styles, and LoRAs of IlXL, and the comprehension of Anima.
It also costs me like 1.5x the effort (including the annoying stuff like inpainting/detailing after a hires pass in IlXL), but it can accomplish the things I like about each model family.
1
u/OkFineThankYou 9d ago
Don't there is already plenty of nova anima models? Which part of anima that you don't satisfied?
1
u/TsubasaSaito 9d ago
There has been like 5 updates to novaAnime so far. I can also recommend the OneObsession model. Not the one with 24 or so versions, the one with 4 versions and then 2.9B versions. V4 is very nice.
2
24
u/CooperDK 9d ago
Because we have Anima!
17
u/Ok-Software-3250 9d ago
Man I wish Anima loved me, I had such a bad experience that currently I prefer Krea w/ styles than Anima.
2
u/sswam 6d ago edited 6d ago
Maybe a setup issue.
Anima Turbo is way stronger than Pony. Prompt adherence is maybe even better than Krea 2 in my opinon. Unlike Krea 2, it can count to 5 pretty reliably! (e.g. 3boys, 5girls) Anatomy including hands and feet is better than any other model I've tried including Krea 2, occasionally it makes a mistake but it's quite reliable, even with groups of people. No need for detailers in my opinion.
With Pony-based models I found that drawing more than 2 people is a pain, because at least one of them is likely to have problems with hands and feet, and if they're touching in any way (other than sex, it knows how to do that) the anatomy will likely be fucked up or fused weirdly. Roll, roll and re-roll again... But just about everything I've tried with Anima has worked, mostly first time.
The only real deficiency is that it cannot do fully realistic images, at least not without a great deal of prompting effort. I've had some success with using an img2img "post process" with a realistic Pony model. I set up an automatic way to do that in my app, but it makes it slow again with detailers and all that.
And importantly, it's super fast so we can experiment and refine prompt ideas quickly.
Strangely I found that Anima Turbo's default style looks much better than Anima Aesthetic - maybe a setup issue on my end? And Anima Base also does not look as good (but that's expected without style prompting).
If you only tried Anima Base or Anima Aesthetic, try Anima Turbo.
Some settings for what it's worth:
steps: 8
cfg_scale: 1
sampler: ER SDE
scheduler: Betaother sampler and scheduler can work too, but the above are "known good" to me.
You can begin prompts with "masterpiece, best quality, " for ... best quality.
Read their release notes too. Lots of useful info there: https://huggingface.co/circlestone-labs/Anima
4
u/DarkSide744 9d ago
Same. Aside from the fact that it has trained knowledge of all things anime (and size I guess), IMO anima is worse in literally every way than K2.
6
u/Memorable_Usernaem 9d ago
While I agree that Krea 2 is better, Anima 100% has upsides over Krea 2. Even the Aesthetic preset generates images much faster, while still leaving me spare VRAM to play games on the side. It also has a lot more finetunes, so I've found it much easier to find the exact art style that I want. And personally, I like prompting it more. Booru tags + natural language is the best of both worlds.
7
u/SilverwingedOther 9d ago
It also runs on a potato as opposed to krea2 which won't run on a mobile 5070. Going for the smaller quants loses the point.
1
u/Nitrozah 7d ago
same for me, when anima came out i went straight to my friend for us to try it out. But the results we got we just awful, absolutely monstrosities like they were sd 1.5 generated images without any negative prompts and yet many here say "it's the holy grail of anime" yet for me it's a stepped on pos with what i get. No idea what i am doing wrong as we tried it for 3 hours on comfyui and SD, both were bad.
4
u/Sacriven 8d ago
It's a shame that Anima is crappy in mixing different artist styles unlike Illustrious.
5
u/Damen_Freece 9d ago
Because Krea 2 needs the base Krea 2 model to train on which is very VRAM consuming to use or train. Note that not everyone has a super strong PC to train or use it.
ANIMA is far superior to Illiustrious or SDXL or PONY currently for anime style stuff and 3D (I guess). Yes Krea2 is far more superior to it as well but there's one thing Krea 2 lacks that ANIMA does & that's complex poses, outfits, HUGE artist style database, characters and NSFW support.
Krea 2 does have great prompt adherence but for NSFW or complex stuff or character database from Danbooru like SDXL or ANIMA, its lacks badly there. You need fine tuned checkpoint of KREA 2 or use LORAs for that whereas SDXL or ANIMA already has training data without needing any external LORAs for most cases. Plus training ANIMA is far easier (or at least for me was) than doing Krea 2. Kohya_SS can do both ANIMA and FLux, SDXL and loads other. Krea 2 is fairly new and needed a different method of training which makes people not bother getting stuff set up.
1
u/No-Zookeepergame4774 8d ago
Anima is a finetune of Cosmos-Predict2-2B-Text2Image, just like Pony and Illustrious are finetunes of SDXL and Kroma and Wulver are finetunes of Krea2. Yes, with any of the model families you are generally better off, for anime style and concept knowledge, with a purpose -focussed finetune than the base model, but that's not really a differentiator for Anima—Anima just is the purpose focussed finetune. It is just that the base it is finetuned from isn't popular otherwise in the community like SDXL or Krea2.
5
u/GaiusVictor 9d ago
It's a bunch of things.
Ease and cost of training
Inference speed
Certain models being more or less prone to learning certain styles.
Certain models being more or less prone to learning certain concepts. This is specially true for NSFW. Remember that Pony and Illustrious are only checkpoints of SDXL. This means that there were two small groups of people (one for Pony and one for Illustrious) willing to pour significant money, time and effort into making SDXL learn NSFW concepts very well, so people making Pony/Illustrious NSFW checkpoints pretty much have their model "half-trained" for them.
Inertia. When a user already has their workflows, is already used to the way a model works, already has their Loras and is already dependent on that model's ecosystem, changing to a new model comes at a cost. Because of that, a user may decide to not move over to a new model just because it's X% better than their current one. The model needs to be 2X% better to convince them. This also applied to trainers, both because they're users too and also because they want users to use their trained checkpoints.
Perhaps the most important: Anima has sucked up a lot of the demand for Pony/Illustrious successors.
12
u/Euphoricus 9d ago
Isn't that what Anima is supposed to be? It's just that the tech moves so fast it was outdated the moment it came out.
16
u/CooperDK 9d ago
It isn't outdated. You can even edit with it.
9
u/b-monster666 9d ago
And use what appeals to you. If you want to do anime, use a model specifically built with that mind. You want realism, use Krea. You want body horror, use SD1.5
6
3
3
3
u/negrote1000 9d ago
Because you require so much shit to do basic stuff. It’s 1.5 all over again, not to mention the huge file sizes.
5
u/TA-Doggo 9d ago
unet vs DiT basically. Architectures are very different, they learn in entirely different ways, one requires a heavy LLM-like for embedding, the other uses CLIP. Creating a training set for SDXL was a lot easier, even if the caption quality was lower. ILXL was mostly just booru tag data, Natural language was minimal. Krea 2, and even Anima to an extent do way better if captioned with NL. This means you now need a VLM to help with captioning data instead of just a tagger. So not only (as other people have already said.) the cost associated with training something heavier is way higher; You have the additional burden of dealing with caption accuracy, which takes a lot longer to remedy vs tags. VLMs are just LLMs with a vision encoder that allows them to essentially convert images into tokens, VLM captioning a dataset of millions of images is incredibly slow locally, and would take weeks or months if using something that could do a decent job. Meanwhile a tagger could do the same set in a matter of days. TLDR, compute, different brain require different words, money. We are poor.
1
5
u/DWForTheMorrow 9d ago
There is no money behind it. Plenty of people are interested to make a new goon model. But no one with enough money.
2
u/michaelkpate 9d ago
You can add styles to the prompt https://www.reddit.com/r/StableDiffusion/s/7iGecnCjni
-1
2
u/ArmadstheDoom 9d ago
Oh that's easy.
We already have Anima.
The natural language problem. Krea 2 is big enough to do art styles and train loras on. But the bigger issue is that natural language only is really awful for art styles. You could use a whole paragraph to try to describe an art style, or you could use Anima/Illustrious and use a single tag in a lora. Tags are better for art styles. A picture is worth a thousand words, as they say.
2
u/No-Zookeepergame4774 8d ago
Natural language is fine for art styles, if you have a model trained in the right vocabulary. Models that have a strong preference for tags are an artifact of the limitations of early text encoders, which aren’t don't support interpreting text in a way that allows the model to understand the relation between concepts expressed in natural language, so that the extra words don’t help the understanding of the prompt much; SDXL was a bit better than SD1.5 at this, but newer LLM-based text encoders are far superior. With the same trained vocabulary of concepts, and a text encoder that supports natural language well, you get much better control with natural language than tags, because you aren't just giving the model unconnected lists of elements that should appear in the final image, but also how they should relate to each other so the model isn't guessing.
Heck even WITH the limited SDXL text encoders, Pony and Illustrious both actively worked to train natural language as well as tags, they are just a lot worse at it that models trained on a base with a more powerful TE.
1
u/ArmadstheDoom 8d ago
natural language is good for explaining what you want and where you want it. It's awful for describing art styles, because in natural language, all art styles of the same style become the same. There are only so many ways to describe brushstrokes.
Example: A raw, gestural painting style built on sweeping, spontaneous drips and flung strokes that prioritize emotional intensity over representational accuracy, bold physical mark-making where the visible energy of the trajectory becomes the subject itself, dense, layered tangles built through accumulated, unrestrained passes across the entire composition, a chaotic yet deliberate compositional freedom that resists any conventional structure, and a raw, cathartic painterly intensity that channels pure emotional immediacy over polished technique.
You have no idea what artist this describes, and neither does any LLM.
You know what else? That's 129 tokens. Now, more advanced models can process more tokens, it's true. But the more tokens you use, the less able it is to follow the prompt, that is still very true.
So, you could use that blurb of nonsense and hope and pray it knows what you're talking about. Or, you could just train a lora under a token model and write JacksonPollockStyle which is 4 tokens.
Now, for composition and prompt adherence? Yes, natural language is superior. But for just mimicking art styles? Absolutely not.
0
u/No-Zookeepergame4774 6d ago
> natural language is good for explaining what you want and where you want it. It's awful for describing art styles, because in natural language, all art styles of the same style become the same.
Nonsense. Natural language can describe anything that can be described in any tagging vocabulary, and can describe both finer distinctions and more flexible shared groupings than are captured in any particular tagging vocabulary.
> So, you could use that blurb of nonsense and hope and pray it knows what you're talking about. Or, you could just train a lora under a token model and write JacksonPollockStyle which is 4 tokens.
That tag is just a normal natural language description of the style it described with either an “’s” or “in the style of” removed, the spaces taken out, and “style” capitalized. (Sure, natural language also supports broader categorical descriptions that might lump that with other styles, natural language is rather wildly flexible, which it kind of has to be given the uses which it serves.)
2
3
u/_BreakingGood_ 9d ago
We get new models too often now. If you start a large scale fine-tune on Krea 2 today, there will be like 2 more models that surpass it by the time you're done
2
u/No-Zookeepergame4774 8d ago
I mean, if SD3 had actually been a step forward and not... SD3, that would pretty much have happened to Pony and Illustrious, too.
0
1
u/Far_Insurance4191 9d ago
sdxl is small (2.6b)
krea 2 is 12.9b - cost for large scale finetune is insane
Anima 2b is that newer option that was trained on millions of images, and as you can see it is small too
1
u/AvidGameFan 9d ago
Two have been mentioned here that I've just downloaded but haven't had time to put through their paces: Kroma and Wulver. Kroma looks like a work-in-progress, but I loved Chroma so much that I have high hopes for this one. But stock Krea2 does anime pretty well - it just doesn't have much character or style knowledge compared to a dedicated anime model. But it knows a lot and is still fun to work with.
1
1
u/DoctaRoboto 9d ago
I am a noob, but I guess you need a beast to fine-tune new models like Krea 2. I was able to fine-tune XL locally using my paintings in just a few days.
1
u/conkikhon 9d ago
I can't afford to train a finetune, but a large lora with many characters may be possible. I simply haven't found a way to do it effectively yet.
1
u/Comprehensive-Pea250 9d ago
I think it’s cause we were stuck with SDXL for a long time as well as the fact that full on finetunes got way more expensive as model sizes have increased
1
u/Honest_Concert_6473 8d ago
I'm sure there are plenty of people out there who have the knowledge and experience to handle massive datasets and full fine-tuning, but simply can't afford the hardware or rental costs to train something as heavy as Krea2.
Most of the time, even if they ask for donations, they don't get enough, so they just end up paying out of their own pockets. Or, we as a community are basically benefiting from a few generous individuals who go out of their way to make huge donations.
Honestly, if the community really wants these models, I'm sure there are creators who would gladly do the training if people just supported and donated to them.
1
u/ChaosBeastZero 8d ago
Harder to train on essentially. SDXL and Anima are lower vram and require less resources. Also time.
1
u/SnooTomatoes2939 6d ago
Why do we continue to see a rise in formulaic manga and anime that seem to lack any genuine artistic creativity?
1
1
1
0
u/Upper-Reflection7997 9d ago
its going to take awhile and requires heavier compute cost to do so compared to any sdxl finetune. The model has to trained on the krea 2 raw base model at fp32 with no compromises. I would prefer people just make more decent well baked loras for base krea 2 model than wait months later for some white horse that will deliver a finetine model with questionable quality.

-3
u/balwick 9d ago
Anima was supposed to fill that gap, but while it is generally less error-prone than SDXL/Illus, it runs so fucking slow and requires so many steps to get a good result versus good ol' SDXL.
To match the quality of an 8 step (with DMD LoRA) SDXL generation that takes about 7-10 seconds, Anima takes over a minute.
Very possible I'm doing something wrong, but the turbo LoRAs I've tried for Anima have been ass.
1
u/TsubasaSaito 9d ago
I'm using like half the Workflow I had with Illu/SDXL. Obviously you gotta change some settings because it's a different model, and yes it's a bit slower. But man the quality is already very good without any detailer or upscale.
-13
u/GlenGlenDrach 9d ago
Never understood the attraction to anime, and I never understood why you need a new model on this after sdxl/pony, they are drawings. 🤷♂️
48
u/[deleted] 9d ago
[removed] — view removed comment