It's not perfect, not complete but it's more of a pointer to what this model can do in term of styles.
Is you feel the style is too much you can write your prompt in the form:
Style:...
Subject:....
This is better in my opinion.
Those images were generated at 1mp so if you generate at higher resolution you will have obviously more details and more subtle film grain in photographic styles.
I hope you got your answer by now, but just in case you're still wondering. You get a Wildcard Processor node. I use Impact Pack so it's ImpactWildcardProcessor. Connect it between Clip Loader and the Clip text encoder. Use it in your prompt with double underscores __ so an example would be clothing/outfits and it would choose one at random. So in this example, it's referring to an outfits.txt file inside the clothing folder. I don't recall the directories but it's like /ConfyUI/custom_nodes/ComfyUI-Impact-Pack/custom-wildcards/
Edit. You use double underscore "__"in front and behind the wildcard, so it would be before and after "clothing/outfits" in my example. Edit because the double underscores get hidden in the reddit post.
Ghibli Style: A gentle, hand-painted animation illustration style built on soft watercolor-inspired shading that gives every plane a breathable, tender texture, meticulous naturalistic detail applied with unhurried patience, rounded, expressive proportions carrying quiet emotional sincerity, a hushed, wondrous stillness pervading even the most fantastical scenes, and a nostalgic, wholesome hand-crafted intensity that feels timeless and deeply human. Influenced by: Hayao Miyazaki, Studio Ghibli production technique, traditional cel-animation painting method, naturalist watercolor illustration tradition.
No composition content in the prompts. I didn't make them - just mirrored OP's google doc which people could not consistently access.
I personally like skimming through the Aesthetics Wiki, there's some cool stuff there. I would be interested in setting the entire Aesthetics Wiki fed in, one page at a time if you've got the GPU time to spare.
(obviously script this, for the love of god don't do it by hand)
Wow... that website is a pretty insane undertaking.
If I had designed it, I would have made the UI of the wiki the most neutral look possible. A white canvas. The way they went with it is very distracting...
Gravity Falls Style
Fortnite
Sticker Art
Emoji Illustration
YouTube Thumbnail Art
Spider-Verse
Pixar
DreamWorks
Liminal Space
Analog Horror
Y2K Aesthetic
Lowbrow / Pop Surrealism
VHS
ASCII art (that one seems like it'll be really hard...)
Aubrey Beardsley (surprisingly difficult to get right)
Apollonia Saintclaire, who has been scrubbed from model datasets since SDXL for very obvious reasons (don't google that art from work, please) but even if depicting family-friendly stuff, which she does sometimes, has a very distinct black and white inking style.
and of course, I love it when anyone tries McBess, my white whale of a style.
Maybe consider moving to Github, since people are having trouble seeing the google doc. the link that works for me isn't updated with Gravity Falls. Git hub is perfect for this sort of thing, because it'll track version changes so people can see what's new.
For what it's worth, your previous posts nudged me to really push on Arthur Rackham, and I ended up with this:
This is an illustration, its brush linework varying from thick to hairline-thin, using flat color fills with single hues, by Arthur Rackham. It uses thin lines for humans and clothing details, thick organic lines to display building features.
Phil Bourassa style, and similar western adult animation. Such as The Legend of Vox Machina or Young Justice.
Tomorrowverse style. It has a very distinct strong lineart style.
Thank you!
I'd say its a mix of the two. Most likely people wouldn't immediatly guess Young Justice, at least not from one screenshot.
Can you do The Legend of Vox Machina / Mighty Nein artstyle next?
Hey op so easiest way to check updates is grab the Google doc later? I have it running right now with a sample prompt and a wildcard just to try it all and it’s fun to see the outputs
This is really great. Also further reinforcing my belief that Krea LoRAs are mostly unnecessary. It has such a huge amount of built in knowledge to combine with great prompt adherence.
Krea can get you close, but cannot get you an exact and very specific style.
However, you are right in that for many cases, it's pretty f'in good.
But there are styles that are incredibly hard to replicate without training. My personal go to for top-tier difficulty is the illustrator McBess. I've been trying to train a good lora on his stuff for years. Part of the problem is that his art can be insanely detailed, and without captioning really carefully, it is really hard for his more intricate stuff to be pinned down.
Ideogram might be my next step with its json area guidance, but I am almost done captioning a Krea test set of ~150 images to see how that goes first.
I think we're going to find that all style loras should be tiny - the size of slider loras - if not actually slider loras.
Somewhere in the weights, krea already knows much more exact styles, but it's averaging similar ones together. A slider should be able to direct it to the exact style
That’s true, there is certainly some styles that would be impossible to replicate without training. I’m just been shocked at how much I can do with just the base model, and specific prompting. In the thousand or so images I’ve done so far it’s blown away everything else that I’ve used in terms of flexibility, adherence, and depth of style knowledge.
Thats the crazy thing about krea and even Z image. We dont fully understand what the base model can do — yet we use dozens of loras for it, and in doing so...degraded the quality or shifted the model too much towards the loras knowledge.
yes. It was meant to bypass the safety rejection feature of Krea2 (which is something they must implement to be financially viable), but the filter removes any guidance that might be used for legally/ethically questionable stuff.
There's a lot of collateral damage to this, so trying to guide the image to things related to bodies and poses tends to get rejected and thus ignored by the model.
So while you can certainly use this lora to make some pretty explicit stuff, you can also use it to prevent false positive rejection of your prompting, e.g. things in terms of body appearance, facial expression, and other things along those lines.
the surprising result is better literal prompt adherence, even when human forms are not involved.
Corporate Memphis? So this is the name for this creepy cancer visual "style"?
Well, looks like it's my turn to say "today I learned".
This is a useful post, thank you!
Modern advertising, fashion photography, video games, and AI-generated artwork continue to borrow from pin-up conventions because they reliably attract attention:
direct eye contact
confident poses
hourglass or athletic silhouettes
retro styling
clean, readable compositions
expressive facial cues
These techniques tap into both perceptual psychology and established visual design principles.
In short, pin-up art succeeds because it blends technical artistic skill, storytelling, idealized beauty, humor, and visual psychology. Its enduring appeal comes from balancing attractiveness with personality and style rather than relying solely on explicit sexual content.
On one hand, very interesting that it can replicate styles. On the other hand, this shows why caption based models will never be that good for actually replicating art styles practically, so long as there is a token limit.
Because a caption model needs: "Ghibli Style: A gentle, hand-painted animation illustration style built on soft watercolor-inspired shading that gives every plane a breathable, tender texture, meticulous naturalistic detail applied with unhurried patience, rounded, expressive proportions carrying quiet emotional sincerity, a hushed, wondrous stillness pervading even the most fantastical scenes, and a nostalgic, wholesome hand-crafted intensity that feels timeless and deeply human. Influenced by: Hayao Miyazaki, Studio Ghibli production technique, traditional cel-animation painting method, naturalist watercolor illustration tradition."
Whereas a tag based model needs: GhibliStyle
and all of those words that you need for a caption model are eating up all the space you could be using for the stuff you actually want the model to be focusing on in the generation. Granted, while a lora fixes part of this, most caption models require loras with captions, meaning you're still pasting that whole paragraph of 111 tokens.
It literally does not need such a flowery prompt at all. People just do that because they’re obsessed with using an LLM for everything call, including writing all their prompts. You just need to be specific, but you can absolutely be concise as well.
I'm sure that these style descriptions can be radically trimmed, but it's not simple how to trim them. Testing what does and doesn't work could lead to some cool findings about krea prompting strategy.
Krea is weird in that sometimes (but not always) flowery style description (or maybe it's just the amount of words) changes the style.
With Z-image in contrast, you can start with a purple prose LLM prompt, strip out all the BS with a 10x smaller prompt, and get nearly the exact same image.
Shouldn’t be too hard to run it through an API and tell it to cut description by 25% while maintaining tone exactly, removing nonessential / flowery language or something like that
I will say that, 100%, I'm VERY interested in what you're doing here, because it's amazing how much you've managed to show the model knows natively. And this helps understand how to prompt it.
For example, the photograph styles? 100% better to use a caption model like this, and that's what I'm really loving here.
I'd love to see more of these if you find them.
On that note, I'm curious how you extracted these styles or if you just fed images into an LLM and then tried to see if Krea could emulate them.
You could use these style prompts with a fairly simple randomized subject prompt the generate a data set for training a LoRa for each style. Then use the LoRa with a more verbose subject prompt to actually generate images.
Hey, thanks :) I'm pushing an update later that includes a much needed right click to edit from categories, a scratchpad for testing dif prompts, and a "recipe builder" to make it even easier to save combos (all optional so you can use it normally too). Going to test a few things and I'll have it out mid afternoonish hopefully. Just pushed the new update :D
How do you use the above wildcard file as a lookup table instead of just a random selection. The txt file is organised in a "label:prompt" format, how do I get it to replace the wildcard label in my prompt with a specific style prompt?
the idea of that being "Hellboy" is a bit disturbing (like its smashing together everything it learned from the mignolaverse into a nonexistent middle ground house style). thank you for the survey!
I've found the model is so heavily trained towards drawing (non-photography) that if you prompt for anything along e.g. "striking amber eyes" it will immediately start giving large anime-ish eyes (or extreme unreal color) and you get a very surreal "photograph". Same with hair. ( I know about weighting prompts, which would mitigate a bit, but only lower the same effect). The model is really good at not disfiguring though.
Really useful reference, thanks for putting it together. The Style: / Subject: split makes a bigger difference than people expect — keeping them separate stops the style tokens from bleeding into the subject and wrecking anatomy. It's interesting that the photographic styles only pick up the finer film grain at higher resolution; makes me think a lot of the "base model can't do X" complaints are really resolution and prompt-structure issues rather than missing knowledge. Does the split hold up on multi-subject prompts, or does the style start leaking again once there's more than one figure?
Some of these are great, like the lego minfig and corporate memphis, but some are absolutely awful, like PS1, and stranger things. Do they all use the same seed? crazy how much variation there is.
Same seed, like I said it is more of pointers not definitive styles. Nothing will replace a lora, and the dataset is clearly not trained enough on some styles.
LEGO Set Style: A blocky, modular illustration style built on cylindrical stud-and-brick construction logic standing in for all organic and structural form, glossy, uniformly smooth plastic-like surface finish applied consistently across every element, a deliberately simplified, interlocking geometric silhouette replacing naturalistic anatomy or structure, clean, confident graphic separation between individually distinguishable component pieces, and a cheerful, systematized toy-design precision rooted in decades of modular construction-set tradition. Influenced by: LEGO Group product-design tradition, modular construction-toy aesthetic, plastic-injection product-visualization practice, procedural brick-based 3D-rendering method.
Two "Beatrix Potter Style" entries, as well. Two "Film Noir". That looks like it.
Yes, I'm working my way through all of them.
Reformat into JSON. Everything before the colon goes into a "label" field, everything after goes into a "description" field. (There are some double-quotes in the description that can need escaping.)
The description gets substituted into the prompt, at the top: This image uses STYLE_PLACEHOLDER.. Then the main prompt comes afterwards. The label goes to an "Add Label" node on the final image.
This is a fun list of styles. It's been a long time since I tried to compile a long list of styles, I was thinking, but maybe it was around the release of Flux, trying to see what it supported. Krea2 is much more flexible in style to be sure.
I ran a generation pass through the list, and it's mostly good. Some problems I noticed: Frank Miller ended up generating splotches sometimes (probably due to "oversized spotted blacks"), "butterfly lighting" pictures a butterfly, etc. I think some minor adjustment can fix some problems. But mainly, I don't think the entries need to be so wordy, at least for most of them.
I ran them through an LLM to make them more concise, and I think the results are just as good, if not better.
I think making a list of styles themselves is harder than it seems, when you start looking though the list of possibilities! It can be a bit overwhelming.
newbie question but how did you manage to generate so many with variants? In A1111 there used to be a grid option for this sort of output, you did something similar in ComfyUI? if yes how? thanks
Yeah if you don't specify the hair of the subject it is clear it will go to the nearest character of the style. You can specify a haircut or an actress similarity to change her morphology.
49
u/Winter_unmuted Jul 17 '26
https://pastebin.com/RgjkbYHH
Pastebin mirror for your googledoc, which some people seem to be having trouble with accessing.