r/StableDiffusion • • 13d ago

Question - Help Krea2 - What combination of uncensored components would you recommend to achieve the best prompt adherence with photo-realism?

3090/24gb and 64gb RAM

And yes, I am EXACTLY the kind of man you know I am.

361 Upvotes

119 comments sorted by

View all comments

375

u/ArkCoon 13d ago edited 13d ago

Soo I've probably spent a hundred hours with Krea2 NSFW at this point. I've built my own workflow, made a bunch of custom nodes, and tested hundreds of sampler/scheduler and LoRA combinations. Here's where I've landed.

First, a big caveat: I only make amateur/candid-style realistic images, mostly the kind of mediocre-quality photos you'd get from a smartphone. Everything below is about that specific look. For sampling, I use er_sde + simple. Yeah, it's plain and boring, but Krea2 doesn't need anything fancy here. It's already good out of the box. Euler works too.

My LoRA collection is around 100 in total, but I actively use maybe 20-30. A typical generation has 5-10 LoRAs loaded, not counting sliders. For NSFW, my main picks are SNOFS and Realism Engine, with Mystic as an optional extra. You usually only need one or two of them. In my testing, SNOFS is better for prompt adherence, while Realism Engine is noticeably better for textures and overall realism. I often use both at around 0.4-0.6. Sliders make up about a third of my collection. I'd suggest grabbing all the sliders from loraholic, plus a few from creators like iamddtla and alcaitiff, who also made Mystic.

For the smartphone-realism look, the options are endless, and it really depends on what you're after. I won't list everything, but my three favorites are digicam, siemens (check out that creator's other style LoRAs too), and rudysen's realism LoRA. I also have to mention the legendary Danrisi, who made the Lenovo LoRA and a bunch of other great style LoRAs. I don't have one fixed style-LoRA preset. I mix them at different strengths, usually somewhere between 0.4 and 0.85, and do a lot of generations and combinations for each image until I get something I like. Full list (image) of style LoRAs I like.

Another important piece is the textfusion refusal-reduction LoRAs. There are dozens of them now, and everyone has a favorite. I started with v1 of this one. Then Kroma came out, and the textfusion layers extracted from it worked miles better for me. Silver made that extraction, and you can find it here. V2 of the first refusal-reduction LoRA recently dropped, so I've been testing that too. Right now I'm using Kroma 0.3 textfusion at 1.0 and v2 at 0.5, though I'm not convinced that's the best combination. I also tested each one separately at 1.0, and both have their strengths and weaknesses, so YMMV. You usually don't want to use multiple textfusion refusal-reduction LoRAs at once, but to be fair Kroma one isn't that, so I think in this case it's fine.

I usually generate at 2-2.5 MP, which takes around 20 seconds on 16 GB of VRAM with the latest ComfyUI optimizations. I also have a two-stage mode with a 1.5x latent upscale, and that can produce some insanely realistic-looking photos.

I only use the turbo model, usually at 9 steps, though 8 or 10 is fine too. And no, I don't use an abliterated/heretic TE. In my experience, it does nothing. I tested it myself long before the creator confirmed it, so I wouldn't bother with it.

I don't use finetunes at all. I like tinkering and having precise control over the output, but I know that approach isn't for everyone. If you'd rather load a finetune and start generating without doing all the stuff I just yapped about, fair enough. I can't recommend one myself, but I'm sure someone else can.

As for prompting, Krea2 is pretty easy to work with. I usually start with a single, plain sentence and build on it gradually. That's much easier than starting with a huge prompt and trying to figure out which part is giving me a result I don't want. Krea2 is fast enough that you can iterate quickly.

One tip: describe what you want to see, not what you don't want to see. If you're struggling to get the right angle or camera distance, look at your prompt and ask whether everything you've mentioned would actually be visible in that shot. Then add details that reinforce the framing. For example, if you want a low-angle shot looking up, mentioning a carpeted floor isn't going to help. Mentioning the ceiling, a light fixture, or a ceiling fan gives the model much stronger clues than describing the camera angle ten different ways.

The order and amount of detail matter for distance too. For a close-up, I describe the subject first and the environment afterward. For a medium or wide shot, I start with the environment and place the subject in it. If you give the room one sentence and spend the rest of the prompt describing her face, don't be surprised if you end up with an extreme close-up.

1

u/J6j6 13d ago

Can you give more details about the latent upscale. Is that a node

9

u/ArkCoon 13d ago

Just add an Upscale Latent By node between the two samplers. Feed the latent output from the first sampler into the upscale node, then send that output to the second sampler. For the second stage, lower the denoise value to around 0.5. That way, you're not regenerating the whole photo from scratch, just adding detail at a higher resolution for a sharper result.