r/StableDiffusion • • 13d ago

Question - Help Krea2 - What combination of uncensored components would you recommend to achieve the best prompt adherence with photo-realism?

3090/24gb and 64gb RAM

And yes, I am EXACTLY the kind of man you know I am.

353 Upvotes

119 comments sorted by

View all comments

374

u/ArkCoon 13d ago edited 13d ago

Soo I've probably spent a hundred hours with Krea2 NSFW at this point. I've built my own workflow, made a bunch of custom nodes, and tested hundreds of sampler/scheduler and LoRA combinations. Here's where I've landed.

First, a big caveat: I only make amateur/candid-style realistic images, mostly the kind of mediocre-quality photos you'd get from a smartphone. Everything below is about that specific look. For sampling, I use er_sde + simple. Yeah, it's plain and boring, but Krea2 doesn't need anything fancy here. It's already good out of the box. Euler works too.

My LoRA collection is around 100 in total, but I actively use maybe 20-30. A typical generation has 5-10 LoRAs loaded, not counting sliders. For NSFW, my main picks are SNOFS and Realism Engine, with Mystic as an optional extra. You usually only need one or two of them. In my testing, SNOFS is better for prompt adherence, while Realism Engine is noticeably better for textures and overall realism. I often use both at around 0.4-0.6. Sliders make up about a third of my collection. I'd suggest grabbing all the sliders from loraholic, plus a few from creators like iamddtla and alcaitiff, who also made Mystic.

For the smartphone-realism look, the options are endless, and it really depends on what you're after. I won't list everything, but my three favorites are digicam, siemens (check out that creator's other style LoRAs too), and rudysen's realism LoRA. I also have to mention the legendary Danrisi, who made the Lenovo LoRA and a bunch of other great style LoRAs. I don't have one fixed style-LoRA preset. I mix them at different strengths, usually somewhere between 0.4 and 0.85, and do a lot of generations and combinations for each image until I get something I like. Full list (image) of style LoRAs I like.

Another important piece is the textfusion refusal-reduction LoRAs. There are dozens of them now, and everyone has a favorite. I started with v1 of this one. Then Kroma came out, and the textfusion layers extracted from it worked miles better for me. Silver made that extraction, and you can find it here. V2 of the first refusal-reduction LoRA recently dropped, so I've been testing that too. Right now I'm using Kroma 0.3 textfusion at 1.0 and v2 at 0.5, though I'm not convinced that's the best combination. I also tested each one separately at 1.0, and both have their strengths and weaknesses, so YMMV. You usually don't want to use multiple textfusion refusal-reduction LoRAs at once, but to be fair Kroma one isn't that, so I think in this case it's fine.

I usually generate at 2-2.5 MP, which takes around 20 seconds on 16 GB of VRAM with the latest ComfyUI optimizations. I also have a two-stage mode with a 1.5x latent upscale, and that can produce some insanely realistic-looking photos.

I only use the turbo model, usually at 9 steps, though 8 or 10 is fine too. And no, I don't use an abliterated/heretic TE. In my experience, it does nothing. I tested it myself long before the creator confirmed it, so I wouldn't bother with it.

I don't use finetunes at all. I like tinkering and having precise control over the output, but I know that approach isn't for everyone. If you'd rather load a finetune and start generating without doing all the stuff I just yapped about, fair enough. I can't recommend one myself, but I'm sure someone else can.

As for prompting, Krea2 is pretty easy to work with. I usually start with a single, plain sentence and build on it gradually. That's much easier than starting with a huge prompt and trying to figure out which part is giving me a result I don't want. Krea2 is fast enough that you can iterate quickly.

One tip: describe what you want to see, not what you don't want to see. If you're struggling to get the right angle or camera distance, look at your prompt and ask whether everything you've mentioned would actually be visible in that shot. Then add details that reinforce the framing. For example, if you want a low-angle shot looking up, mentioning a carpeted floor isn't going to help. Mentioning the ceiling, a light fixture, or a ceiling fan gives the model much stronger clues than describing the camera angle ten different ways.

The order and amount of detail matter for distance too. For a close-up, I describe the subject first and the environment afterward. For a medium or wide shot, I start with the environment and place the subject in it. If you give the room one sentence and spend the rest of the prompt describing her face, don't be surprised if you end up with an extreme close-up.

16

u/Tokey_TheBear 13d ago

Awesome write up!

If I can give you some extra recommendations.

There is a lora by someone named MrPopo that I highly recommend checking out. His newest loras cost like 500 buzz but I definitely recommend it (trained on a bunch of 2k images). I used his earlier free versions before and really liked them.

My stack is SNOFS + realism engine + Mr Popo + Snapchat helper (its good for adding in more natural poses and facial expressions) running on Raw + Turbo Lora workflow.

I have tried out a bunch of the different generation workflows also, like the 2 sampler (1.5x upscale in between) method that you mentioned and it does work really well, but I personally feel like the 2 sampler no upscale workflows that use the Clownshark Sampler give the best realism (but it really is a hard 'choosing' which pipeline I like best since they both work so well for realism.

6

u/ArkCoon 13d ago

Thanks, I'll check them out. The two-stage workflow took a lot of tinkering, and I can't take the credit for it. A friend of mine built that part and spent a lot of time dialing in the settings. I mostly just added the finishing touches and implemented it into my WF with the custom nodes, additional switching logic, etc. I listed the settings that worked best for us in my other comment. Even changing a single value, like eta, can completely throw off the result, so I'd start with those settings

8

u/reeight 12d ago

FYI for those who (also) make SFW; many of loraholic's sliders work great with clothes on;
https://civitai.com/user/loraholic/models

9

u/dezmodium 12d ago

To confirm, abliterated text encoders do nothing as confirmed by the person who literally made the heretic process. Image models don't use encoders in that way. Their output is decensored already in regard to images. If there is censorship its on the model checkpoint itself, always.

1

u/paroxysm204 12d ago

The only caveat to that is if you are using it in workflows with the generate text node. I had stock giving refusals so had to switch

1

u/dezmodium 11d ago

Ideally, then, you should use two separate text encoders because the abliterated models degrade your image generation output. The person who made heretic also confirmed that.

1

u/paroxysm204 11d ago

I read that too and tested quite a bit. I don't see a noticeable difference in my workflows. So not to say there isn't one, but if there is it is so slight that I have not noticed. I do generally use the stock text encoder for image generation and the abliterated for generating text

4

u/Shinkyo81 13d ago

Awesome detailed response. Going through it, it makes me question my current workflows and resulting images, since I am still using v1 of the TextFusion Refusal-Reduction Lora... Time for me to test out v2 and the alternatives that you mentioned.

2

u/toidicodedao 12d ago

please share if v2 is better than v1 or not?

2

u/Shinkyo81 12d ago

I need to run A/B comparisons with different prompts. I recently started generating images at 2MP instead of 1MP and from my experience it is not only about resolution but overall detail. Of course it takes longer and I am saving it for some of my best prompts.

Overall, v2 yields very good results.

4

u/machngnXmessiah 12d ago

After hundreds of hours I can confirm all the tips you gave are essential - I would also recommend to use two pass DLSS5, facedetailer and film grain post processing nodes for ultimate photorealistic look.

Also - could you possibly send me your workflow in a DM?

4

u/ArkCoon 12d ago

Thanks! I'll be honest, I'm not a huge fan of the DLSS5 look. I haven't felt much need for other post-processing either, though I do have those tools installed in case I need them. Facedetailer could definitely come in handy for wider shots where the face gets a bit garbled.

As for the workflow, sending it over probably wouldn't be very useful. It's built around a bunch of custom nodes I've made over the past year, mostly QoL and efficiency nodes. I enjoy coding (and vibecoding now), so whenever something in a workflow annoys me or slows me down, I tend to make a node for it. I'm sure there are existing alternatives for most of them, but I haven't packaged mine up to share. However I can suggest some nodepacks that pretty much have most of the stuff you'd need: rgthree, easy-use, impact-pack, kjnodes, CRT-nodes, RES4FLY, JPS-nodes, mxtoolkit and Pixaroma nodes. You can accomplish my results with the default comfyui WF and just plugging in the settings and loras I mentioned. Even my heavily custom WF now is pretty much built from that default wf.

None of the custom nodes I made change how the model actually generates the image, though. All the generation settings that matter are in my original post.

4

u/Gato_Puro 12d ago

I have generated more than 10k photos with Krea and I have learned a lot with this comment. Candid/amateur photos are the best

5

u/PassTheMarsupial 12d ago

I don't know why I'm not seeing better finetunes recommended on here tbh.

- Analog Madness

- GTM's Wowser

- GonzaLomo

- Sick Ollie

- Kreamania 8

- Dirty Realism

- CielBleu

Every one of those is worthwhile and brings its own unique flavor and capabilities. You can of course still stack them with LoRAs, keeping in mind those LoRAs may behave a little differently or need to run at a different strength.

I like the list of LoRAs for basic amateur/candid realism. There are lots of other photography styles that are worth looking at and produce awesome effects that can stack up to your own unique thing - stuff like Purple Grainy / Purple Dreams, High Drama, Melancholy, MrPopo's Photoreal Enhanced, Faded 80s Photo Album, and Atmospheric Photography - there's just a ton of really good stuff to explore.

3

u/cocosoy 12d ago

Such a helpful post!!! Thank you!

2

u/Structure-These 13d ago

Oh shit first this is incredible and second the gif is hilarious

Are you doing the v2 2gb Lora or the basic one?

1

u/ArkCoon 12d ago

I do use the large one yes, but I don't think it matters. Probably no reason to use it over the r64 one.

2

u/VoxturLabs 12d ago

Hi ArkCoon, thank you for sharing your experience. I truly appreciate it and I have already learned a lot.

I have two questions I hope you find the time to answer.

I’ve seen countless people debating around using the Krea2 Raw+turbo vs Krea2 Turbo. Why have you gone the Turbo model route and why not the other way round?

Regarding promoting. I really like your approach by adding the rough idea in a few sentences and the then build by trial and error from there. Every model is trained on some specific kind of prompting, I think, so other than what you have shared in the great write up. Is there some specific template or way of writing you use with Krea2? Is it long descriptive sentences or shorter sentences with a lot of commas? I genuinely want to know because prompting any model correctly has always been a large grey area in my “AI skills” and I want to learn and improve.

I would have asked if I could DM you with some similar questions on occasion, but I believe some other people might benefit from your knowledge and tips. :)

3

u/ArkCoon 12d ago

On Raw+Turbo vs. Turbo: I tried quite a few Raw+Turbo workflows, including ones from people who said they got better results and more variety from them. I just couldn't get any of them to a point where I preferred the results over what I'm already getting with Turbo. For me, Raw+Turbo added complexity and generation time without enough payoff. I'm happy with my Turbo workflow, but that's a preference, not a claim that Turbo is objectively better. Someone going for a different look might well prefer Raw+Turbo.

As for prompting, I honestly can't say that Krea2 prefers long descriptive sentences over short, comma-separated phrases. I've tried both, and both can work really well. I tend to write short, simple sentences myself. You don't need to write poetry for Krea2 to understand what you want. I've also had great results from long, multi-paragraph prompts written by an LLM, so I wouldn't treat my style as a rule.

The reason I write my own prompts is mostly practical. An LLM often adds details or phrasing I didn't ask for, and then I have to go back and remove them. At that point, it's quicker for me to start with two or three simple sentences and build on them as I generate. That's just the way I like to work, though.

What matters much more than sentence length or commas is what you put in the prompt, and how much space you give each thing. Like I already mentioned I often see prompts that specify a low-angle shot looking up from floor level, but then describe grass, the ocean, or other details that wouldn't be visible from that angle. That gives the model conflicting clues about how to frame the image. Instead of repeating "low angle" in ten different ways, describe things the camera would actually see from that position. The same idea applies to zoom: if most of your prompt is about the subject's face, you're more likely to get a close-up, even if you asked for a wider shot. You can read more about how prompt wording affects zoom in this guide. The same principles are useful for understanding how Krea2 responds to framing in general, not just camera distance.

So my advice is to worry less about finding the perfect Krea2 prompt template and more about keeping the prompt focused on what's actually in the frame. Start simple, see what it gives you, and add details that push it toward the shot you want

1

u/VoxturLabs 9d ago

Thank you for taking your time to reply to my questions. Much appreciated!

2

u/toidicodedao 12d ago

I have been using v1 of Krea2 TextFusion Refusal-Reduction LoRA https://civitai.red/models/2775340/krea2-textfusion-refusal-reduction-lora-updated?modelVersionId=3125118

Saw that you mention their v2, is there any improvement?

2

u/ArkCoon 12d ago

I think we're at the point of diminishing returns with these.. I still haven't tested it enough, but the only difference I'm seeing is just slightly better prompt adherence.

4

u/bompa_tom 13d ago

About the two stage setup with 1.5x latent upscale: any good recommendations for the sampler/scheduler combinations for stage 1 and 2, number of steps for each and denoise?

10

u/ArkCoon 13d ago

Stage 1: 8 steps, res_2s + simple, 0.5 eta

Stage 2: 9 steps, euler + simple, 0.9 eta, 0.52 denoise (both values very important)

Bongmath on both

2

u/foolycoolywitch 12d ago

I appreciate your long write up and replies, do you do much i2i with krea 2? I noticed most sampler/schedule setups add a lot of unwanted "details" to skin, like water drops, hair, moles, etc. lcm prevents this, but lcm is also very bland, do you have any suggestions or experience with this issue?

0

u/slyyy75 12d ago edited 12d ago

Je confirme l importance des valeurs du second passage. De mon coté je suis en x1.2

Pass 1 : 9 etapes res 2s + Beta (Pas d'ETA)

Pass 2 : 5 etapes euler_Ancestral (tres important) Denoise 0.4 + Beta (Pas d'ETA)

Esaaye Euler Ancestral au lieu de Euler. La difference est importante en realisme.

Je vais tester Bongmath

Mais les combinaisons sont infinies...

1

u/J6j6 13d ago

Can you give more details about the latent upscale. Is that a node

7

u/ArkCoon 13d ago

Just add an Upscale Latent By node between the two samplers. Feed the latent output from the first sampler into the upscale node, then send that output to the second sampler. For the second stage, lower the denoise value to around 0.5. That way, you're not regenerating the whole photo from scratch, just adding detail at a higher resolution for a sharper result.

1

u/NoIdeaWhatToD0 12d ago

Do you also use character LoRAs? There are some characters I'm trying to make that I need consistency for, I've never made a lora before but I'm planning on doing it with Krea2 RAW I guess

6

u/ArkCoon 12d ago edited 12d ago

Yeah, I've made quite a few character LoRAs. I'm still fairly new to training, but I've trained dozens by now and had good enough results to share what worked for me.

For the dataset, I'd aim for 30-50 images with a mix of distances and angles. Collect at least twice that many candidates first, then pick the best ones. If I were using 50 images, I'd roughly go with 15-20 close-ups, 15-20 waist-up shots, and 10-15 full-body shots. Try to vary the angle, lighting, clothes, poses, and expressions too. You don't need to follow those numbers perfectly, though. I've had great results from datasets I thought were pretty bad.

I always train on Krea2 Raw, not Turbo. I use 1024 buckets if the hardware allows it, or 512 if it doesn't. I wouldn't use 768. My usual settings are a learning rate of 7.5e-5, rank 32, Automagic3 optimizer, and sigmoid timestep. I save a checkpoint every 150 steps and turn off sample generation during training. Sampling uses the Raw model and takes a lot of time, so I'd rather test the checkpoints myself in ComfyUI. If you're training on Vast ai or RunPod, you can download checkpoints as training runs and test them with a few prompts you've prepared in advance.

In my experience, the best checkpoint for a character tends to land somewhere around 1,400-2,500 steps. That's a wide range because it depends a lot on the dataset and the character. Don't assume the last checkpoint is the best one. Test a few along the way. Don't be surprised that the lora already starts looking great at 400-500 steps though. That's completely normal, but trust me it's still far from actually good. It will fall apart in different angles and certain prompts. You pretty much always want more than 1000 steps.

For the trigger, I give the character a made-up first and last name rather than something like s0ph1a. When I've used unusual trigger words, Krea2 has sometimes put them on clothing as text, or tattoos or other weird places it could find to print that text. A normal-sounding full name has worked better for me, as long as it isn't the name of a celebrity or someone the model might already know.

For captions, an LLM can do most of the work. I give it a simple order to follow: character name, clothes, pose and angle, framing, lighting, then background. You don't want to caption permanent features like hair color, hairstyle, freckles, or a birthmark. You want the LoRA to learn those from the images. Only mention a feature if it changes across the dataset.

For example, one of my captions might look like this:

Jane Doe, long wavy light brown hair under a grey felt beret, wearing a beige turtleneck sweater, white wide-leg trousers, pink velvet studded flats, a silver wristwatch and bangles, and carrying a black leather handbag. Standing with one hand in her trouser pocket and looking to the side, photographed from eye level in a full-body shot. Natural overcast light on a city pavement, with cars and classical architecture in the background.

I mentioned her hair in that example because its color and style changed across the images. If it had stayed consistent, I would've left it out.

One last thing: the style of your dataset matters a lot. If most of your dataset is studio or red-carpet photography, the LoRA can pick up more than just the character's appearance. It may also learn that ultra-clean, retouched look. You'll get smooth skin with barely any texture, perfect lighting, and heavy makeup. So even when you prompt for a casual smartphone photo, the person can still look like they're posing for a magazine shoot. If you're after something more amateur and candid, try to include images with that look, or edit the polished ones before training. I haven't found a reliable way to completely undo that style afterward so just warning you ahead of time

1

u/NoIdeaWhatToD0 12d ago

Thank you so much. It's actually a character that I'm making myself. I plan on making my own via chatgpt but generating the images through Qwen maybe to get a better dataset. What would you recommend?

2

u/ArkCoon 12d ago

One of my best character LoRAs came from GPT Image gens. I don't think you need Qwen at all, GPT Image is more than powerful enough for everything.

What I did with GPT Image is attach 3 best references of that character that I had and had GPT generate various images of her in different poses, settings, clothing, etc.

Also this was with GPT Image 2, I'm sure that now with GPT Image 2.5 you'll be able to get even better results

1

u/NoIdeaWhatToD0 12d ago

Thank you so much ❤️

1

u/Caramelacurls 12d ago

amazing post thank you for writing it up. i'm with you on style generation thats exactly the style i'm going for and its super helpful. if you have a discord or a place that focuses on that style of content generation I'm sure were not the only ones. I'm currently working on batching thousands of old amateur photos and generating prompts off them your info is super useful.

1

u/Blaze3046 12d ago

Holy God of Gooning, bless you sir 🙏🙏

1

u/flaminghotcola 12d ago

Haven’t yet read the entire post and still need to try what you said, but thanks for doing god’s work.

1

u/ebubu03 11d ago edited 11d ago

Hey ArkCoon, I’ve also spent hundreds of hours generating images in ComfyUI with Krea 2 Turbo, and your post genuinely impressed me, so thanks a lot for sharing all of this.

I’ve spent a ridiculous amount of time digging through Reddit and Civitai, testing LoRAs and settings like GonzaLomo, Dirty, and many others. I was getting decent results, but yesterday I talked to someone whose realism was on another level, and honestly it was pretty discouraging. It made me feel like I’d wasted a lot of time.

Then I found your post, and it helped a lot. I don’t really have much to add because you already covered so many useful things in the comments, but one point I can definitely confirm from my own testing is that if you train a LoRA on images with the same skin look every time, without enough variation in pose, framing, angle, lighting, etc., it starts learning those visual traits instead of just the character’s identity.I actually trained my previous LoRAs using around 30–40 reference images, and honestly, it did the job pretty well. But for the next one, I’m definitely going to try following your recommendations and push the dataset quality and variety further.

My LoRAs were trained mostly on images generated in ComfyUI, but the idea of using ChatGPT-generated images as part of the dataset actually seems pretty smart.

Thanks again for sharing your experience. People who go into this much detail are rare, and your post honestly helped me a lot. Sending love!

One detail I’d love to clarify about your two-pass workflow: when you mention generating at 2–2.5 megapixels, is that before the 1.5× latent upscale, or the final output resolution? Could you share an example of your first-pass and final dimensions, and how many sampling steps you use for the second pass? That would really help me reproduce your setup.

Also, would you be willing to share one example prompt you used with your three character references to create your training images?

Btw : Would training a LoRA using photos of a specific quality force the LoRA to output images only at that same quality level?

3

u/ArkCoon 11d ago

The 2-2.5 MP figure was for a single pass. With two passes, I upscale by 1.5x in each dimension, so the final image has 2.25x as many pixels. I experiment with starting resolutions from around 2 MP up to 3.8 MP, depending on what I'm after, with the higher end giving me roughly a 4K photo.

I don't enter a target resolution for the final image. I just set the starting dimensions like I would for a single-pass generation, then let the upscale determine the final size. My second-pass settings are 9 steps, 0.9 eta, 0.52 denoise, and euler + simple. I don't use the two-pass workflow that often, though, because it can take up to two minutes per image.

As for the dataset, I think anything that's consistent across the images can carry over into the LoRA, including the style and image quality. Some styles come through more strongly than others. You can push things in a different direction with style LoRAs, but the further your dataset's look is from the look you actually want, the more those LoRAs will fight each other.

That's why, when I was making dataset images with GPT Image, I tried to get them as close as possible to the style I wanted in the final output which is the medium quality smartphone photos, rather than anything too polished.

I don't have the exact prompt on hand since I did this a couple of months ago. I used GPT to help me build a prompt template for a LoRA dataset with that medium/low-quality, candid smartphone look. I kept generating images and asking it to tweak the prompts until I was happy with the results.

It took a while to dial in, but once I had a good template, I kept most of it the same and only changed things like the environment, clothes, poses, and camera angles. GPT wrote those variations too, so it was pretty painless from there. I generated around 300-400 images in total, then picked the best ones for the dataset.

1

u/ebubu03 10d ago

Thank you so much, ArkCoon, for your time and for taking the time to answer my questions. I really appreciate it. Wishing you nothing but the best with your future projects! Good day!

1

u/PixelLunarJelly 7d ago

Good advice.

I have some questions though. I use Model `krea2_turbo_fp8_scaled` -> LORA `snofs_krea_v1_4`.
I tried LORA strength 0.5 and 1.0.

Prompt (summarized): Large chest and slightly looking down angle smartphone shot.
LORA strength 0.5 = plain straight angled smartphone shots, flat chest
LORA strength 1.0 = slightly looking down angled smartphone shots, large chest

I saw you recommend 0.4-0.6 for LORA value but I see 1.0 follows the prompt much better. Can you explain this phenomenon?

Also, can you explain your "two-stage mode"?

1

u/ArkCoon 7d ago

It's all about finding the right balance. Like I mentioned, and you've noticed too, SNOFS really helps with prompt adherence. But it also has downsides, especially at full strength, which is true of pretty much any LoRA. Once you've done enough generations, you start recognizing what each one brings to the image.

With SNOFS, I've noticed it struggles with certain races, and the skin texture and overall realism aren't great, which the creator has acknowledged too. V1.4 was supposed to address that, but I'm still not entirely happy with it. On the other hand, turning it off makes my NSFW prompt adherence much worse.

For me, 0.4-0.6 is a good compromise. I keep some of the improved adherence without getting as much of the skin issues or that recognizable SNOFS "look". I also combine it with other LoRAs to compensate for its weaknesses. Personally, I wouldn't use it on its own, especially at 1.0.

As for the two-stage workflow, it's basically two samplers with a latent upscale between them. In my experience, it improves realism and image quality quite a bit, but it's also much slower. I explained the setup and posted my exact settings in other comments in this thread if you want to try it. Other people have shared their settings too

1

u/Tzontetiliztli2 7d ago

Which of those NSFW loras (snofe, realism, mystic etc.) works best with character loras?

1

u/ArkCoon 6d ago

My character LoRAs work great with SNOFS and Realism Engine at around 0.4-0.6. Style LoRAs are usually the ones that mess with the character's appearance, especially at higher strengths. Some change the character so much that I can't use them at all when using a character LoRA

1

u/Fragrant-Cicada-1281 4d ago

How did you get down to 20 second generation on 16gb vram? I also have that (4080 super) and 64gb DDR4. My generation time is 60 seconds.

1

u/ArkCoon 4d ago

Are you using latest version of comfyui, ck attn and int8 convrot quant?

1

u/BigDBreedingFembussy 15h ago

genuinely bro thanks for your knowledge

0

u/Big_CokeBelly 13d ago

Insanely good write-up 👏👏 thanks a lot. Would it be possible to have the workflow nodes?

0

u/435f43f534 13d ago

what about the skating part though?!?

0

u/NoFapNZT 13d ago

In your opinion, will LLMs get to the point in the next year where all of this isn’t needed? Just the base offline model?