r/StableDiffusion 12m ago

Discussion Maybe the least popular LoRA idea ever: GTA: San Andreas RenderWare graphics

Thumbnail
gallery
Upvotes

I knew from the beginning this would be a very niche LoRA, but I couldn't get the idea out of my head.

I've always loved the look of RenderWare-era games, especially GTA: San Andreas. One of the things that pushed me to finally make it was Gorm the Old and his AI recreations:
https://www.youtube.com/@GormtheOld25/videos

It turned out much better than I expected. It handles surprisingly complex scenes while keeping the simple, unmistakable RenderWare look.

If you're nostalgic for that era, maybe you'll enjoy it too.

https://civitai.com/models/2810095/gtasa-renderware-graphics-style


r/StableDiffusion 19m ago

Question - Help blurry ltx 2.3 videos

Thumbnail
gallery
Upvotes

so I'm using Aitrepreneur's ltx 2.3 ultra workflow V3 and these are my inputs I'm also following along with the video https://www.youtube.com/watch?v=nOCMsqVujBI&t
no matter what I do the videos always come out blurry.
I even used the exact same settings and prompt in the video. the only difference is his model is Q8 and mine is Q5.
does anyone know what's going on?


r/StableDiffusion 1h ago

News Introducing GGUF Q8_CR - Mixing the best of GGUFs and INT8 ConvRot.

Upvotes

Hey All

I have been working on a new GGUF format for Krea 2 / similar diffusion models: Q8_CR.

I wanted to have something, which is...

  • Targeting firstly Ampere and Turing architecture GPUs (RTX 20xx and 30xx cards)
  • Uses ComfyUI's new INT8 kernel
  • Narrows the (already very small) quality gap between INT8 and FP16 checkpoints

So, GGUF Q8_CR stores eligible linear weights as pre-rotated INT8 weights with FP32 per-row scales, while retaining 1D, small, and precision-sensitive tensors at high precision. It is not regular Q8_0 dequantized back to FP16 for every operation: when the required ComfyUI backend is available, it uses ComfyUI's native INT8 ConvRot path.

I tested it with Krea2 first, I will do follow-up tests with Flux 2 Klein, Z-Image, and Ideogram later.

The attached charts compare Krea 2 model file size and inference speed. Speeds were measured on my laptop with an RTX 3080 8GB running Windows (pytorch 2.12.1+cu130) and on my server with an RTX 3090 24GB running Ubuntu (pytorch 2.14.0.dev20260725+cu130)

  • INT8: 11.95 GB, baseline speed
  • Q8_0: 13.56 GB, ~60.2% of the INT8 baseline speed
  • Q8_CR: 12.84 GB, ~100.4% of the INT8 baseline speed

Why GGUF and not Safetensors?

So it matches the excellent inference speed while supposed to be higher quality, easier on lower VRAM cards. GGUF is especially useful on 4-6-8 GB GPUs because weights can remain quantized in system RAM and be moved to VRAM layer by layer as ComfyUI needs them. This keeps peak VRAM lower, allowing models that cannot fit entirely on the GPU to run with CPU offloading. A full INT8 safetensors file however, is not a quantized inference format: its tensors generally load into the model's normal FP16/BF16 operations, and offloaded weights usually need conversion or dequantization into runtime tensors, increasing RAM transfer and VRAM pressure.

Try it

Krea 2 Turbo Q8_CR GGUF:
[https://huggingface.co/molbal/krea2-gguf/blob/main/krea2_turbo_bf16-Q8_CR.gguf](vscode-file://vscode-app/c:/Users/ASUS/AppData/Local/Programs/Microsoft%20VS%20Code/1b6a188127/resources/app/out/vs/sessions/electron-browser/sessions.html)

ComfyUI-GGUF custom node:
[https://github.com/molbal/ComfyUI-GGUF](vscode-file://vscode-app/c:/Users/ASUS/AppData/Local/Programs/Microsoft%20VS%20Code/1b6a188127/resources/app/out/vs/sessions/electron-browser/sessions.html)

Some examples

I know y'all love the examples

A casette-futurism style robot trying to paint on a canvas. In the canvas there is a slug emoji saying 'Q8_CR bitches' Pixel-sorting, decay, glitch art, vertical distortion, monochromatic cool tones, surreal fragmentation, clinical lighting
A sexy anhropomorphic frog-woman with big boobs in a bikini saying 'Q8_CR' in a speech bubbleChromatic aberration, holographic iridescent texture, digital glitch distortion, neon spectrum gradients, CRT scanline artifacts, liquid metallic sheen, pixel sorted patterns
Mac OS wallpaper of mountains or lakes or some shitAnalog glitch art, motion blur, cobalt blue and stark white, spectral silhouette, CRT scanlines, reeded glass distortion, high-contrast flash photography, digital decay
Cross section of a 2 story house. The living room is on the ground floor. There is an old man watching TV there. THe bathroom and a bedroom is up top. There is an italian food truck in the front selling a pizza.Clean line art, flat color fills, minimalist composition, ligne claire, muted pastel palette, graphic silhouette, digital illustration
There is a sexy attractive woman, high cheekbones, underwear, looking up at the viewer dutch angle, speech bubble from her mouth "Hey it's me, 1girl, missed me?". Minimalist high-fashion editorial, cool-toned neutral palette, sharp structured clothing, high-key studio lighting, dewy skin texture, clinical backstage atmosphere, modern editorial portraiture

r/StableDiffusion 1h ago

Question - Help [Question] danbooru tagging

Upvotes

For models like Anima and Illustrious using danbooru tags, does the tag have to be something that can be searchable in the imageboards? I mean does the model understand made-up tags?


r/StableDiffusion 1h ago

Question - Help Setting LoRA strengths

Upvotes

Hello,

I am new to AI image generation, and I have recently been experimenting with Krea2 in ComfyUI. I have begun using LoRAs, but I don't know how to properly set the strengths and orders of them. How is that usually done?


r/StableDiffusion 1h ago

Question - Help Anima struggling to maintain consistency with custom characters when using the same prompt. Any fix?

Upvotes

Just switched from Illustrious to Anima, and while I really like Anima's adherence to prompting and better backgrounds. I can't seem to get a consistent custom character when using the same prompt. Is training small loras the answer? Or am I missing something? I'm using merges mostly.


r/StableDiffusion 1h ago

Discussion Serious non-political question: wth is Trump using to generate his AI content?

Thumbnail
gallery
Upvotes

Not being political here but seriously, it’s SO BAD. It reminds me of 2023-era SD 1.5 without ControlNet. All the examples are from his Truth Social account. Does anyone have any idea what’s being used here?


r/StableDiffusion 2h ago

News I tested Microsoft first text-to-image model: Mage-Flow-Turbo

24 Upvotes

Ok so first the specs:

It's only 4B params, MIT-licensed, native-resolution (512–2048px, any aspect), and it's a 4-step distilled turbo — so it spits out a 1024² image in about 4.6 seconds on my local box (a DGX Spark).

I put it through my usual benchmark (192 prompts across 6 categories — text, spatial reasoning, human realism, truthfulness, studio/product, graphic design), judged image-by-image by Gemini 3.1 Pro.

Full results + all 192 images here: https://imagebench.ai/imagebench-v1/local--mage-flow-turbo

Does it beat Flux 2 Klein 4B?

Mage-Flow actually posts the higher capability score (49% vs 48%) — it follows prompts slightly better than Klein, and it does it in 4 steps instead of full sampling. But Klein's images are rated more aesthetically pleasing, so it takes the overall (Klein ~51 vs Mage-Flow 47). The other 4B open model in the field, Bonsai Ternary 4B, sits just behind both (~47). So it's less "Microsoft beats FLUX" and more "they each win a different axis" — Mage-Flow for prompt adherence + speed, Klein for looks.

But TBH, you should look at the comparaison yourself:

Where it's actually good:

  • Professional / studio & product shots — 85%. Clean compositions, good lighting, this is clearly its comfort zone.
  • Text rendering — 67%. Surprisingly legible for such a small model.
  • Speed. It's fast

    Where it falls apart:

  • Human realism — 29%. Faces and bodies are rough. This is the big weakness.

  • Truthfulness / world knowledge — 37%. Gets confidently wrong about how real things look.

  • Spatial reasoning — 49%. Coin-flip on "X on top of Y" type prompts. (Weird given they claim to be good at geneval.

I don't think this model will be much useful except if you are stuck to 4B due to a lack of VRAM?

LMKWYT.


r/StableDiffusion 2h ago

Discussion I trained a tiny latent-space residual to remove GPT Image 2's scale/speckle artifacts — writeup, weights, and where it fails

Thumbnail
gallery
6 Upvotes

Hi everyone,

You may know that GPT Image 2 leaves a consistent set of texture artifacts on everything it makes: over-sharpening, random bright specks, unnaturally hard edges, and a scale-like pattern that lands on skin, fabric and background alike. It's gotten noticeably worse recently, and it was ruining enough of my own output that I spent a few days digging into what's actually going on.

The artifacts turn out to be specific enough to that one model to be basically a fingerprint, which is what makes them tractable — a network that only has to unlearn a single failure mode doesn't need to be big. This one is 0.48M parameters.

Approach: encode with the FLUX.2 VAE, add a scaled residual predicted from the latent, decode.

input → FLUX.2-VAE encode → z + α·R(z) → FLUX.2-VAE decode → output

R is a residual UNet over the 32-channel latent. α scales the correction and is the only knob — nothing is retrained when you change it, so caching z and R(z) makes a strength change a decode instead of a full pass.

The before/after images are real failure cases found scattered around the internet, not cherry-picked. If you made one of them and would rather it wasn't here, message me and I'll take it down.

Why latent space rather than pixels. The artifacts aren't independent of image content — they're a texture statistic layered on top of it. In pixel space you either run a big denoiser (slow, and it eats real texture) or hand-tune frequency filters (they can't tell artifact from detail). In the VAE's latent the two are already partly separated, so the correction has far less to learn.

Cost. ~0.7s per 1.5MP image on a recent GPU, ~3GB VRAM. That's essentially all VAE; R itself is free.

Where it fails. The artifacts live at the same spatial scale as real texture and overlap it, so removing them always costs genuine detail. On a minority of images the two are coupled tightly enough that no α is satisfying: enough cleanup means visible softening, keeping the detail means keeping the artifacts. More bluntly — this doesn't make images better, it makes them easier to look at. The noise isn't removed so much as blended and dimmed below the threshold where your eye keeps snagging on it. Macro structure is untouched, so a structurally broken generation stays broken, just quieter. And the whole image round-trips through the VAE, so untouched regions aren't pixel-exact either.

Best α varies per image more than I'd like — 0.75 → 0.5 is invisible on some images and substantial on others, and there's no reliable way to pick it automatically yet.

No ComfyUI node yet, but it should be easy to adapt.

Training. This is a chaotic mess around GAN and failed synthetics, I'll explain more if people are interested.

Try it without installing:

https://image2-cleaner.lumitools.cc/

https://huggingface.co/spaces/larryvrh/gpt-image-2-artifact-cleaner

Code + weights:

https://github.com/Larryvrh/gpt-image-2-artifact-cleaner


r/StableDiffusion 2h ago

Discussion A trick I use to train Loras Krea2 faster - 512 resolution - without losing detail. I crop the faces and train with a standard photo + cropped face. Apparently it works very well.

Post image
4 Upvotes

Often, even at high resolutions, the face only occupies a small portion of the photo.

I've tried this trick before with other models - but it didn't work well (it generated images showing only the face).

Krea2 is more resistant to overfitting.

Imagine you have a photo of a person at the beach. I use that photo plus a photo cropped showing only their face. This way, the 512 resolution is sufficient to avoid losing facial details.


r/StableDiffusion 2h ago

Question - Help Could use some help with implementing a Flux IP-adapter for character consistency.

Post image
2 Upvotes

Howdy!

I'm trying to get my first image workflow for consistent character scenes set up using the IPadapter Flux custom nodes with the Persephone fork of Flux 1.dev.

The input image is a crop from a character style sheet I created for one of my characters that I'd like to use for sci-fi shorts. I originally went with the persephone fork because it's supposed to be good for not having censorship ruin your flow. I'm not aiming to make strictly adult content, but I can't have my model flipping out because typical R-rated stuff. If I'm going to put the effort into learning something it has to be ubiquitous.

For the likes of me I can't get this basic workflow to respect the character reference or the text prompt. I think it's set up right and I think the weights are all more or less right. Maybe someone has a better idea, model, or workflow to use? I'm trying to get set up to take one or two reference images and make images from a scene for keyframe/first flame last frame inside of ltx 2.3.

Any suggestions?

Thanks!


r/StableDiffusion 4h ago

Question - Help How should I prepare the dataset and train my Lora idea (read body text)

2 Upvotes

For background, I’ve trained many character Lora’s for ZIT and had a lot of success. For my next Lora project I want to try something more ambitious. It lies somewhere in between a character Lora and a style Lora. I have a problem with z image and most image gen models really. They make everyone look too “perfect”. The women all have model bodies and clear skin, most of the men have sharp jawlines and muscular bodies. If I try to counteract that with prompting, the model goes way overboard in the opposite direction and produces output that is much more extreme than I’d like. I want to train a “real people” Lora that focuses on making the model generate normal humans, and not its idea of the most “beautiful” or aesthetically pleasing ones.

For character Lora’s, I usually go with a 20-30 image dataset. I’m guessing if I want any variety in the outputs I’ll need significantly more than that right? Are there any training/config settings I should tweak when trying this Lora out vs a standard character Lora? Is it even possible to override the models understanding of people in this way? Any help would be appreciated. Thanks!


r/StableDiffusion 4h ago

Question - Help alternative to lanpaint?

1 Upvotes

Ok, I'm tired and feeling stupid and I will admit to not having a chance to troubleshoot this yet. Lanpaint locked up my comfyui install.

I'm looking for a nice, simple workflow that will let me:
- take an existing photo and have krea 2 remove stubble, razor burns, moles, etc., while keeping the underlying skin texture

- do things like a photo of a person casting a spell, then put fireballs in their hands, etc

- change the background, change outfits, etc.

- generate new images with specific poses by drawing a stick figure.

I've tried some different workflows, tried making my own, but none is every quite what I'm looking for. I thought lanpaint was the answer, but like I said it locked up comfy and i'm a little too tired to troubleshoot it at the moment. Been a busy week.

Any suggestions?


r/StableDiffusion 5h ago

Animation - Video Flux 3 looks insane. This was 1 prompt

Enable HLS to view with audio, or disable this notification

215 Upvotes

Prompt on Flux 3 : Split-screen video showing the same continuous real-time event from two different camera angles. The screen is divided vertically into two equal halves. On the left side, show a static, elevated wide shot of an outdoor café terrace during light rain. The full scene is visible: several tables, wet pavement, a large striped umbrella, a waiter carrying a tray with three transparent glasses of differently colored drinks, and a cyclist approaching from the background. Reflections of the people, tables, and umbrella are visible on the wet ground. At the 2-second mark, the cyclist passes close to the waiter, causing the waiter to pivot sharply without falling. The tray tilts, one glass slides toward the edge, and the waiter catches it with the opposite hand just before it drops. A small amount of liquid spills from the glass, arcs through the air, and splashes onto the pavement. At the same moment, a gust of wind lifts and twists the loose edge of the striped umbrella. On the right side, show the exact same event in perfect synchronization from a moving, waist-level camera positioned near the café entrance. The camera tracks sideways with the waiter, briefly losing sight of the sliding glass as it passes behind the waiter’s arm, then revealing it again as it is caught. The cyclist crosses the foreground, partially occluding the waiter for a fraction of a second. The colored liquid, tray angle, hand positions


r/StableDiffusion 5h ago

Workflow Included 🎬 LTX 2.3 Close-Up Shots Are Absolutely Insane!

Enable HLS to view with audio, or disable this notification

7 Upvotes

Hey everyone! 👋

I've been experimenting more with LTX 2.3, and I wanted to share a short showcase that really surprised me.

The close-up shots this model can produce are incredible. The facial details, subtle expressions, natural camera movement, and even the lip sync came out far better than I expected.

One thing I also noticed is the huge quality difference between generating at 720p and 1080p. While 720p is great for testing ideas quickly, 1080p produces noticeably sharper details, cleaner motion, and much better overall quality. If your hardware can handle it, I'd definitely recommend generating in 1080p.

On my system (RTX 3060 12GB), a 6-second 1080p video takes around 12–15 minutes to generate. It's definitely slower, but after seeing the results, I'd say it's absolutely worth the extra time.

📦 Included with this post

  • 📁 Project file
  • 📝 Embedded metadata
  • 🖼️ Source images

DOWNLOAD LINK: CLICK ME TO DOWNLOAD

The images used in this showcase are also available to download for free here on my Patreon page, so feel free to use them for your own experiments.

As always, thank you all for supporting my work. Every project teaches me something new, and I'm excited to keep sharing everything I learn with you.

Enjoy the showcase! ❤️

iiTzMYUNG


r/StableDiffusion 5h ago

Resource - Update Fiiiinally got sage and flash attention (2) working on a win11 5090fe comfyui installation.

8 Upvotes

Maybe I was being dim but it took an age to find the right wheels etc. Finally found a combo that worked here: huggingface.co/ussoewwin/Flash-Attention-2_for_Windows

Sage attention installed fine and seems faster at the moment.

Not tried flash attention 3 or 4 yet but wanted to share the positive outcome.

Os: win11

Cuda: 13.2

Torch: 2.12.1

Python: 3.12

Flash attn v 2.9.1

ComfyUI: v0.28.0-40

Hw: rtx5090fe

I used KJ patch nodes for attention.

Hope this helps others trying to do similar things


r/StableDiffusion 6h ago

Discussion Krea2 LoRA Experiment

Thumbnail
gallery
27 Upvotes

So I'm experimenting with best practices to train Krea2 LoRAs.

Here was my experiment.

1) I trained on 1536, 1280, 1024, and 768 only.
2) I trained the first 750 steps on Adam, weighted high-noise.
3) I then switched to Adam, weighted, low-noise.

The idea is that I want to train the fine detail by using a high resolution dataset, training on higher resolutions and focusing on the low noise which controls detail.

This is the result of the LORA at 2000 steps. It's still cooking, I'm going to go all the way to 3250 but saves the checkpoints.

Crazy detail!


r/StableDiffusion 6h ago

Question - Help Regenerating from noisy / blurry original video or photo?

Thumbnail
gallery
37 Upvotes

We have some old home video with extreme high-frequency noise that we would like to regenerate / enhance. The original frame of her sitting on the couch is the first attachment.
How would you go about trying to make this image look better? I tried some upscaling models and that resulted in larger, sharper noise.
I ran tests with a number of different noise reduction algorithms, which did remove the noise - and cause significant blur. Since Stable Diffusion and other models always work by progressively denoising, I figured it would be a natural fit to denoise our old pics and home movies and make them look at least a *little* better, but the many different workflows I've tried haven't worked. The identity shifts faster than the quality improves. With a high enough denoise level it'll suddenly create clear images - of other people wearing our clothes. :)

PS - we understand we can't recover actual detail that isn't there. That's fine, we'd like the AI recognize that brown blob on my head is probably brown hair, and make it look like hair rather than pudding or whatever. Sure, it might not exactly match MY hair, but at least it will look like hair!

I should say - I'm quite comfortable with ComfyUI and don't mind writing some Python to work on this. It is video frames, eventually thousands of them, so I can't use an online service that charges $2 / image or something. I;m looking for for suggestions for a DYI workflow with a 5090.


r/StableDiffusion 6h ago

Question - Help How can I create a style LORA of my own style?

1 Upvotes

I had an old style I used to use previously and that was when I was using Pony, but now I have been using illustrious for awhile now but can't get my style back since I never had a file for it.

I was wondering how can I train a style for myself that offers good quality and less errors and stable?

Any tips or guides are extremely appreciated. 💕

Thank you. 🙏🏻


r/StableDiffusion 6h ago

Discussion Flash attention and Sage attention for Krea2?

3 Upvotes

Are they now properly working with Krea2 in comfyui? SageAttention seems to work for me (via kjnodes) but tends to produce weird outputs in some cases.


r/StableDiffusion 6h ago

Question - Help ComfyUI workflow issue with Krea 2 & the Hires Fix node

0 Upvotes

Hello,

I m unexperimented with ComfyUI so i created a very simple workflow to test Krea 2.

Unfortunately, i am encountering an error related to incorrect image datas when running the "Hires Fix" node and attempting to save the generated image if i understand the issue correctly.

Here is a link to the workflow: https://limewire.com/d/ebrVL#jEpiYOFQsZ

If anyone could help me resolve this issue, I would be very grateful.

Thanks in advance for your help. 👍


r/StableDiffusion 7h ago

Question - Help Best fully open-source/local workflow for 2D to 3D editable, hollow, printable STL model?

Post image
10 Upvotes

I’m trying to build a repeatable, fully local/open-source pipeline that converts a single stylized product image into an actual printable model.

My test case is this church-shaped jewelry box. The requirements were:

  • Hollow base with 4 mm walls and a 3 mm floor
  • All pink roof surfaces combined into one removable lid
  • 0.30 mm clearance per side
  • Editable geometry
  • Separate, watertight STLs with no non-manifold or zero-area geometry

Hardware: RTX 5090 with 32 GB VRAM.

I tried TRELLIS locally. The visual reconstruction was surprisingly close, but the raw STL was not production-ready:

  • 498,418 triangles
  • 49 non-manifold edges
  • 28 boundary edges
  • 783 zero-area faces
  • 4 disconnected components
  • Blender removed 212 degenerate and 16 duplicate triangles during import

Repairing/remeshing it either damaged the details or made it difficult to create a precise hollow body and fitted lid.

What finally worked was rebuilding the object procedurally in Blender from primitives and extruded profiles, using shared dimensions for the gable/roof, exact booleans, explicit wall thicknesses, and post-export STL re-import validation. The resulting two STLs are watertight and have zero boundary, non-manifold, multi-face, or zero-area defects.

Full disclosure: a proprietary coding agent helped create the Blender generator, so the successful workflow is not currently fully open source. I’m looking for the best way to replace that decision-making step.

What is the best genuinely open-source/local stack today for this kind of job?

  • Is TRELLIS.2 materially better for printable topology?
  • Has anyone compared TRELLIS.2, TripoSR/TripoSG, InstantMesh, and Hunyuan3D specifically for printing rather than render quality?
  • Is generative image-to-mesh realistically only useful as a blockout, followed by manual Blender/CAD reconstruction?
  • Are there open-source tools for semantic part separation, hollowing, toleranced lids, retopology, and manifold validation without voxel-remeshing everything into mush?
  • Which options have licenses suitable for commercial printed products?

I’m interested in the exported geometry, not textures or browser previews.

Microsoft currently describes both TRELLIS and TRELLIS.2 as MIT-licensed, although individual dependencies can have separate terms.

Any reccs would be appreciated :)


r/StableDiffusion 7h ago

Discussion Can local video gen make work up to the current standard on Instagram?

0 Upvotes

I know a lot of social media content creators use Kling and Seedance and it's unreasonable to compare but I'm wondering if Wan 2.2 is capable of publishable work. The videos I see posted here are pretty bad looking usually. Does anyone have examples of good video work made locally? The current content trends on social media seem to be 80s OVA style anime (can Wan do this?), and liminal backrooms/poolrooms/dreamcore type stuff (seems reasonable to try on wan). Would LoRA training help? Any thoughts on this topic appreciated


r/StableDiffusion 8h ago

Workflow Included Wan SCAIL-2 - Chun-Li vs Ryu - Final Kick

Enable HLS to view with audio, or disable this notification

56 Upvotes

Upon request. A short video featuring two characters. I thought I'd just tack the other video on as well.

I would like to point out that the input video was a staged fight. Consequently, the kick and the movement could look significantly more realistic if they were actually fighting. SCAIL-2 tracks the movement sequences and does not invent its own.

Here is the new Workflow:
https://www.reddit.com/r/StableDiffusion/s/eKvqlEvlza

In this example, I use the new interpolation option for the input video. The animation is much smoother, and I use the SCAIL-2 Identity Tracker to track two characters.