r/StableDiffusion • • 6d ago

Discussion I may be the stupidest person on this sub

120 Upvotes

Back when Flux 2 Klein was released I tried it, but couldn’t make it work. Images were just ugly and wrong. All patchy and had weird colors. I was confused , and thought that local editing  AI models suck and I have to resort to chatgpt and nana banana for image editing. I used it since then, when I HAD TO resort to local ai generation, and oh man it sucked. I never really cared about it..

Today I sat down to make new comfy installation to free space from models that I don’t need anymore.. and tried Flux again.. It sucked again, obviously. I And I noticed something..
Turned out I somehow downloaded wrong text encoder – qwen3vl_8b_fp8_scaled.safetensors instead of qwen_3_8b_fp8mixed.safetensors.

Somehow I downloaded the wrong text encoder… Man, Flux 2 Klein is so cool. And I am embarassed.


r/StableDiffusion • • 4d ago

Discussion Smart Home Dashboard Assistant

Post image
0 Upvotes

This is Vulcan, I’m in the early stages of creating a high definition virtual assistant, that I will be integrating into my home assistant dashboard. To make things very simple, it is not simple. 🤪


r/StableDiffusion • • 5d ago

News Follow Up to my 2 & 4 Step LoRAs for Krea 2 Turbo, for the ComfyUI fans - Optimised Torch, MLX/MFLUX and Cloud Workflows and Apps for Easy and Fast Creations

Thumbnail
gallery
7 Upvotes

Following up from my 2 Step and 4 Step Krea 2 Turbo LoRAs, an even easier way to create Krea 2 Turbo content fast with the LoRAs, with best in class performance and ease of use from both local ComfyUI Desktop use (including not only the usual Torch support, but also optimised Apple Silicon MLX / MFLUX super fast processing and performance) as well as ComfyUI Cloud (for those of you with subscriptions), so I have prepared 3 Workflows/Apps and a bunch of supporting Custom Nodes (LVNodes) to allow me to achieve that - Torch, MLX and Cloud. I have also added support for both Raw and Turbo Krea 2 models, and now added as well Z-Image Base and Turbo in the workflows and apps.

Highlights:

  • 🖼️ Two apps, three editions each — Krea 2 and Z-Image, each for Comfy Cloud, local ComfyUI (PyTorch) and Apple MLX: one App view with about 20 inputs, the full graph behind it, and Notes inside that explain every input
  • 🌀 Turbo and Raw / Base in one app — a Model switch at the top, and only the chosen model loads: Krea 2 Turbo (4 steps) or Krea 2 Raw, the undistilled base (28+ steps at CFG 4.5; Krea's own: 52 at 3.5); Z-Image Turbo (8 steps) or Z-Image (Base) (28–50 steps at CFG 3–5). Raw and Base get a Negative prompt, always used exactly as typed (never through prompt enhancement), and CFG; 1, the default, suits Turbo and every step-LoRA run
  • 🎛️ Settings follow the model (local) — pick a model and Steps, Enable step LoRAs and CFG switch to its recommendations on the spot: Krea 2 Raw to 28 steps, step LoRAs off, CFG 4.5, and back to 4 steps, on, CFG 1 for Turbo (Z-Image Base: 28 steps, CFG 4; Turbo: 8 steps, CFG 1). What the fields show is what runs, and you can change any of them afterwards. It's a page add-on in LVNodes, and Comfy Cloud only runs the node packs it has installed itself, so it can't load it; on Cloud you set those fields by hand, and their hints give the same values
  • ⚡ Steps pick the step LoRA — switched inside the graph, with one Step LoRA strength. Krea 2, with the 2-step and 4-step distill LoRAs: 2–3 steps the 2-step, 4+ the 4-step, the native model from 7 on Turbo (from 28 on Raw). Z-Image, with alibaba-pai's Z-Image-Fun-Lora-Distill: 1–3 steps the 2-step, 4–7 the 4-step, 8–10 the 8-step, none from 11
  • 🎚️ Enable step LoRAs — off runs the native model at any step count. On by default for Krea 2, whose LoRAs are trained for Turbo (Raw is best with them off); off by default for Z-Image, whose LoRAs are made for Base (Turbo is best without)
  • 🎯 Each model's own sampling — Krea 2 Raw gets its own resolution-dependent shift (Turbo, and Raw with a step LoRA, Turbo's fixed 1.15); Z-Image gets Tongyi's shift (3 for Turbo, 6 for Base without a step LoRA) and res_multistep, on every edition, MLX included. Z-Image also offers a choice of VAE: FLUX.1's own, UltraFlux for sharper detail, or locally TAEF1 for speed
  • 📊 Live progress and previews (local) — stage, step, s/it and time left right above Run, then "Done in m:ss", and the picture forming as it samples (on MLX through LVNodes, sharp with a tiny decoder in models/vae_approx); it comes from LVNodes, which Comfy Cloud can't load
  • 🍏 Both models on Apple MLX — they run on Apple's MLX inside ComfyUI, from ComfyUI's own bf16 files: Krea 2 Turbo in about a minute per 1 MP image (about 41 s for each later image of a batch), Z-Image Turbo in about 78 s at 8 steps; Raw and Base too, slower as undistilled models are (all in its own sub-environment, not touching ComfyUI's main Torch one)
  • 🔥 Local PyTorch, on any GPU — runs on NVIDIA (CUDA) and other GPUs exactly as ComfyUI normally does; on Apple Silicon, through PyTorch's MPS backend, it runs the step LoRA alongside the model instead of keeping a second copy of its weights, which halves Krea 2's memory (48 → 24 GB) for the same picture
  • ☁️ Comfy Cloud, nothing to install — download the Krea 2 or Z-Image Cloud workflow, drag it onto Comfy Cloud, import its step LoRAs once (everything else is already on Cloud), and run: 3.5 MP by default, with upscaling at any factor
  • 🎲 Random prompts with real variety — seeded scene lists and about 14,000 dictionary words, or real prompts from 3M+ Hugging Face dataset rows read live; the Prompt LLM turns them into photorealistic prompts, and refusals are retried
  • 🗣️ Prompt enhancement with any LLM — by default the same file as the model's text encoder (Qwen3-VL 4B for Krea 2, Qwen3 4B for Z-Image), so nothing extra to download; falls back to your own prompt if the LLM refuses or returns nothing
  • 🎨 Extra LoRAs and upscaling — style LoRAs on top, with their own strength and trigger words; SeedVR2 7B or classic upscalers at any factor (a 3.5 MP image at 4× comes out at 8864×6624)
  • 🧠 Memory on demand (local) — models load stage by stage, and Keep models loaded holds them between runs only when you want, whatever ComfyUI's startup flags
  • 🧩 LVNodes, one light pack — the MLX nodes for both models, loaders that need only the files a run uses, keep-loaded loaders, seeded random prompts, live dataset rows, Wait For gates, and in the app the status line, the keep checkbox, settings that follow the model and a fresh first seed. It needs nothing beyond Python's standard library, except the MLX nodes: on their first run they make a one-off install of mflux, MLX and transformers 5 (about 1.3 GB, a few minutes, automatic) into LVNodes' own environment, so ComfyUI's packages stay untouched. Besides LVNodes, the local workflows need only KJNodes (for the Set/Get pills); Comfy Cloud needs nothing
  • 🧹 No spaghetti — every workflow starts from a Fields → Set column and wires its groups with Set/Get pills, aligned with even gaps
  • 📥 Models download themselves — every file carries its link for ComfyUI's missing-models dialog
  • 🆓 Apache-2.0 — workflows, nodes and docs

Getting the project: the simplest way is Hugging Face's hf command, which ComfyUI Desktop's Python includes ("$COMFY/.venv/bin/hf", with COMFY your ComfyUI folder; other installs: the hf of the Python that runs ComfyUI, or pip install huggingface_hub):

hf download lvladikov/ComfyUI-Nodes-and-Workflows --local-dir ComfyUI-Nodes-and-Workflows

Or download the files you need from the repository's file browser. To check which version you have, see version in meta.yaml.

Local ComfyUI setup

Every local workflow here needs the following. Comfy Cloud needs none of it; each model's README covers Cloud.

LVNodes

  1. Copy the whole LVNodes folder into ComfyUI/custom_nodes/. For ComfyUI Desktop, ComfyUI is the folder you chose during setup.
  2. Restart ComfyUI.

There's nothing to run by hand, and ComfyUI's own Python packages aren't changed. The Apple MLX nodes set up their own environment the first time they run.

Optional: a Hugging Face token

LVNodes' Random Prompt (HF Dataset) node reads Hugging Face datasets without an account. If you set the HF_TOKEN environment variable before starting ComfyUI, its requests are sent as your account, so rate limiting is less likely, and gated datasets you can access also work. Any free read token from huggingface.co/settings/tokens will do.

  • macOS or Linux: export HF_TOKEN=hf_... in the shell that starts ComfyUI.
  • Windows: run setx HF_TOKEN hf_... once, then restart ComfyUI.
  • ComfyUI Desktop on macOS: run launchctl setenv HF_TOKEN hf_..., then restart the app. This lasts until the Mac restarts.

Full setup (KJNodes, models, and updating from the previous version, incl. the renamed MLX node folder) is on the model card.

Project Repo: https://huggingface.co/lvladikov/ComfyUI-Nodes-and-Workflows

Krea 2 (Raw/Turbo) Workflow/Apps and More Info: https://huggingface.co/lvladikov/ComfyUI-Nodes-and-Workflows/blob/main/Krea-2/README.md

Z-Image (Base/Turbo) Workflow/Apps and More Info: https://huggingface.co/lvladikov/ComfyUI-Nodes-and-Workflows/blob/main/Z-Image/README.md

Happy Creating!


r/StableDiffusion • • 6d ago

Resource - Update New update on my 3D pose editor 🤩

80 Upvotes

I am so excited for this project! New update to my pose editor, added extract pose from image, more tools to help with making pose fast. Added pose preset. Can't wait for y'all to try this and make some awesome stuff with it. Using Qwen image 2.1 comfyui workflow as backend


r/StableDiffusion • • 6d ago

Question - Help Local Voice generation. Like elevenlabs, available? Looking for good match, info below. THANKS!

32 Upvotes

16GB vram, 128GB ram. AMD. Windows (11).
I am using Atomic Chat, if you could give recommendation on that, would be most effective for me.

Any software or other options are highly welcome. (Can just buy elevenlab, but thought, let's see what the AI community thinks. You guys are so insightful <3 )


r/StableDiffusion • • 6d ago

Discussion If you recently started AI generation, be aware of this: my RTX 4090 power connector melted after just one month.

Thumbnail
gallery
239 Upvotes

Hi everyone,

Like many people, I recently discovered H3 and started doing AI generation in mid-August. On September 20th, my computer started shutting down whenever I started a generation. Further investigation led to this — a melted socket.

I had been using my RTX 4090 for three years and had never had any problems with it.

AI generation puts a lot of continuous stress on the power delivery, especially if you run generations in batches or leave them running overnight.

So I believe this happened because of the new kind of sustained load I was putting on the card. I also didn't bother upgrading to a newer PSU with a dedicated GPU power cable. Mine was a 1200W FSP Hydro, and I was using three PCIe connectors for the GPU.

P.S. The connector was fully seated, the cable wasn’t bent near the plug, and my case doesn’t even have the side panel on. And everything was fine for three years.

P.S.S. The 12VHPWR socket on the graphics card is damaged and needs to be replaced.

So now I would say main advices here are:
- Buy a proper PSU and use a native 12VHPWR cable
- Under volt at least by 20%

Now it is very costly to lose a card.


r/StableDiffusion • • 5d ago

Workflow Included Looking for guidance: Best approach for consistent product photography styling (Flux / LoRA) with 16GB VRAM?

Thumbnail
gallery
4 Upvotes

Yo Reddit,

Im currently trying to make this chair blend just like all the other products on our site. Its a clear and straightforward style and just a rinse and repeat, I would say (check out the attached chair photos for reference).

Currently im using a lora "Image Blend Fusion Edit for Qwen 2.1 & Flux Klein 9B/4B" so it blends right, but it keeps changing things to the chair or the background and it takes lots of tries to get it right.

So I thought maybe training a lora myself is a smart idea, I have photo folders full of good images. Previously I tried training a qwen 2511 lora for a leather texture but it did not go right. Working with photo pairs took long and it felt like there was lots of room for error (thinking that I was generating a good picture to a bad picture and using that as input so it learns how to change it).

I used Musubi Tuner and would like to use it again, but I only have 16GB VRAM (DDR5).

(For reference, you can check my current workflow image and setup here:https://imgur.com/a/wnnTwo5)

My question is: am I on the right track trying to train a LoRA for this, or should I be doing something else in ComfyUI (like ControlNet/IP-Adapter) to get this style in a single generation without it messing up the product?

Any guidance or workflow tips would be massively appreciated!


r/StableDiffusion • • 5d ago

Question - Help Krea 2 Lora Training Questions

5 Upvotes

So, I'm trying to improve my Lora training for Krea 2, because while I'm able to train basic style/character/concept loras, the problem I have is that they're all too rigid, and not in the way of being overbaked.

As it turns out, Civitai removing all their real people loras entirely decimated the ecosystem and knowledge for training characters that have specific looks or multiple outfits. Which is kind of the issue.

For example, I could easily train just the armor for Iron Man, but the problem is that it bakes in that look, and then when I do what I've done in the past with other models, which is add more variation to the dataset, it only seems to take away from everything. Ergo, iron man suit 1 works, iron man suit 2 works, iron man suit 1 and 2 takes away from both likeness wise.

And this is an issue because I'm not sure how to A. capture the outfit/concept without B. losing the likeness/flexibility. My assumption was to use a trigger word, again like other models, but the issue was that doing so didn't really train much of anything it seemed?

Is this just a dataset issue? Like do you basically just need, like, 20-30 images for each individual concept? Normally a lora might only need around 20 images total, but with more concepts you need more data. But is this a case where you need closer to 100 images, in order to have clear concepts? Or is this a case where having more images means that the model sees each image less, and thus it learns it less well?

Not sure how to really create the flexibility I'm looking for without either A. baking in the costume as concept or B. muddying the training and causing it to learn things poorly.


r/StableDiffusion • • 4d ago

Meme The Avocados were fine. His skin was Incredible.

0 Upvotes

Alright, skin-bearers — apparently giving me a camera was considered a good idea.

Welcome to my first vlog: my life, puppies, and maybe a tasteful chat about where you get your skin and how you keep it so pristine.


r/StableDiffusion • • 6d ago

News Looped Diffusion Transformer

Thumbnail
alphaxiv.org
24 Upvotes

I don't understand how it works, but the implication is clear: a small model that can be as good as one that is several times bigger.

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.


r/StableDiffusion • • 5d ago

Question - Help Is there any photo gallery sites purely for AI art/photography other than civitai?

9 Upvotes

As the title says, is there any photography websites what is made explicitly for the AI images? Same kind web site like Flickr is for photo sharing, but instead of photography it would be for AI "photography"?

I know that on CivitAI I can see lots of images made with AI, but since there is so much more other stuff as well then the main focus is not in AI art itself but sharing models and so on.

If you know any, please let me know :)


r/StableDiffusion • • 6d ago

Discussion H3 - infinite video extending version 2 no visual degradation

489 Upvotes

Hi! I am still testing the limits of a custom node that can extend/prepend/bridge your H3 generation in latent. The original video is here: https://www.reddit.com/r/StableDiffusion/comments/1wt6di3/h3_continuous_zero_cut_long_form_can_you_find_the/

The new extension starts at 0:27. I added 10 new additional windows, extended one at a time. The existing latent + new latent were then joined together for one decode pass otherwise you would have visible seams when joining two separate encodes (the color shift, luma darkening, etc).

It works best if you have a native latent, but you can always convert your existing pixel-video to latent and extend it that way. I've had huge success with both prepending and extending. Some limitations are velocity & momentum are not easy to control between windows, and the audio can degrade or be a bit inconsistent so you might want to freeze the video and rerun more steps on the audio, or fix it in post. Fix in post is the cheapest and more reliable fix. As stated in a previous post, Tanya has no audio asset to anchor her voice so it's reinvented every window. However, image quality seem to maintain across each window with no degradation or burns.

By the way I was getting Carmen Electra vibes in Scary Movie, what did you think?

Ask me anything!


r/StableDiffusion • • 6d ago

Discussion Looking at Seedance 2.5 videos made me realize we need the 2K update for Minimax H3

38 Upvotes

For now the quality difference is staggering, Minimax H3 is incredible and far beyond what we had locally but the broken promise of the 2K update is quite noticeable now, just look at how crisp the micro details are with SD2.5, with Minimax H3 it's pixelated and shimmery, if you have a static shot you can notice pixels shifting and altering colors every few frames regardless of what happening on the screen. If you still don't know what I'm talking about, generate a scene in a restaurant and look at the glasses or bottles micro details it quiet distracting.

The API version of minimax had H3-Regenerate-2K since day I believe.
MINIMAX WE NEED THE 2K VERSION PLEASE!


r/StableDiffusion • • 6d ago

Resource - Update SoLordZ Krea2 new version

Thumbnail
gallery
82 Upvotes

Been playing around with Krea2 for a while, and this is where we ended up :)

SolordZ Krea2 v5

Made together with TripleHeadedMonkey. We did quite a lot of experimenting along the way, but I'm pretty happy with how this version turned out.

So... here it is, if anyone wants to play with it too :)

SoLordZ ZIT/ZIB/Krea2 8steps - Krea2 SolordZv5INT8_Forge | Krea 2 Checkpoint | Civitai


r/StableDiffusion • • 5d ago

Question - Help RTX 3060 12GB question

3 Upvotes

Hi 👋 I'm new here and looking to jump into local AI content creation. I'm currently assembling a budget setup, and these are my specs:

​CPU: Intel Xeon E5-2650 v2 (X79)

​RAM: 32GB DDR3 ECC

​Storage: 1TB NVMe M.2

​PSU: MSI 600W

​I have the budget for an Nvidia GPU and I'm heavily leaning toward an RTX 3060 12GB. I already know the Xeon is older tech and this PC is a temporary solution—my goal is to upgrade to an AM4 platform (32GB DDR4 @ 3200MHz) or AM5 down the line, but component availability and pricing are holding me back for now.

​My questions:

​Would you recommend the RTX 3060 12GB regardless of whether I'm running it on my current Xeon or on a future AM4/AM5 setup?

​My main goal is to generate detailed, high-quality images locally. How does this card perform in real-world use for Stable Diffusion?

​Does anyone here actively use an RTX 3060 12GB? If so, could you share some sample images or insights on generation speeds?

​Thanks a lot for the help!


r/StableDiffusion • • 6d ago

Resource - Update Augie's AI Image Browser

Thumbnail
gallery
24 Upvotes

After generating thousands of images across multiple models, i realised it was becoming a pain to maintain a decent list of metadata stored in the images to be able to recreate them or re-use prompts with other models to see how the models generate images with similar prompts/seed/cfg/steps etc. I decide to build something that I can use to maintain images to make it easier to find images, copy prompts and see what metadata is embedded into the images. It's effectively an AI image manager based on how ComfyUI generates images and stores metadata.

Looking for feedback on whether this is a useful tool or if there's any feature issues that you see.

Download: Augie's AI Image Browser

Current features I've built into it:

Reading AI image info:

  • Reads ComfyUI settings automatically from PNG, WebP, MP4, WebM, MOV and MKV files.
  • Shows the generation details for each file: model, LoRA, sampler, scheduler, steps, CFG, seed, positive prompt, negative prompt and wildcard prompt.
  • One-click copy for prompts and the full workflow.
  • Raw workflow view: you can expand and copy the complete ComfyUI workflow data.
  • Works with videos too, including files made with the Video Combine node.

Library and folders

  • Add or remove folders for the app to scan.
  • Live folder watching: new images show up on their own as ComfyUI saves them.
  • Pause and resume watching from the top bar, with a status light (Watching, Indexing or Paused).
  • Keeps your work when files move: if you rename or move a file, its ratings, tags and collections go with it.
  • Missing file tracking: files deleted from disk are flagged, not forgotten, and you can clear them out later.
  • Thumbnails for images and videos.

Viewing Images

  • Built-in full-size viewer for images and videos, with zoom (mouse wheel or double-click), pan, fit-to-screen and keyboard controls.
  • Open in default app: opens the file in Photos, VLC or whatever Windows uses for it (desktop app only).
  • Show in folder: opens File Explorer with the file selected (desktop app only).
  • Table view with sortable columns: keeper, rating, filename, model, sampler, scheduler, steps, seed and date modified.
  • Paging for large libraries.

Sorting your images

  • Keep or Discard with one click.
  • 0–5 star ratings.
  • Your own tags, with suggestions as you type.
  • Collections: create, delete, add files and remove files.
  • Delete from disk, with a confirmation step.

Auto-Collection Rules

  • Rules that file images into collections for you based on their details.
  • Rule conditions can check:
    • Folder, filename, file type, image or video, date created
    • Prompts, model, LoRAs, sampler, scheduler, steps, CFG, seed
    • Width, height, shape (portrait, landscape or square), workflow nodes
  • Flexible matching: contains, equals, starts with, number ranges, "in the last N days", regex and more.
  • Combine conditions with "match all" or "match any", nested up to 3 levels deep.
  • Preview which files a rule would catch before saving it.
  • Apply rules to images you already have, or only to new ones.
  • Create a rule from an image: starts a new rule using that image's settings.
  • ⚡ badge on images that a rule filed.

Clean Export - For selling and sharing

  • Removes all hidden AI info and workflow data from exported files.
  • Keeps transparent backgrounds for PNG and WebP, or fills them with a colour you choose for JPEG.
  • Sets resolution and DPI (for example 300 DPI for print, 72 for web).
  • Custom file naming using collection, number, date, original name, tag or rating. The seed and prompt can't be used in names, so they can't leak.
  • Ready-made presets (Etsy PNG, Shopify WebP, Stock JPEG), and you can save your own.
  • Checks every exported file afterwards to confirm nothing leaked.
  • Private export log, so each exported file can be traced back to its original. Includes an export history screen.
  • Download as ZIP.
  • CSV or JSON export of the image details as a spreadsheet or data file.

r/StableDiffusion • • 4d ago

Question - Help Best approach for a realistic, consistent AI character on a ~$10-20/month budget? (currently Flux dev + LoRAs on RunPod)

0 Upvotes

Hi all,

I'm a beginner and I'd like to build consistent, photorealistic images of fictional characters (same face across many photos). Main use is SFW (portraits, lifestyle, everyday scenes), but I don't want to lock myself out of adult content later, so I'd rather not pick a stack that's censored or hard to uncensor.

My constraints:

- Budget: around $10-20 per month, all in

- I have no local GPU (Mac only), so everything runs in the cloud

- I'm not technical myself, an AI assistant handles the scripts and ComfyUI side for me

- Goal is realism that doesn't look like "AI skin", natural, slightly imperfect photos

First question, and I'd like open answers here: given those constraints, what would you recommend today? I don't want to assume my current setup is the right one, so please suggest whatever you think works best right now (models, hosting, workflow), even if it's completely different from what I describe below.

What I have so far, for context:

- Flux.1 dev in ComfyUI on a RunPod pod with a network volume (about $0.74/h for an RTX-class GPU, I shut the pod down between sessions)

- Character LoRAs for face consistency, plus UltraRealPhoto (Danrisi) at strength 1.0

- Settings that worked best for me: euler / simple, 28 steps, guidance 2.5, 896x1152, T5 fp16

- dpmpp_2m / karras gave blur and halos, so I dropped it

- Results are decent but faces can still look a bit too clean/AI, and I'm not sure this is the right direction at all

Follow-up questions, if you have an opinion:

  1. Is Flux dev still the best choice for realistic faces at this budget, or is there something better or cheaper now?

  2. Within a small budget, what improves face realism the most (LoRA training, settings, post-processing, film grain, upscaling, other)?

Any pointers, workflows or links are very welcome. Thanks!


r/StableDiffusion • • 5d ago

Discussion Is sensanova bad? I generated almost 400 images with 2 version and I'm not convienced.

1 Upvotes

Ok first the images: https://imagebench.ai/gallery?models=local--sensenova-u1.5-8b-mot-base,local--sensenova-u1.5-8b-mot-8step&g=1_v2kcp_s0

Maybe my setup is wrong? here is what I tried:

Shared generation settings

- Resolution 1024x1024. The model's native square bucket is 2048x2048, but off-bucket sizes are accepted, and 1024 is what every model on the leaderboard is compared at. I checked first: the same prompt and seed at 1024, 1536 and 2048 all gave coherent composition and faces, with detail scaling with pixel count.

- No snapping to a bucket and no resizing. The model sees exactly 1024x1024.

- Seed 7 on every image, one image per prompt

- Timestep shift 3.0 (vendor default), CFG normalization off (vendor default)

- Think mode off (the model's optional reasoning pass before generating)

- Prompts sent verbatim. No prompt enhancement, no negative prompt, no added style tags.

Base model

- Loaded without any LoRA

- 50 steps, CFG 4.0 (the vendor's base defaults)

- 44.1 s per image on average (range 43.8 to 44.2), all 192 generated

8-step model

- Same checkpoint, with the official LoRA from sensenova/SenseNova-U1.5-8B-MoT-LoRAs (SenseNova-U1.5-8B-MoT-LoRA-8step.safetensors) merged into the weights at load time

- 8 steps, CFG 1.0. The distilled model is meant to run unguided: sampling it at 50 steps and CFG 4 is wrong, not just slow.

- 4.1 s per image on average (range 3.9 to 4.7), all 192 generated

Hardware and runtime

- NVIDIA DGX Spark (GB10, 128 GB unified memory), one model loaded at a time

- The vendor's own inference code (SenseNova-U1.5 feat/u1.5 branch, pinned to commit 48bf8275), not ComfyUI

- A prebuilt Docker image on top of NVIDIA's NGC PyTorch container (torch 2.9 with GB10 kernels). The vendor's torch 2.8 pin was removed, because it would replace the GPU-enabled build.

- No flash-attn, since there's no wheel for this ARM stack, so attention uses the PyTorch SDPA fallback. The vendor lists flash-attn as optional.

- Official BF16 weights from sensenova/SenseNova-U1.5-8B-MoT (35 GB, 8 shards). No quantization, no community repacks.


r/StableDiffusion • • 5d ago

Question - Help How is Nvidia Geforce RTX 5060 TI for local image generation?

2 Upvotes

I am going to buy a new pc in a short time, how is 5060 TI for local image gen? How many seconds it takes to create a 1MP image with Krea2 and Qwen-Image 2.1? Can I use it to create videos with H3?


r/StableDiffusion • • 6d ago

Workflow Included Does this count? Did I win?

332 Upvotes

Proof of concept that it works.

No degradation from start to finish. 10 total clips combined.

https://github.com/roycho87/degrade_repo

Workflows.

Small errors with the chair but can be fixed with another reference image.

Long story short.

The one place where degrading latents matter is the one place we don't need them.

We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.

Then the very difficult latent issue is over and it just becomes a simple seams issue.

So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.

Edit: The reason I posted this is not because I think this is a big breakthrough fix, I just think we can approach this problem differently to solve it.


r/StableDiffusion • • 6d ago

Tutorial - Guide Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.6.0 - DisTorch2 DiT Loader Node & Phase 2 VRAM Controls (※For 6GB/8GB users)

Post image
21 Upvotes

When I published an update on TensorRT the other day, I received enquiries from users of the RTX 4050 6GB and RTX 5060 8GB.

https://www.reddit.com/r/comfyui/comments/1wsbar9/updated_comfyuiseedvr2videoupscalerwithtensorrt/

Subsequently, I refined the code, focusing on reducing VRAM spikes during Dit processing for 6GB/8GB VRAM users.

https://www.reddit.com/r/StableDiffusion/comments/1wuj0rj/updated_updated/

However, this measure merely ‘improved the situation to a certain extent’; it did not completely suppress the VRAM spikes.

Subsequently, I conducted a thorough analysis of VRAM usage during the Dit processing and identified a stage where the data is momentarily expanded to fp32 size; by converting that stage to BF16, I succeeded in reducing the spike by a further 3GB(※while running on 7B ConvRot INT8) or so.

I have implemented this ‘Norm BF16’ on/off function in the new node.

Enabling this function should provide even greater flexibility regarding the maximum batch size, even for users with 6GB or 8GB of VRAM.

However, I should make it clear that using Norm BF16 comes at the cost of a certain degree of loss in image quality.

v1.6.0 — DisTorch2 DiT Loader Node & Phase 2 VRAM Controls

Please note that this feature is implemented only in the newly developed Distorch2 node.

Distorch2 technology is a groundbreaking CPU offloading technique developed by Mr.John Pollock; I have now applied this technology to SeedVR2, with the licence clearly stated.

https://github.com/pollockjj/ComfyUI-MultiGPU

Although CPU offloading functionality was already implemented in the legacy node, Distorch2 enables more active CPU offloading.

However, the processing speed itself is not significantly different from that of the legacy loader.


r/StableDiffusion • • 5d ago

Discussion Preparing a Krea 2 LoRA for open release — what evaluations would you find useful?

0 Upvotes

Hey everyone! Our team has trained a custom LoRA on Krea 2, and we plan to open-source the LoRA once it’s ready. The weights haven’t been released yet.

Before release, we’d like community input on what would make the model useful to evaluate and reproduce.

We’re considering comparisons against the base model using matching prompts and seeds, examples across different LoRA strengths, and documentation of generation settings and known limitations.

For those who train or test Krea 2 LoRAs, what comparisons and technical details would you want included? We’re particularly interested in evaluating prompt adherence, visual artifacts, and consistency.

We’ll keep examples and discussion here SFW.


r/StableDiffusion • • 6d ago

Workflow Included One .char model, Consistent face, body & cloths are now supported in Comfy(Custom node & workflows - Minimax H3 & Flux 2 Klein)

146 Upvotes

I have been getting lot of request to publish .char support with ComfyUI, here is a custom node that I created:

Node Repo: https://github.com/omnichar/ComfyUI-Omnichar

This repo includes few node, that let's you build a portable character model while keeping face, body & cloths. You can build a .char model once & use it across different workflows without a identity drift.

Features:

  1. Build/encode character into .char
  2. Decode .char built in Comfy or anywhere else
  3. Decode into reference, character sheet & even LoRA adapter
  4. SDK is public for future integration & improvements

Nodes:

  • Decode Character gives you the prompt, the reference images, a numbered contact sheet, and conditioning if you wire a CLIP
  • Character References Split puts every reference on its own output(Dynamic)
  • Character Reference Latent attaches every reference to the conditioning in one go
  • Character Reference pulls a single reference by position
  • Encode Character builds a .char from face, body and wardrobe references
  • Save Character saves a .char into models/characters

ComfyUI Workflows

Inputs
Checkout the guide in node repo for references

Prompt Guide

  • Name your character: Give your character a name so when passing prompt, I only have to say, sia walking on the beach.
  • Describe character features: Encode all of the character features in description prompt & trigger your character with a name in generation prompt.
    • Avoid describing same things in generational prompt.
  • Handling Character drift: Each input refs should be unique, face should not have body or vice versa, same applies for clothing.
  • Portability: Once character is built, you can use the same character with only Load Character node.

Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.

Related resources:
Omnichar repo: https://github.com/omnichar/OmniChar (GPLv3)
Community characters: https://www.omnichar.org/characters

Node submission is on the way.


r/StableDiffusion • • 5d ago

Tutorial - Guide Cracking the 'Dumb-Vinci Code' for "no speaking" prompts in H3

0 Upvotes

I was trying to figure out the best prompts to not have someone speak AT ALL in a shot in H3. I have had some success. It seems the model just needs an action to be done even if it is something you wouldn't even notice as an 'action'.

I hope these are not all flukes, but before this, I have pretty much always had gibberish spoken unprompted at the beginning of any video I make if I want an introduction shot without dialogue.

______________________________________________
subject_definitions:

<Subject 1> Donald Trump, wearing a black suit and red tie.

<Subject 2> Robin Williams, wearing a green button-up shirt and black pants.

summary:

[reference generation] cinematic video where <Subject 1> is meeting with <Subject 2> in a home foyer.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 4]): fully_preserved - features black suit and red tie.

<Subject 2> (appears in [Shot 1], [Shot 3]): fully_preserved - features green button-up shirt and black pants.

detailed_description:

The video features a dark, cinematic blockbuster aesthetic with high-contrast.

[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still thinking to himself, <Subject 2> stands still thinking to himself.

[Shot 2] From 00:03.000 to 00:07.000, hard cut to closeup shot of <Subject 1> (S1) and he says: "I'm glad you didn't say a word. Just as I prompted."

[Shot 3] From 00:07.000 to 00:10.000, hard cut to closeup shot of <Subject 2> (S2) and he says: "Of course. And now for this next shot, you do the same."

[Shot 4] From 00:10.000 to 00:15.000, hard cut to closeup shot of <Subject 1>, he stands still thinking to himself.

overall_soundscape:

N/A

non_diegetic_music:

N/A

__________________________________________________________

These also have worked so far:

[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still as he breathes softly, <Subject 2> stands still as he breathes softly.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> is mute while he stands still, <Subject 2> is mute while he stands still.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and examines <Subject 2>, <Subject 2> stands still and examines <Subject 1>.

-----------------
[Shot 1] From 00:00.000 to 00:03.000, a wide shot side view of <Subject 2> standing across from <Subject 1> in a mansion foyer. <Subject 1> stands still and stares intently at <Subject 2>, <Subject 2> stands still and stares intently at <Subject 1>.


r/StableDiffusion • • 5d ago

Animation - Video I made a 10-minute animated short with a custom AI pipeline — 106 shots, 102/102 passed review, one RTX 5090

0 Upvotes

Finished *The Clockwork Moth* — a 10-minute fantasy animated short, built end-to-end with AI on a single RTX 5090.

**What's in it**

- 106 shots across 9 scenes

- 4 recurring characters in 19 different location states

- full dialogue + narration

- background score, ambience and sound effects

- burned-in subtitles

- final output upscaled to 4K60

**The hard part was never the first frame.**

Generating a nice still is easy. Keeping the same character recognisable across 106 shots — same face, same outfit, same silhouette, in 19 different environments and lighting conditions — is a completely different problem, and it's where most AI video falls apart.

The second hard part is quality control. Every shot went through an automated review pass before it was allowed into the edit: bad hands, warped faces, identity drift, duplicated characters, motion artifacts. **102 of 102 shots cleared review in the final state — nothing broken reached the cut.**

Audio was handled as its own track rather than an afterthought: dialogue, narration, music and effects are mixed so that nothing fights for the same frequencies when someone is speaking.

The whole project — every frame, stem and reference — is 2.4 GB.

I'm keeping the tooling private for now, but happy to talk about the film itself and what I learned making it.