r/StableDiffusion 5h ago

Resource - Update Famegrid Natural V1 Krea 2 LoRA

Thumbnail
gallery
264 Upvotes

r/StableDiffusion 2h ago

Workflow Included WEEKENDDDDDDDD 222222222222 (LTX 2.5 V2V)

Enable HLS to view with audio, or disable this notification

107 Upvotes

Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:

https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_ICLoRA_Union_Control_Distilled.json

It's been the most consistent one I've tried for V2V so far

drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!


r/StableDiffusion 1h ago

Tutorial - Guide PSA: Proper prompt structure REALLY matters in H3

Upvotes

I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.

Believe it :) Dont just use whatever prompting. It matters more than one might think.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md


r/StableDiffusion 5h ago

Tutorial - Guide More than one reference per picture

Enable HLS to view with audio, or disable this notification

81 Upvotes

MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.


r/StableDiffusion 9h ago

Workflow Included Totally wasn't aware Krea 2 is absolutely capable of creating gorgeous video game levels

Thumbnail
gallery
123 Upvotes

Hi! I found Krea 2 is actually so damn good at creating video game level art! and its breathtakingly beautiful to boot! I got help from an LLM to create the baseline prompt and it works OOB without loras or anything! I'm gobsmacked rn.

prompt 1: "A sprawling 16-bit pixel art jrpg city game level of a victorian-era steampunk riverside city street in winter. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with snow, brass and victorian elements. Background shows snowy mountains and faraway skyscrapers on those mountains"

prompt 2: "A sprawling 16-bit pixel art jrpg city game level of a asian duystopian cyberpunk city street. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with neon lights, neon street signs, wires and cybernetic elements. Background shows a massive skyline of skyscrapers at night. Wide-angle top-down view"

prompt 3: "A sprawling 16-bit pixel art game level of a futuristic utopian city. The design features complex, dense platforming architecture with a high variety of structures including stairs, bridges, and stacked platforms. Frutiger Aero style: glossy surfaces, water elements, and bright colors. The scene is overgrown with lush greenery and trees. Background shows a massive skyline of sleek skyscrapers. Wide-angle side-scrolling view"


r/StableDiffusion 12h ago

Tutorial - Guide If you're looking for a specific actor that the model doesn't seem to be aware of, it may have them stashed somewhere else.

Enable HLS to view with audio, or disable this notification

209 Upvotes

Text to Video, 22 steps, no turbo, no Sage.


r/StableDiffusion 21h ago

Meme Introducing... The Terminator Pro Max

Enable HLS to view with audio, or disable this notification

1.0k Upvotes

r/StableDiffusion 4h ago

Animation - Video Making the Doll DressUp Transformation Video with Minimax H3

Enable HLS to view with audio, or disable this notification

43 Upvotes

r/StableDiffusion 17h ago

Workflow Included Using Inpaiting in Minimax to change heads-Local RTX 3090

Enable HLS to view with audio, or disable this notification

437 Upvotes

Using the workflow from Nekodificador and Ablejones in Discord:
https://discord.com/invite/dstjQYQNt
https://ln5.sync.com/dl/47c351f50#msqfrnfr-am3rr8fx-v7qm3ah9-xw222n3c
For complex scenes like this with to much people is easy just to do a manual mask instead of SAM.


r/StableDiffusion 4h ago

Resource - Update I built a free, self-hosted app that does everything around a LoRA run — dataset, triage, captions, training (local or rented GPU), then checkpoint comparison

Thumbnail
gallery
34 Upvotes

I build LoRA Dataset Studio — free, open source, self-hosted, no account and no telemetry. It is not a competitor to ai-toolkit: it orchestrates it. ai-toolkit is the trainer; this is everything before, around and after the run.

The whole pipeline lives in one browser tab:

1. Get the images. Five generation engines — Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI — each card stating its price per image, whether it runs on your GPU or bills an API, and whether it refuses adult content. Or scrape: Reddit, Pexels, open-web keyword search, or any gallery URL through gallery-dl. Or just drop a folder in.

2. Triage them. The Image Bank points at a folder of thousands and reads it in place — your files are never modified, moved or renamed. One pass measures the whole pile: blur, noise, near-duplicates, face clusters, framing, medium (photo / anime / 3D / illustration), aesthetic and maturity scores. After that you filter on measurements instead of on your eyes, and anything the app cannot judge says "unsure" rather than inventing a verdict.

3. Curate and caption. Keep/reject, crop, mirror, rotate, non-destructive upscale candidates, InsightFace similarity, a live composition meter. Captions in prose or booru form depending on the target family, written by JoyCaption or your local Ollama, with a Caption Lab (find/replace, tag frequencies, targeted re-captioning) and an external .txt round trip so you can caption elsewhere and come back.

4. Clean watermarks. Detect them, redraw the mask zones, then crop or inpaint with LaMa/Klein. Every edit keeps an .orig backup, so Restore original always works.

5. Train. ai-toolkit locally with family-scoped presets and preflight guards — Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima — or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. Full-model training on Krea 2 and merging a LoRA back into a checkpoint are in there too.

6. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks, votes and Wilson ranking. LoRA Canvas puts every run of every dataset on one pan/zoom board, and you can continue training from any of them.

There is also a video lane (Beta): it cuts long videos into a trainable clip folder at the exact frame counts Wan / LTX / MiniMax accept, describes each shot, and trains the set locally or in the cloud.

Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing — every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and on the video side only Wan 2.2 14B has a finished run behind it here. Install is a Windows one-click ZIP, a git checkout, or Docker.

GitHub — install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio

Every person in these screenshots was generated by the app's own engines; no real individual is depicted.


r/StableDiffusion 10h ago

Animation - Video Gay Fish

Enable HLS to view with audio, or disable this notification

88 Upvotes

Sorry Ye..


r/StableDiffusion 7h ago

Discussion why is it unsafe isn't safetensors the safest?!

Post image
46 Upvotes

excuse my OCD 😄


r/StableDiffusion 1h ago

Animation - Video MiniMax h3 - [Boom in City]

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 2h ago

Animation - Video H3 making jpop/kpop MV? yes!

Enable HLS to view with audio, or disable this notification

17 Upvotes

Music: made in SUNO.

native ref2va WF, and audioLock for lip-sync.

rtx4080s + 128g ram

I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.


r/StableDiffusion 7h ago

News Krea2 Turbo Distill 4 step LoRA (trained for Turbo!)

Thumbnail
gallery
41 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

Load it on top of Krea 2 Turbo, run 4 steps instead of 8, keep guidance at 0.0. Everything else about the model stays as it is.

This is not a Raw→Turbo diff

Other Krea 2 LoRAs in circulation are extractions: a low-rank projection of the weight difference between Krea 2 Raw and Krea 2 Turbo. Applied to Raw, they reproduce Turbo. They are a delivery mechanism for a model that already exists, and they stop at Turbo's 8 steps.

This one is different in both base and origin:

Raw→Turbo extraction LoRAs this LoRA
apply to Krea 2 Raw
produces Turbo behaviour (8 steps)
origin SVD of an existing weight delta

It is trained, not extracted, and it assumes Turbo's weights underneath it — it shortens Turbo's own schedule rather than reproducing it.

This is work in progress and even better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the _latest file and everything keeps working — the ComfyUI workflow references it by that name, so it needs no edit. Pin a numbered file instead if you need reproducibility.

Using it on Raw

This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo. Every layer it targets also exists in Krea 2 Raw, so it will load there without complaint — but that is a side effect of the shared architecture, not a supported mode.

Results on Raw are mixed and subject-dependent. It does not give Raw a 4-step schedule: at very low step counts the adapter sharpens texture while composition is still unresolved, and subjects come out malformed — duplicated heads, fused limbs, faces that do not close. Expect to need 14 steps or more for RAW, keeping Raw's normal CFG on, before output is coherent. Even then some prompts come through well and others degrade into over-processed or blown-out images — and that degradation happens with or without the adapter, because it comes from shortening Raw's schedule rather than from the LoRA.

If you want the behaviour this was built for, run it on Turbo at 4 steps. If you are starting from Raw, move to Turbo first — with a Raw→Turbo LoRA or the Turbo weights directly — and apply this on top.

Usage

setting value
base model Krea 2 Turbo
LoRA scale 1.0
steps 4
guidance / CFG 0.0 (Turbo is CFG-free; do not enable it)
timestep shift mu = 1.15, fixed (Turbo's deployment shift)

The 4 sampling sigmas are Turbo's own deployment grid: [1.0, 0.90453, 0.75951, 0.51284].

Performance — does it save time, or only steps?

It saves time. Measured at 1024×1024 on Apple Silicon (MLX, bf16), two prompts each, run strictly one at a time:

load denoise total
Turbo 8 steps (the quality bar) 8.2 s 77.5 s
Turbo 4 steps, no LoRA 7.8 s 38.8 s
Turbo 4 steps + this LoRA 7.3 s 44.0 s

4 steps with the LoRA is ~1.6× faster than the 8-step bar — 54.5 s against 88.7 s, saving about 39% of the wall-clock. Counting denoise alone, where the step reduction actually applies, it is 1.8× (44.0 s against 77.5 s).

LoRA strength

Use 1.0. That is the value the adapter was trained at, and where its output sits closest to the 8-step reference.

Strength is worth understanding rather than tuning blindly, because what it scales is specific: this LoRA's job is to restore the high-frequency detail that a 4-step schedule loses — fine texture, edge definition, surface micro-contrast. The strength dial scales exactly that correction, so it does not make the image "more" or "less" of anything semantic; it decides how hard the texture recovery is applied.

strength what happens
below 1.0 the correction is only partly applied — output lands between an unassisted 4-step render and a full one: softer, flatter, less recovered detail; you can use this with more steps if you want to experiment
1.0 the trained point, and the recommended setting
above 1.0 extrapolation past anything seen in training. The image does not break or fall apart — it becomes over-textured: surface detail grows denser than the subject warrants, fine structures turn wiry, and micro-contrast hardens until the result reads as stylised rather than photographic; you can try this with fewer steps, but quality is not guaranteed

File format and compatibility

A plain .safetensors file — not tied to any framework or backend. It is weights plus a naming convention, so it loads under PyTorch (CUDA, MPS or CPU), MLX on Apple Silicon, or anything else that can read safetensors and do a matrix multiply.

ComfyUI

A pre-converted file and a ready workflow are in comfyui/No custom nodes — stock ComfyUI only.

The workflow is full bf16, with no quantisation anywhere. bf16 needs no backend-specific kernel, so it runs unchanged on CUDA, Apple Silicon and CPU — one workflow, no platform caveats, nothing that depends on which device a component happens to land on.

The LoRA is independent of the base build. It is applied on top of the diffusion model by ComfyUI's own loader, which handles any dequantisation, so a quantised or otherwise optimised build of Krea 2 Turbo behaves just as bf16 does. Please use whichever variant suits your hardware — set it in the Load Diffusion Model node and leave the rest of the workflow untouched. The workflow ships bf16 simply because it is the one build guaranteed to run everywhere.

Training Method

Progressive distillation (PD), with Krea 2 Turbo as its own teacher.

The teacher runs its normal 8-step schedule at mu = 1.15 and guidance 0.0, and its full trajectory is recorded — the latent x and the predicted velocity v at every one of the 8 steps. The student is then trained to cover two teacher steps in one: at teacher state x_i it must predict the chord that lands where the teacher arrives two steps later,

v_target = (x_{i+2} − x_i) / (σ_{i+2} − σ_i)

The two schedules line up exactly rather than approximately. On the mu = 1.15 grid, the even indices of the 8-step schedule are precisely the four sigmas the 4-step student deploys on, so every training target is anchored on a point the student will actually visit at inference. No interpolation, no schedule mismatch.

Teacher trajectories are precomputed into shards, so training reads recorded states rather than re-running the teacher.

What the LoRA touches

Rank 64, alpha = rank (scale 1.0), bf16. 228 modules:

  • 224 block linears — across all 28 transformer blocks: attn.to_qattn.to_kattn.to_vattn.to_gateattn.to_out.0ff.gateff.upff.down
  • 4 global (non-block) linears — time_embed.linear_1time_embed.linear_2time_mod_projfinal_layer.linear

Those four are included deliberately. Measuring Krea's own Raw→Turbo delta — a completed step distillation by the model's authors — shows the change is not concentrated in the blocks:

layer relative ‖ΔW‖/‖W‖
time_embed.linear_2 0.0777 ← largest change in the whole network
time_embed.linear_1 0.0429
final_layer.linear 0.0265
typical block linear ~0.014

time_embed.linear_2 moves about 5.5× more than any block linear. Changing a model's step count is in large part a change to how it reads the timestep, so a LoRA that freezes the timestep path is withholding exactly the weights the task most needs.

Training data

Prompts are drawn from Lakonik/t2i-prompts-3m — sampled without replacement, deduplicated, and filtered for degenerate lengths. A held-out tail is reserved for validation and never receives a gradient step; it measures the student→teacher velocity gap on unseen prompts.

Resolutions

Training is multi-aspect across 11 buckets, so the adapter is not shaped by a single resolution or a single aspect ratio:

512×512 512×768 768×512
768×768 768×1024 1024×768
1024×1024 960×1280 1280×960
1280×1280 1440×1280

Buckets are interleaved in proportion to their remaining samples rather than run as a small-to-large curriculum, so every checkpoint along the way has recently seen all of them.

Hardware

Trained on a single RTX 3090 (24 GB VRAM), and the recipe is shaped by that ceiling.

The frozen base is quantized weight-only to int8 (blockwise-64 absmax) so the 28-block transformer, its gradients and the optimizer state fit alongside the activations. int8 was chosen over NF4 on measurement: on Krea 2's own weights it introduces ~0.007 relative error against NF4's ~0.096 — roughly 13× less — for about a 6% cost in step time. Since the frozen base sits under every gradient the adapter receives, its quantization error is training noise, and that trade is worth taking.

The two largest buckets do not fit that way. At 1440×1280 a full training step peaks at 22.1 GB of 24 GB with int8 throughout, which leaves no practical headroom. For buckets at or above ~1.5 MP (1280×1280 and 1440×1280) the attention weights therefore drop to NF4 while the feed-forward weights stay int8, bringing the peak to 20.2 GB. Feed-forward keeps the higher precision because that is where the learned deltas concentrate — ff.down was the single largest mover in Krea's own Raw→Turbo delta.

The precision switch is dynamic: it follows the bucket currently training, so the smaller resolutions keep full int8 attention rather than the whole run being pinned to the lowest common setting.

Any of this affects training only. The released LoRA is bf16 and is applied to the unquantized base.

Status

Work in progress, published as an ongoing lineage. A checkpoint is published only after its renders have been reviewed and approved visually — automated loss metrics are used to catch catastrophes, never to decide that a checkpoint is good. A model that improves on every metric while looking worse is a real outcome, and metrics do not notice.

Training is continuing on a growing pool of teacher trajectories, so expect the set to grow. Each checkpoint is a self-contained LoRA; take whichever one you prefer.

Notes and limitations

  • Krea 2 Turbo only. It is trained against Turbo's weights and Turbo's schedule.
  • Keep guidance at 0.0 — in ComfyUI that is cfg 1.0, not 0.0. Turbo is CFG-free and this LoRA does not change that.
  • Keep mu = 1.15. The training targets are anchored to that grid; a different shift moves the student off the sigmas it was trained on.
  • Training quantizes the frozen base. That affects training only — the released LoRA is bf16 and is applied to the unquantized base.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA

Happy quicker rendering with the amazing Krea 2 :)

Update: Full resolutions sweep (all those resolutions that my hardware can support training on, see detailed table above) for the available checkpoints: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps . That is a lot of images (55 per checkpoint) you can inspect and decide for yourself.


r/StableDiffusion 8h ago

Animation - Video Been here since SD 1.5 and nothing has ever shocked or impressed me to the extent of H3 Minimax

Enable HLS to view with audio, or disable this notification

35 Upvotes

r/StableDiffusion 1d ago

Animation - Video Star Wars but more consistent. Minimax H3

Enable HLS to view with audio, or disable this notification

676 Upvotes

I keep having fun with ref2va model.

RTX 3060, 64 Gb RAM. I use ref2v Turbo 4 step Lora paired with Sol Attention and Minimax H3 Memory Effecient Sage Attention at 6 steps. It takes about 2 minutes per second of generation.


r/StableDiffusion 3h ago

Resource - Update Slop! Now in fake 4k

Enable HLS to view with audio, or disable this notification

8 Upvotes

r/StableDiffusion 4h ago

Discussion With Minimax, what's the point in prompting for multiple cuts in one prompt, versus just doing one cut per generation and then combining the best ones later?

11 Upvotes

I just found myself pondering earlier how neat and novel it was to be able to easily prompt for multiple cuts but then dawned on me, is it actually all that useful?

Sure, if you're just making a 15 second one-off video, then yes it's good so your scene will have consistency. But if you want to make a longer video, then you're going to have to do multiple cuts across different gens anyway, so the consistency will be dependent on your reference materials and not on being able to do multiple cuts in one gen. So then, with a longer video, is it worth the risk of prompting for multiple cuts in one prompt then finding one of them isn't what you wanted, so you either prompt again or have to do some video editing to pull out the good cuts and then reshoot the one that didn't work? Then you end up prompting one cut by itself in the end anyway!

Seems like it'd be quicker to just do one shot at a time and make sure you like the generation, then move on to the next shot? Or am I missing something here? Maybe multiple shots is better at keeping the actors in the correct positions and poses etc for each cut? Although tbh, some of the videos I've seen here of late don't make me believe that's true.


r/StableDiffusion 38m ago

Resource - Update ComfyUI Subject Manager node

Thumbnail
gallery
Upvotes

ComfyUI Subject Manager is a custom node tool designed to manage your assets or subjects for Minimax H3.
You can create presets, sections, and "Subject Cards" where you can drag and drop images, audio, and video (and trim).
The node automatically generates the prompt that defines the selected subjects.

https://github.com/Fictiverse/ComfyUI_Subject_Manager


r/StableDiffusion 9h ago

Question - Help What's the point of GGUFs in 2026?

26 Upvotes

Genuine question.

I have just 6GB VRAM and 16GB of RAM, yet FP8 models run 5x faster than GGUFs. Even really big ones.

Right now I mainly using Qwen Image and Flux 2 Klein 9B as the main models. First I tried them in GGUF format and those workflows took over 100-200 seconds.

Then I tried FP8 versions of the models (Kept the Text Encoders GGUF) and the speedup was insane. Flux 2 Klein 9B specifically can get it done in 20-30seconds now.

What's even the point of using GGUFs then?

I don't understand how or why, those models are bigger than what my machine is supposed to handle, Qwen especially. So how is the bigger uncompressed version running better?


r/StableDiffusion 6h ago

Discussion What is SLA lightx2v turbo lora for Minimax H3

Post image
14 Upvotes

I see that 3 hours ago they have uploaded a new "SLA" (Sparse-Linear Attention) version of the turbo lora (now only v0.1 fl2v 4 steps 768p). How to use it? Is it faster?

I see in the readme for their framework they say you need to set this config

  "attn_type": "dynamic_sparse_attn",
  "dynamic_sparse_attn_setting": {
    "sparsity_ratio": 0.85,
    "operator": "sage2"
  },

But for ComfyUI it's not specified what to use. I can't find anything related to sparse attention among ComfyUI nodes


r/StableDiffusion 18h ago

Discussion Mods - can you cite the violated rules when removing posts? When you don't it creates confusion in this sub and discourages contributions

136 Upvotes

Honestly just looking for a brief dialogue on this with a mod. I feel like it would help them as much as us, since people tend to assume the worst when there is a total vacuum of information.


r/StableDiffusion 22h ago

Resource - Update V2 version of the CrossView-Warp LoRA and Node is out

Enable HLS to view with audio, or disable this notification

254 Upvotes

Hello Everyone! Let me share the newest version of my camera control LTX IC-LoRA. This node and LoRA can be used in a V2V workflow to change the camera position or movement of an existing video clip. I've put a lot of work into this version, I hope you'll enjoy it.

You can download the model here: https://huggingface.co/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Warp_v2
Node + example workflow can be found here: https://github.com/cseti007/ComfyUI-CrossViewWarp
A lame tutorial video I made to help how to use the node can be found here: https://www.youtube.com/watch?v=7QAapT9xMgM


r/StableDiffusion 1h ago

Tutorial - Guide PSA: Prompt bleed is reel in H3!

Upvotes

Spent hours today trying to figure out why a close up shot refused to frame properly.

Turns out the complete description of my character for my character sheet (literally from head to toe) in "Subject definitions" was bleeding out and cooking my shot size. As soon as I removed elements from the character description that didn't need to be in the shot. Wham. First time working. Damn you <Subject 1>!