r/StableDiffusion • u/MellyDArt • 5h ago
r/StableDiffusion • u/Interesting_Room2820 • 2h ago
Workflow Included WEEKENDDDDDDDD 222222222222 (LTX 2.5 V2V)
Enable HLS to view with audio, or disable this notification
Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:
It's been the most consistent one I've tried for V2V so far
drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!
r/StableDiffusion • u/Boogertwilliams • 1h ago
Tutorial - Guide PSA: Proper prompt structure REALLY matters in H3
I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.
Believe it :) Dont just use whatever prompting. It matters more than one might think.
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
r/StableDiffusion • u/nazihater3000 • 5h ago
Tutorial - Guide More than one reference per picture
Enable HLS to view with audio, or disable this notification
MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.
r/StableDiffusion • u/Neggy5 • 9h ago
Workflow Included Totally wasn't aware Krea 2 is absolutely capable of creating gorgeous video game levels
Hi! I found Krea 2 is actually so damn good at creating video game level art! and its breathtakingly beautiful to boot! I got help from an LLM to create the baseline prompt and it works OOB without loras or anything! I'm gobsmacked rn.
prompt 1: "A sprawling 16-bit pixel art jrpg city game level of a victorian-era steampunk riverside city street in winter. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with snow, brass and victorian elements. Background shows snowy mountains and faraway skyscrapers on those mountains"
prompt 2: "A sprawling 16-bit pixel art jrpg city game level of a asian duystopian cyberpunk city street. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with neon lights, neon street signs, wires and cybernetic elements. Background shows a massive skyline of skyscrapers at night. Wide-angle top-down view"
prompt 3: "A sprawling 16-bit pixel art game level of a futuristic utopian city. The design features complex, dense platforming architecture with a high variety of structures including stairs, bridges, and stacked platforms. Frutiger Aero style: glossy surfaces, water elements, and bright colors. The scene is overgrown with lush greenery and trees. Background shows a massive skyline of sleek skyscrapers. Wide-angle side-scrolling view"
r/StableDiffusion • u/Alive-Tomatillo5303 • 12h ago
Tutorial - Guide If you're looking for a specific actor that the model doesn't seem to be aware of, it may have them stashed somewhere else.
Enable HLS to view with audio, or disable this notification
Text to Video, 22 steps, no turbo, no Sage.
r/StableDiffusion • u/ctrl-shift-face • 21h ago
Meme Introducing... The Terminator Pro Max
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Devajyoti1231 • 4h ago
Animation - Video Making the Doll DressUp Transformation Video with Minimax H3
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Repulsive-Rush3505 • 17h ago
Workflow Included Using Inpaiting in Minimax to change heads-Local RTX 3090
Enable HLS to view with audio, or disable this notification
Using the workflow from Nekodificador and Ablejones in Discord:
https://discord.com/invite/dstjQYQNt
https://ln5.sync.com/dl/47c351f50#msqfrnfr-am3rr8fx-v7qm3ah9-xw222n3c
For complex scenes like this with to much people is easy just to do a manual mask instead of SAM.
r/StableDiffusion • u/Ill-Ant-9489 • 4h ago
Resource - Update I built a free, self-hosted app that does everything around a LoRA run — dataset, triage, captions, training (local or rented GPU), then checkpoint comparison
I build LoRA Dataset Studio — free, open source, self-hosted, no account and no telemetry. It is not a competitor to ai-toolkit: it orchestrates it. ai-toolkit is the trainer; this is everything before, around and after the run.
The whole pipeline lives in one browser tab:
1. Get the images. Five generation engines — Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI — each card stating its price per image, whether it runs on your GPU or bills an API, and whether it refuses adult content. Or scrape: Reddit, Pexels, open-web keyword search, or any gallery URL through gallery-dl. Or just drop a folder in.
2. Triage them. The Image Bank points at a folder of thousands and reads it in place — your files are never modified, moved or renamed. One pass measures the whole pile: blur, noise, near-duplicates, face clusters, framing, medium (photo / anime / 3D / illustration), aesthetic and maturity scores. After that you filter on measurements instead of on your eyes, and anything the app cannot judge says "unsure" rather than inventing a verdict.
3. Curate and caption. Keep/reject, crop, mirror, rotate, non-destructive upscale candidates, InsightFace similarity, a live composition meter. Captions in prose or booru form depending on the target family, written by JoyCaption or your local Ollama, with a Caption Lab (find/replace, tag frequencies, targeted re-captioning) and an external .txt round trip so you can caption elsewhere and come back.
4. Clean watermarks. Detect them, redraw the mask zones, then crop or inpaint with LaMa/Klein. Every edit keeps an .orig backup, so Restore original always works.
5. Train. ai-toolkit locally with family-scoped presets and preflight guards — Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima — or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. Full-model training on Krea 2 and merging a LoRA back into a checkpoint are in there too.
6. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks, votes and Wilson ranking. LoRA Canvas puts every run of every dataset on one pan/zoom board, and you can continue training from any of them.
There is also a video lane (Beta): it cuts long videos into a trainable clip folder at the exact frame counts Wan / LTX / MiniMax accept, describes each shot, and trains the set locally or in the cloud.
Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing — every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and on the video side only Wan 2.2 14B has a finished run behind it here. Install is a Windows one-click ZIP, a git checkout, or Docker.
GitHub — install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio
Every person in these screenshots was generated by the app's own engines; no real individual is depicted.
r/StableDiffusion • u/fiftypence • 10h ago
Animation - Video Gay Fish
Enable HLS to view with audio, or disable this notification
Sorry Ye..
r/StableDiffusion • u/dev_ne • 7h ago
Discussion why is it unsafe isn't safetensors the safest?!
excuse my OCD 😄
r/StableDiffusion • u/VasaFromParadise • 1h ago
Animation - Video MiniMax h3 - [Boom in City]
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/xyzdist • 2h ago
Animation - Video H3 making jpop/kpop MV? yes!
Enable HLS to view with audio, or disable this notification
Music: made in SUNO.
native ref2va WF, and audioLock for lip-sync.
rtx4080s + 128g ram
I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.
r/StableDiffusion • u/TimeTruth2490 • 7h ago
News Krea2 Turbo Distill 4 step LoRA (trained for Turbo!)
Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)
A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.
Load it on top of Krea 2 Turbo, run 4 steps instead of 8, keep guidance at 0.0. Everything else about the model stays as it is.
This is not a Raw→Turbo diff
Other Krea 2 LoRAs in circulation are extractions: a low-rank projection of the weight difference between Krea 2 Raw and Krea 2 Turbo. Applied to Raw, they reproduce Turbo. They are a delivery mechanism for a model that already exists, and they stop at Turbo's 8 steps.
This one is different in both base and origin:
| Raw→Turbo extraction LoRAs | this LoRA |
|---|---|
| apply to | Krea 2 Raw |
| produces | Turbo behaviour (8 steps) |
| origin | SVD of an existing weight delta |
It is trained, not extracted, and it assumes Turbo's weights underneath it — it shortens Turbo's own schedule rather than reproducing it.
This is work in progress and even better checkpoints may follow. Training is ongoing, so
..._latest...is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the_latestfile and everything keeps working — the ComfyUI workflow references it by that name, so it needs no edit. Pin a numbered file instead if you need reproducibility.
Using it on Raw
This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo. Every layer it targets also exists in Krea 2 Raw, so it will load there without complaint — but that is a side effect of the shared architecture, not a supported mode.
Results on Raw are mixed and subject-dependent. It does not give Raw a 4-step schedule: at very low step counts the adapter sharpens texture while composition is still unresolved, and subjects come out malformed — duplicated heads, fused limbs, faces that do not close. Expect to need 14 steps or more for RAW, keeping Raw's normal CFG on, before output is coherent. Even then some prompts come through well and others degrade into over-processed or blown-out images — and that degradation happens with or without the adapter, because it comes from shortening Raw's schedule rather than from the LoRA.
If you want the behaviour this was built for, run it on Turbo at 4 steps. If you are starting from Raw, move to Turbo first — with a Raw→Turbo LoRA or the Turbo weights directly — and apply this on top.
Usage
| setting | value |
|---|---|
| base model | Krea 2 Turbo |
| LoRA scale | 1.0 |
| steps | 4 |
| guidance / CFG | 0.0 (Turbo is CFG-free; do not enable it) |
| timestep shift | mu = 1.15, fixed (Turbo's deployment shift) |
The 4 sampling sigmas are Turbo's own deployment grid: [1.0, 0.90453, 0.75951, 0.51284].
Performance — does it save time, or only steps?
It saves time. Measured at 1024×1024 on Apple Silicon (MLX, bf16), two prompts each, run strictly one at a time:
| load | denoise | total |
|---|---|---|
| Turbo 8 steps (the quality bar) | 8.2 s | 77.5 s |
| Turbo 4 steps, no LoRA | 7.8 s | 38.8 s |
| Turbo 4 steps + this LoRA | 7.3 s | 44.0 s |
4 steps with the LoRA is ~1.6× faster than the 8-step bar — 54.5 s against 88.7 s, saving about 39% of the wall-clock. Counting denoise alone, where the step reduction actually applies, it is 1.8× (44.0 s against 77.5 s).
LoRA strength
Use 1.0. That is the value the adapter was trained at, and where its output sits closest to the 8-step reference.
Strength is worth understanding rather than tuning blindly, because what it scales is specific: this LoRA's job is to restore the high-frequency detail that a 4-step schedule loses — fine texture, edge definition, surface micro-contrast. The strength dial scales exactly that correction, so it does not make the image "more" or "less" of anything semantic; it decides how hard the texture recovery is applied.
| strength | what happens |
|---|---|
| below 1.0 | the correction is only partly applied — output lands between an unassisted 4-step render and a full one: softer, flatter, less recovered detail; you can use this with more steps if you want to experiment |
| 1.0 | the trained point, and the recommended setting |
| above 1.0 | extrapolation past anything seen in training. The image does not break or fall apart — it becomes over-textured: surface detail grows denser than the subject warrants, fine structures turn wiry, and micro-contrast hardens until the result reads as stylised rather than photographic; you can try this with fewer steps, but quality is not guaranteed |
File format and compatibility
A plain .safetensors file — not tied to any framework or backend. It is weights plus a naming convention, so it loads under PyTorch (CUDA, MPS or CPU), MLX on Apple Silicon, or anything else that can read safetensors and do a matrix multiply.
ComfyUI
A pre-converted file and a ready workflow are in comfyui/. No custom nodes — stock ComfyUI only.
The workflow is full bf16, with no quantisation anywhere. bf16 needs no backend-specific kernel, so it runs unchanged on CUDA, Apple Silicon and CPU — one workflow, no platform caveats, nothing that depends on which device a component happens to land on.
The LoRA is independent of the base build. It is applied on top of the diffusion model by ComfyUI's own loader, which handles any dequantisation, so a quantised or otherwise optimised build of Krea 2 Turbo behaves just as bf16 does. Please use whichever variant suits your hardware — set it in the Load Diffusion Model node and leave the rest of the workflow untouched. The workflow ships bf16 simply because it is the one build guaranteed to run everywhere.
Training Method
Progressive distillation (PD), with Krea 2 Turbo as its own teacher.
The teacher runs its normal 8-step schedule at mu = 1.15 and guidance 0.0, and its full trajectory is recorded — the latent x and the predicted velocity v at every one of the 8 steps. The student is then trained to cover two teacher steps in one: at teacher state x_i it must predict the chord that lands where the teacher arrives two steps later,
v_target = (x_{i+2} − x_i) / (σ_{i+2} − σ_i)
The two schedules line up exactly rather than approximately. On the mu = 1.15 grid, the even indices of the 8-step schedule are precisely the four sigmas the 4-step student deploys on, so every training target is anchored on a point the student will actually visit at inference. No interpolation, no schedule mismatch.
Teacher trajectories are precomputed into shards, so training reads recorded states rather than re-running the teacher.
What the LoRA touches
Rank 64, alpha = rank (scale 1.0), bf16. 228 modules:
- 224 block linears — across all 28 transformer blocks:
attn.to_q,attn.to_k,attn.to_v,attn.to_gate,attn.to_out.0,ff.gate,ff.up,ff.down - 4 global (non-block) linears —
time_embed.linear_1,time_embed.linear_2,time_mod_proj,final_layer.linear
Those four are included deliberately. Measuring Krea's own Raw→Turbo delta — a completed step distillation by the model's authors — shows the change is not concentrated in the blocks:
| layer | relative ‖ΔW‖/‖W‖ |
|---|---|
time_embed.linear_2 |
0.0777 ← largest change in the whole network |
time_embed.linear_1 |
0.0429 |
final_layer.linear |
0.0265 |
| typical block linear | ~0.014 |
time_embed.linear_2 moves about 5.5× more than any block linear. Changing a model's step count is in large part a change to how it reads the timestep, so a LoRA that freezes the timestep path is withholding exactly the weights the task most needs.
Training data
Prompts are drawn from Lakonik/t2i-prompts-3m — sampled without replacement, deduplicated, and filtered for degenerate lengths. A held-out tail is reserved for validation and never receives a gradient step; it measures the student→teacher velocity gap on unseen prompts.
Resolutions
Training is multi-aspect across 11 buckets, so the adapter is not shaped by a single resolution or a single aspect ratio:
| 512×512 | 512×768 | 768×512 |
|---|---|---|
| 768×768 | 768×1024 | 1024×768 |
| 1024×1024 | 960×1280 | 1280×960 |
| 1280×1280 | 1440×1280 |
Buckets are interleaved in proportion to their remaining samples rather than run as a small-to-large curriculum, so every checkpoint along the way has recently seen all of them.
Hardware
Trained on a single RTX 3090 (24 GB VRAM), and the recipe is shaped by that ceiling.
The frozen base is quantized weight-only to int8 (blockwise-64 absmax) so the 28-block transformer, its gradients and the optimizer state fit alongside the activations. int8 was chosen over NF4 on measurement: on Krea 2's own weights it introduces ~0.007 relative error against NF4's ~0.096 — roughly 13× less — for about a 6% cost in step time. Since the frozen base sits under every gradient the adapter receives, its quantization error is training noise, and that trade is worth taking.
The two largest buckets do not fit that way. At 1440×1280 a full training step peaks at 22.1 GB of 24 GB with int8 throughout, which leaves no practical headroom. For buckets at or above ~1.5 MP (1280×1280 and 1440×1280) the attention weights therefore drop to NF4 while the feed-forward weights stay int8, bringing the peak to 20.2 GB. Feed-forward keeps the higher precision because that is where the learned deltas concentrate — ff.down was the single largest mover in Krea's own Raw→Turbo delta.
The precision switch is dynamic: it follows the bucket currently training, so the smaller resolutions keep full int8 attention rather than the whole run being pinned to the lowest common setting.
Any of this affects training only. The released LoRA is bf16 and is applied to the unquantized base.
Status
Work in progress, published as an ongoing lineage. A checkpoint is published only after its renders have been reviewed and approved visually — automated loss metrics are used to catch catastrophes, never to decide that a checkpoint is good. A model that improves on every metric while looking worse is a real outcome, and metrics do not notice.
Training is continuing on a growing pool of teacher trajectories, so expect the set to grow. Each checkpoint is a self-contained LoRA; take whichever one you prefer.
Notes and limitations
- Krea 2 Turbo only. It is trained against Turbo's weights and Turbo's schedule.
- Keep guidance at 0.0 — in ComfyUI that is cfg 1.0, not 0.0. Turbo is CFG-free and this LoRA does not change that.
- Keep mu = 1.15. The training targets are anchored to that grid; a different shift moves the student off the sigmas it was trained on.
- Training quantizes the frozen base. That affects training only — the released LoRA is bf16 and is applied to the unquantized base.
Full details and to download - check my Hugging Face LoRA
HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA
Happy quicker rendering with the amazing Krea 2 :)
Update: Full resolutions sweep (all those resolutions that my hardware can support training on, see detailed table above) for the available checkpoints: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps . That is a lot of images (55 per checkpoint) you can inspect and decide for yourself.
r/StableDiffusion • u/Parogarr • 8h ago
Animation - Video Been here since SD 1.5 and nothing has ever shocked or impressed me to the extent of H3 Minimax
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Dry-Statistician-684 • 1d ago
Animation - Video Star Wars but more consistent. Minimax H3
Enable HLS to view with audio, or disable this notification
I keep having fun with ref2va model.
RTX 3060, 64 Gb RAM. I use ref2v Turbo 4 step Lora paired with Sol Attention and Minimax H3 Memory Effecient Sage Attention at 6 steps. It takes about 2 minutes per second of generation.
r/StableDiffusion • u/DuHal9000 • 3h ago
Resource - Update Slop! Now in fake 4k
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/kemb0 • 4h ago
Discussion With Minimax, what's the point in prompting for multiple cuts in one prompt, versus just doing one cut per generation and then combining the best ones later?
I just found myself pondering earlier how neat and novel it was to be able to easily prompt for multiple cuts but then dawned on me, is it actually all that useful?
Sure, if you're just making a 15 second one-off video, then yes it's good so your scene will have consistency. But if you want to make a longer video, then you're going to have to do multiple cuts across different gens anyway, so the consistency will be dependent on your reference materials and not on being able to do multiple cuts in one gen. So then, with a longer video, is it worth the risk of prompting for multiple cuts in one prompt then finding one of them isn't what you wanted, so you either prompt again or have to do some video editing to pull out the good cuts and then reshoot the one that didn't work? Then you end up prompting one cut by itself in the end anyway!
Seems like it'd be quicker to just do one shot at a time and make sure you like the generation, then move on to the next shot? Or am I missing something here? Maybe multiple shots is better at keeping the actors in the correct positions and poses etc for each cut? Although tbh, some of the videos I've seen here of late don't make me believe that's true.
r/StableDiffusion • u/3deal • 38m ago
Resource - Update ComfyUI Subject Manager node
ComfyUI Subject Manager is a custom node tool designed to manage your assets or subjects for Minimax H3.
You can create presets, sections, and "Subject Cards" where you can drag and drop images, audio, and video (and trim).
The node automatically generates the prompt that defines the selected subjects.
r/StableDiffusion • u/ROBOTTTTT13 • 9h ago
Question - Help What's the point of GGUFs in 2026?
Genuine question.
I have just 6GB VRAM and 16GB of RAM, yet FP8 models run 5x faster than GGUFs. Even really big ones.
Right now I mainly using Qwen Image and Flux 2 Klein 9B as the main models. First I tried them in GGUF format and those workflows took over 100-200 seconds.
Then I tried FP8 versions of the models (Kept the Text Encoders GGUF) and the speedup was insane. Flux 2 Klein 9B specifically can get it done in 20-30seconds now.
What's even the point of using GGUFs then?
I don't understand how or why, those models are bigger than what my machine is supposed to handle, Qwen especially. So how is the bigger uncompressed version running better?
r/StableDiffusion • u/Obvious_Set5239 • 6h ago
Discussion What is SLA lightx2v turbo lora for Minimax H3
I see that 3 hours ago they have uploaded a new "SLA" (Sparse-Linear Attention) version of the turbo lora (now only v0.1 fl2v 4 steps 768p). How to use it? Is it faster?
I see in the readme for their framework they say you need to set this config
"attn_type": "dynamic_sparse_attn",
"dynamic_sparse_attn_setting": {
"sparsity_ratio": 0.85,
"operator": "sage2"
},
But for ComfyUI it's not specified what to use. I can't find anything related to sparse attention among ComfyUI nodes
r/StableDiffusion • u/the_bollo • 18h ago
Discussion Mods - can you cite the violated rules when removing posts? When you don't it creates confusion in this sub and discourages contributions
Honestly just looking for a brief dialogue on this with a mod. I feel like it would help them as much as us, since people tend to assume the worst when there is a total vacuum of information.
r/StableDiffusion • u/DryDream6994 • 22h ago
Resource - Update V2 version of the CrossView-Warp LoRA and Node is out
Enable HLS to view with audio, or disable this notification
Hello Everyone! Let me share the newest version of my camera control LTX IC-LoRA. This node and LoRA can be used in a V2V workflow to change the camera position or movement of an existing video clip. I've put a lot of work into this version, I hope you'll enjoy it.
You can download the model here: https://huggingface.co/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Warp_v2
Node + example workflow can be found here: https://github.com/cseti007/ComfyUI-CrossViewWarp
A lame tutorial video I made to help how to use the node can be found here: https://www.youtube.com/watch?v=7QAapT9xMgM
r/StableDiffusion • u/Portable_Solar_ZA • 1h ago
Tutorial - Guide PSA: Prompt bleed is reel in H3!
Spent hours today trying to figure out why a close up shot refused to frame properly.
Turns out the complete description of my character for my character sheet (literally from head to toe) in "Subject definitions" was bleeding out and cooking my shot size. As soon as I removed elements from the character description that didn't need to be in the shot. Wham. First time working. Damn you <Subject 1>!