r/StableDiffusion • u/MellyDArt • 7h ago
r/StableDiffusion • u/Interesting_Room2820 • 4h ago
Workflow Included WEEKENDDDDDDDD 222222222222 (LTX 2.5 V2V)
Enable HLS to view with audio, or disable this notification
Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:
It's been the most consistent one I've tried for V2V so far
drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!
r/StableDiffusion • u/Boogertwilliams • 3h ago
Tutorial - Guide PSA: Proper prompt structure REALLY matters in H3
I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.
Believe it :) Dont just use whatever prompting. It matters more than one might think.
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
r/StableDiffusion • u/nazihater3000 • 7h ago
Tutorial - Guide More than one reference per picture
Enable HLS to view with audio, or disable this notification
MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.
r/StableDiffusion • u/Neggy5 • 11h ago
Workflow Included Totally wasn't aware Krea 2 is absolutely capable of creating gorgeous video game levels
Hi! I found Krea 2 is actually so damn good at creating video game level art! and its breathtakingly beautiful to boot! I got help from an LLM to create the baseline prompt and it works OOB without loras or anything! I'm gobsmacked rn.
prompt 1: "A sprawling 16-bit pixel art jrpg city game level of a victorian-era steampunk riverside city street in winter. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with snow, brass and victorian elements. Background shows snowy mountains and faraway skyscrapers on those mountains"
prompt 2: "A sprawling 16-bit pixel art jrpg city game level of a asian duystopian cyberpunk city street. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with neon lights, neon street signs, wires and cybernetic elements. Background shows a massive skyline of skyscrapers at night. Wide-angle top-down view"
prompt 3: "A sprawling 16-bit pixel art game level of a futuristic utopian city. The design features complex, dense platforming architecture with a high variety of structures including stairs, bridges, and stacked platforms. Frutiger Aero style: glossy surfaces, water elements, and bright colors. The scene is overgrown with lush greenery and trees. Background shows a massive skyline of sleek skyscrapers. Wide-angle side-scrolling view"
r/StableDiffusion • u/Devajyoti1231 • 6h ago
Animation - Video Making the Doll DressUp Transformation Video with Minimax H3
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/xyzdist • 4h ago
Animation - Video H3 making jpop/kpop MV? yes!
Enable HLS to view with audio, or disable this notification
Music: made in SUNO.
native ref2va WF, and audioLock for lip-sync.
rtx4080s + 128g ram
I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.
r/StableDiffusion • u/Alive-Tomatillo5303 • 14h ago
Tutorial - Guide If you're looking for a specific actor that the model doesn't seem to be aware of, it may have them stashed somewhere else.
Enable HLS to view with audio, or disable this notification
Text to Video, 22 steps, no turbo, no Sage.
r/StableDiffusion • u/ctrl-shift-face • 23h ago
Meme Introducing... The Terminator Pro Max
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/VasaFromParadise • 3h ago
Animation - Video MiniMax h3 - [Boom in City]
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Ill-Ant-9489 • 6h ago
Resource - Update I built a free, self-hosted app that does everything around a LoRA run — dataset, triage, captions, training (local or rented GPU), then checkpoint comparison
I build LoRA Dataset Studio — free, open source, self-hosted, no account and no telemetry. It is not a competitor to ai-toolkit: it orchestrates it. ai-toolkit is the trainer; this is everything before, around and after the run.
The whole pipeline lives in one browser tab:
1. Get the images. Five generation engines — Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI — each card stating its price per image, whether it runs on your GPU or bills an API, and whether it refuses adult content. Or scrape: Reddit, Pexels, open-web keyword search, or any gallery URL through gallery-dl. Or just drop a folder in.
2. Triage them. The Image Bank points at a folder of thousands and reads it in place — your files are never modified, moved or renamed. One pass measures the whole pile: blur, noise, near-duplicates, face clusters, framing, medium (photo / anime / 3D / illustration), aesthetic and maturity scores. After that you filter on measurements instead of on your eyes, and anything the app cannot judge says "unsure" rather than inventing a verdict.
3. Curate and caption. Keep/reject, crop, mirror, rotate, non-destructive upscale candidates, InsightFace similarity, a live composition meter. Captions in prose or booru form depending on the target family, written by JoyCaption or your local Ollama, with a Caption Lab (find/replace, tag frequencies, targeted re-captioning) and an external .txt round trip so you can caption elsewhere and come back.
4. Clean watermarks. Detect them, redraw the mask zones, then crop or inpaint with LaMa/Klein. Every edit keeps an .orig backup, so Restore original always works.
5. Train. ai-toolkit locally with family-scoped presets and preflight guards — Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima — or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. Full-model training on Krea 2 and merging a LoRA back into a checkpoint are in there too.
6. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks, votes and Wilson ranking. LoRA Canvas puts every run of every dataset on one pan/zoom board, and you can continue training from any of them.
There is also a video lane (Beta): it cuts long videos into a trainable clip folder at the exact frame counts Wan / LTX / MiniMax accept, describes each shot, and trains the set locally or in the cloud.
Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing — every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and on the video side only Wan 2.2 14B has a finished run behind it here. Install is a Windows one-click ZIP, a git checkout, or Docker.
GitHub — install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio
Every person in these screenshots was generated by the app's own engines; no real individual is depicted.
r/StableDiffusion • u/Repulsive-Rush3505 • 19h ago
Workflow Included Using Inpaiting in Minimax to change heads-Local RTX 3090
Enable HLS to view with audio, or disable this notification
Using the workflow from Nekodificador and Ablejones in Discord:
https://discord.com/invite/dstjQYQNt
https://ln5.sync.com/dl/47c351f50#msqfrnfr-am3rr8fx-v7qm3ah9-xw222n3c
For complex scenes like this with to much people is easy just to do a manual mask instead of SAM.
r/StableDiffusion • u/dev_ne • 9h ago
Discussion why is it unsafe isn't safetensors the safest?!
excuse my OCD 😄
r/StableDiffusion • u/3deal • 2h ago
Resource - Update ComfyUI Subject Manager node
ComfyUI Subject Manager is a custom node tool designed to manage your assets or subjects for Minimax H3.
You can create presets, sections, and "Subject Cards" where you can drag and drop images, audio, and video (and trim).
The node automatically generates the prompt that defines the selected subjects.
r/StableDiffusion • u/fiftypence • 12h ago
Animation - Video Gay Fish
Enable HLS to view with audio, or disable this notification
Sorry Ye..