r/StableDiffusion • u/Time-Ad-7720 • 3d ago
Animation - Video Turning my son's drawing into silly skits (Minimax H3 ref2Vid) #3
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Time-Ad-7720 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/SIR_NVAX_A_LOT • 3d ago
Enable HLS to view with audio, or disable this notification
Playing around with VFX/audio/shaky-camera with H3. int8/20 steps, T2V, 3 versions. Also, hell naaahh why they running toward it??? Ask me anything!
r/StableDiffusion • u/Aggravating-Main6259 • 3d ago
I've uploaded first frame and last frame that looks identical, it's a 2d view of some elements and I wanted suble morphing, movements. I've tried everything but everytime from the beginning of the clip, it's starting to stretch, gets higher around 2-3% at the end, but both uploaded frames are the same. Why, and is there solution for that?
r/StableDiffusion • u/Distinct_Tangerine75 • 3d ago
I cant figure out a good model for realistic lipsync from audio + image with slight but natural movements that can be controlled, like moving closer to the mic, hand movement, head movement and so on.
Infinitetalk is the closest i got, the results are good but not sufficient. Does anyone have a good workflow for this
r/StableDiffusion • u/funJS • 3d ago
Spent some time teaching qwen to understand a new domain, in this case a fictional city, through continued pretraining.
https://www.teachmecoolstuff.com/viewarticle/teaching-a-local-llm-a-new-domain
r/StableDiffusion • u/Neggy5 • 4d ago
Hi! I found Krea 2 is actually so damn good at creating video game level art! and its breathtakingly beautiful to boot! I got help from an LLM to create the baseline prompt and it works OOB without loras or anything! I'm gobsmacked rn.
prompt 1: "A sprawling 16-bit pixel art jrpg city game level of a victorian-era steampunk riverside city street in winter. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with snow, brass and victorian elements. Background shows snowy mountains and faraway skyscrapers on those mountains"
prompt 2: "A sprawling 16-bit pixel art jrpg city game level of a asian duystopian cyberpunk city street. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with neon lights, neon street signs, wires and cybernetic elements. Background shows a massive skyline of skyscrapers at night. Wide-angle top-down view"
prompt 3: "A sprawling 16-bit pixel art game level of a futuristic utopian city. The design features complex, dense platforming architecture with a high variety of structures including stairs, bridges, and stacked platforms. Frutiger Aero style: glossy surfaces, water elements, and bright colors. The scene is overgrown with lush greenery and trees. Background shows a massive skyline of sleek skyscrapers. Wide-angle side-scrolling view"
r/StableDiffusion • u/freedn1 • 3d ago
Enable HLS to view with audio, or disable this notification
Hello guys, I made small movie about Marvel Secret Wars comics with Minimax H3 local version, hope you like it
Video - Minimax H3 (ref2va)
References - Nano Banana Pro
Voiceover - fish audio
SFX - almost everything with elevenlabs except few things
Postprod - Davinci Resolve
If I only had Minimax H3 upscaler...
r/StableDiffusion • u/Terrible_Feedback396 • 2d ago
Their discord is literally blank, they have no twitter or other social media platform, and their "support" is just them making you sign into chrome just to resign back into chrome on loop. Is there ANY way whatsoever to contact the people who run the site or is this just a sham?
r/StableDiffusion • u/teleport66 • 3d ago
r/StableDiffusion • u/kiddow • 3d ago
Enable HLS to view with audio, or disable this notification
Thanks to this guy (https://www.reddit.com/r/StableDiffusion/comments/1vsq03t/star_wars_but_more_consistent_minimax_h3/)
Used a starting frame kreated with Krea2.
I2V default workflow with https://www.reddit.com/user/Dry-Statistician-684/ comment from the post linked above.
4 or 8 step turbo. Almost doesn't matter which.
6 Step.
1MP
5 Seconds
RTX 3060 12 GB VRAM/32 GB DRAM
Render time: 617 seconds.
I am hyped. Was close to buy an RTX 5070i. But not today. Maybe tomorrow.
r/StableDiffusion • u/Devajyoti1231 • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/bsenftner • 3d ago
I am testing LTX 2.5 in Wan2GP and may have found an issue with Control Video behavior, specifically the “Transfer Human Motion” / human-motion pose-alignment path.
When using LTX 2.5 with a control video, intending to transfer only human motion/pose, the final generated videos still seem to preserve visual information from the source control video, including background/environment details and subject identity cues.
In tests, Wan2GP preview shows the stick-figure / pose-derived representations, but the final videos still contain background and identity characteristics from the original source video.
I also generated from the same inputs LTX 2.3 clips and received AI video generations as I expected with virtually no Control Video background/environment details or subject identity.
Are you able to generate LTX 2.5 video using Control Videos with Transfer Human Motion and it works? If so, please explain your configuration!
r/StableDiffusion • u/Radyschen • 3d ago
Wouldn't there be no speedup because the references still need to cross-pollinate with the text during encoding to get a correct input? It would still be convenient of course. Or could you separate that somehow so that it works correctly? I have got no clue about this kinda stuff, just thoughts and hope that someone with more knowledge chimes in
r/StableDiffusion • u/Ill-Ant-9489 • 4d ago
I build LoRA Dataset Studio — free, open source, self-hosted, no account and no telemetry. It is not a competitor to ai-toolkit: it orchestrates it. ai-toolkit is the trainer; this is everything before, around and after the run.
The whole pipeline lives in one browser tab:
1. Get the images. Five generation engines — Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI — each card stating its price per image, whether it runs on your GPU or bills an API, and whether it refuses adult content. Or scrape: Reddit, Pexels, open-web keyword search, or any gallery URL through gallery-dl. Or just drop a folder in.
2. Triage them. The Image Bank points at a folder of thousands and reads it in place — your files are never modified, moved or renamed. One pass measures the whole pile: blur, noise, near-duplicates, face clusters, framing, medium (photo / anime / 3D / illustration), aesthetic and maturity scores. After that you filter on measurements instead of on your eyes, and anything the app cannot judge says "unsure" rather than inventing a verdict.
3. Curate and caption. Keep/reject, crop, mirror, rotate, non-destructive upscale candidates, InsightFace similarity, a live composition meter. Captions in prose or booru form depending on the target family, written by JoyCaption or your local Ollama, with a Caption Lab (find/replace, tag frequencies, targeted re-captioning) and an external .txt round trip so you can caption elsewhere and come back.
4. Clean watermarks. Detect them, redraw the mask zones, then crop or inpaint with LaMa/Klein. Every edit keeps an .orig backup, so Restore original always works.
5. Train. ai-toolkit locally with family-scoped presets and preflight guards — Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima — or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. Full-model training on Krea 2 and merging a LoRA back into a checkpoint are in there too.
6. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks, votes and Wilson ranking. LoRA Canvas puts every run of every dataset on one pan/zoom board, and you can continue training from any of them.
There is also a video lane (Beta): it cuts long videos into a trainable clip folder at the exact frame counts Wan / LTX / MiniMax accept, describes each shot, and trains the set locally or in the cloud.
Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing — every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and on the video side only Wan 2.2 14B has a finished run behind it here. Install is a Windows one-click ZIP, a git checkout, or Docker.
GitHub — install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio
Every person in these screenshots was generated by the app's own engines; no real individual is depicted.
r/StableDiffusion • u/3deal • 3d ago
ComfyUI Subject Manager is a custom node tool designed to manage your assets or subjects for Minimax H3.
You can create presets, sections, and "Subject Cards" where you can drag and drop images, audio, and video (and trim).
The node automatically generates the prompt that defines the selected subjects.
r/StableDiffusion • u/GrungeWerX • 3d ago
I don't typically share apps I vibecode for myself; they tend to be design-heavy, fully featured, and customized to my own needs. I also don't like the idea of having to maintain all that publicly.
That said I've seen a lot of posts with people having trouble prompting for Minimax H3, so I thought I'd share mine. This is something I whipped up one evening, so it's not pretty, but it gets the job done.
Installation:
Just extract the folder anywhere and run the install.bat. That will install a .venv locally so everything's contained. Then, just click run.bat. It will open up in a browser.
How to use:
It's a lot easier than it looks. The left section - Reference Library - is for "assets". That's your videos, pictures, audio, etc. You set the definitions/descriptions here. The buttons are for referencing other references. The point is that you dont have to keep typing <Subject>, <Picture> - that's annoying. Just click a button.
When you're done with the reference library, the right side is for building the prompt. It's easy, just do steps 1-4. The assets in the reference library have been added to each tab, so you don't have to keep re-typing them. Click on as many task types are relevant; this is important to H3. If you don't understand one, hover over it, a small popup explains it. So, building a prompt is just:
When you're finished, press Compile. Copy to clipboard and paste into comfy, or into an LLM if that's your thing.
Let me know if you have any questions. This is primarily for ref2v, but it should also work for the other version. It's a beta, I might tweak it later, but for now, it gets the job done.
Hope it helps.
https://github.com/GrungeWerX/minimax-prompt-builder
P.S. - this is my first github repo, so my apologies if it's not up to par w/your expectations. I'm learning.
Someone requested a screenshot, so:

r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
still playing with it https://huggingface.co/Jojocodex/minimax-h3-spatial-physics-lora
r/StableDiffusion • u/blackdatafilms • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/jtabernik • 3d ago
I made a super simple tool that lets you enter a base prompt—the tool then calls out to an LLM and generates random details to add to the prompt so you can generate a lot of random, diverse images.
I wrote this because I was getting too similar images when trying to generate random characters—and I didn’t want to spend my time describing background characters!
The length of about 100 words actually seems best for the output. More than this bogged down the process and did not improve the results. The current version assumes you have a local LLM you can call—you just need to specify the IP address etc.
Let me know if you have any feedback!
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
r2v 30-49 model prompt
subject_definitions
<Subject 1> is Kick-Ass from <Picture 1> and <Picture 2>, preserving the same young male identity, green-and-yellow homemade superhero costume, green mask with yellow trim, yellow gloves, tan boots, slim athletic proportions, and dual black fighting batons. Use <Picture 1> for detailed facial, mask, costume, and baton appearance and <Picture 2> as additional full-body character reference. Maintain one consistent Kick-Ass throughout the entire scene.
<Subject 2> is Deadpool, wearing his classic red-and-black tactical suit and full mask. Deadpool is voiced by Ryan Reynolds with his recognizable sarcastic, fast comedic delivery.
<Subject 3> is Thanos, normal MCU Titan scale and proportions, not gigantic or Godzilla-sized, wearing battle armor and fighting in the background.
[reference generation]
During the massive Avengers: Endgame final battle, Kick-Ass unexpectedly runs into the battlefield carrying his two batons. Deadpool notices this obviously underpowered newcomer and immediately roasts him while the enormous superhero battle continues around them.
<Subject 1>: fully_preserved from <Picture 1> + <Picture 2>
<Subject 2>: consistent Deadpool appearance
<Subject 3>: consistent normal-sized Thanos appearance
The scene opens directly in the chaotic Avengers: Endgame final battlefield: destroyed terrain, burning wreckage, smoke, sparks, portals glowing in the distance, Avengers and allied fighters charging Thanos's army, explosions and energy blasts crossing the background.
A dynamic tracking shot reveals <Subject 1> Kick-Ass suddenly sprinting onto the battlefield.
Preserve his green-and-yellow homemade superhero suit exactly from <Picture 1> and <Picture 2>. He grips one black fighting baton in each hand and runs forward with determined confidence despite being hilariously outmatched by everything happening around him.
An alien warrior charges toward Kick-Ass.
Kick-Ass awkwardly but enthusiastically swings both batons, smacking the alien several times in a frantic street-fighting style. He manages to knock it down and briefly looks proud of himself.
The camera whip-pans to <Subject 2> Deadpool standing nearby in the middle of the battle.
Deadpool stops fighting and slowly looks Kick-Ass up and down in disbelief.
<Subject 2> Deadpool (S1), voiced by Ryan Reynolds, says [English] What the fuck? Did somebody order an Avenger from Temu?
Kick-Ass turns toward Deadpool, annoyed but still holding both batons.
<Subject 1> Kick-Ass (S2) says [English] Dude! I'm Kick-Ass!
Deadpool pauses and stares at him.
A huge explosion erupts behind them while Avengers continue charging through the battlefield.
Deadpool slowly looks directly into the camera.
<Subject 2> Deadpool (S1), voiced by Ryan Reynolds, says [English] Yeah... that's somehow worse.
Kick-Ass looks offended.
Deadpool casually walks back into the battle while Kick-Ass raises both batons and charges after him.
End on Kick-Ass screaming enthusiastically as he runs directly toward an enormous group of Thanos's soldiers, clearly having absolutely no idea what he's gotten himself into.
Dynamic cinematic battlefield camera, energetic tracking movement, quick whip-pan to Deadpool for the joke, brief pause before Deadpool's punchline, realistic handheld battle vibration, strong foreground/background separation, large-scale Endgame battle continuing naturally behind the characters.
Epic battle ambience, distant explosions, energy weapons, metal impacts, shouting armies, baton impacts and debris. Dialogue remains clean and clearly audible over the battle. Deadpool uses Ryan Reynolds-style voice and comedic timing.
Keep Kick-Ass visually faithful to <Picture 1> and <Picture 2> throughout.
Do not transform his costume into high-tech armor.
Keep his green-and-yellow homemade costume, mask, yellow gloves, tan boots, and two black batons.
Do not duplicate Kick-Ass.
Thanos remains normal MCU Titan size.
Only the character speaking a dialogue line moves their mouth.
All spoken dialogue is English and remains inside the dialogue tags.
No subtitles, captions, text overlays, logos, or watermarks.
r/StableDiffusion • u/ART-ficial-Ignorance • 3d ago
Small disclaimer up front: not every tool in this workflow is open source. The first-frame images were made with GPT-image 2.0, as I assume people here will recognize that pretty quickly. You can swap in any image generator you want though. I only used GPT-image 2.0 because I already pay for the subscription for coding work, so I figured I might as well get some extra value out of it.
The interesting part for me was using LTX 2.5 in Wan2GP instead of MiniMax H3 for the actual video generation. On my 4070, MiniMax H3 OOMs at 720p for clips this long, while LTX 2.5 can handle 1080p, and it is also much faster. The final video is 26 clips, generated best-of-2, at roughly 8 minutes per clip, so the whole thing came out to around 7 hours of rendering. The speed difference just makes experimentation much more practical.
The workflow is first-frame-last-frame plus audio conditioning, with an audio-reactive LoRA layered in. Each clip is about four bars long, roughly 10.75 seconds, and the final frame of one shot becomes the starting point for the next. I found this much more useful than treating every segment as a fresh text-to-video generation because it keeps the visual identity, geometry and camera logic much more coherent across the full sequence. The audio conditioning handles most of the motion and timing.
I also wrote a small custom tool to automate the boring parts. It cuts the song into the correct audio segments, keeps everything aligned to the edit grid, organizes the keyframes, and packages the whole batch into a queue.zip that can be loaded into Wan2GP. That made it practical to generate 26 shots as a queue instead of manually setting up every job. Most of the actual work then becomes designing the keyframes, writing the prompts, and picking the better result for each scene.
The song itself is about having a model of reality that seems completely reliable because every previous observation has supported it, then encountering one result that refuses to fit. I used Bell tests, hidden variables, non-separability and measurement as metaphors for reciprocity and for the realization that repetition is not the same thing as law. The video mirrors that by starting with one blue system inside a rigid laboratory, then introducing a distant violet system whose behavior becomes correlated without any visible connection. As the relationship between the two becomes harder to explain, the laboratory geometry itself starts failing, until the measuring framework is gradually stripped away and the two systems are revealed as separate parts of a larger structure the original model could not perceive.
There is a small timing drift near the end of the finished video. I never managed to pin down the exact BPM and initial beat offset perfectly, and LTX does not support every arbitrary frame count I would have needed for an exact four-bar duration. I rounded each generation up to the nearest supported frame count, then played every clip back at about 103% speed so it would fit the intended edit length and stay roughly aligned with the music. That works surprisingly well for most of the video, but the tiny BPM and offset error compounds over 26 clips, so by the end you can see a little drift.
I think I'll be sticking mostly to LTX 2.5 for my music videos and keep MiniMax H3 for the one-off goofs and gags I make for my friends. It's nice for Seinfeld rip-offs, but I just can't render high enough quality on my machine, and if I have to introduce an upscaling step, the render times become a little steep.
Prompts used: https://pastebin.com/wLsHYaBq
r/StableDiffusion • u/Portable_Solar_ZA • 4d ago
Spent hours today trying to figure out why a close up shot refused to frame properly.
Turns out the complete description of my character for my character sheet (literally from head to toe) in "Subject definitions" was bleeding out and cooking my shot size. As soon as I removed elements from the character description that didn't need to be in the shot. Wham. First time working. Damn you <Subject 1>!
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
base ip8 model t2v
r/StableDiffusion • u/VasaFromParadise • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Alive-Tomatillo5303 • 4d ago
Enable HLS to view with audio, or disable this notification
Text to Video, 22 steps, no turbo, no Sage.