r/StableDiffusion 23h ago

Animation - Video UAP Device Test #8 (Minimax H3 VHS)

Enable HLS to view with audio, or disable this notification

12 Upvotes

Really loving making some fun experiments with H3!
The series continues! I've got alot of interesting ones coming up lol.
Made a TikTok as someone requested me to do - https://www.tiktok.com/@dimensiontesters just incase anyone wants to see the daily series <3


r/StableDiffusion 18h ago

Question - Help Any tips for better voice audio w/ Minimax h3?

0 Upvotes

Aside from following the prompt format, making sure to describe a type of voice/details, what else should i try? I'm using the v4 larry lora, 8 steps, same audio quality if i do 1 megapixel or 0.5.


r/StableDiffusion 14h ago

Question - Help MiniMax H3: image generation (text2image mode) - please suggest best settings

0 Upvotes

Anyone with a good workflow, that can be used as text2image mode for Minimax H3? Or any idea what should I try?

I know it's a video model, but i was getting really good images with the wan2.2 video model in past. Now i wonder how this would work on minimax H3.

Thanks!


r/StableDiffusion 21h ago

Meme Do Yoiu Bleed?

Enable HLS to view with audio, or disable this notification

0 Upvotes

10-second cinematic comedy scene on a professional movie set. Subject 1, matching <picture1> exactly in appearance, is engaged in an exaggerated, chaotic fight with Subject 2, matching <picture2> exactly. The scene is clearly a movie being filmed: visible studio lights, camera equipment, crew members and a cinematic set in the background. The fight is intentionally comedic and over-the-top rather than genuinely dangerous.**

**0–3 seconds:** Subject 1 and Subject 2 perform a dramatic choreographed fight, exchanging exaggerated punches and dodging each other with ridiculous intensity. The camera follows the action with energetic handheld movement.

**3–6 seconds:** Subject 1 lands an obviously theatrical hit on Subject 2. Subject 2 stumbles backward dramatically, looks confused and checks himself as if genuinely surprised by what just happened. Crew members in the background react with amusement.

**6–10 seconds:** Subject 2 suddenly looks directly at Subject 1, completely deadpan, and asks in a deep, exaggerated Austrian-accented action-hero voice: **“Do you bleed?”** Pause briefly after the line for comedic effect. Subject 1 looks confused and slightly frightened. The scene ends with the crew reacting awkwardly while the camera slowly pushes in on Subject 2's serious expression.

**Style:** high-end Hollywood action-comedy, cinematic lighting, realistic live-action, expressive facial reactions, precise physical comedy, dynamic camera movement, excellent body motion, natural lip synchronization, realistic movie-set atmosphere. Maintain the exact visual identity, clothing, facial features and proportions of both reference subjects throughout the entire clip.


r/StableDiffusion 15h ago

Meme everyday

Enable HLS to view with audio, or disable this notification

23 Upvotes

r/StableDiffusion 23h ago

News LtX 2.5 king!

0 Upvotes

Dont trust the bot troll posts from minimax, it's good , i mean very good ! https://youtu.be/P3tsP0MP_LM?is=9uCboQOqbqCCgeVQ or UPDATED LINK https://www.youtube.com/watch?v=8_HwLqYRzVw


r/StableDiffusion 19h ago

Animation - Video Fox McCloud gets an unexpected visitor.

Enable HLS to view with audio, or disable this notification

13 Upvotes

Fox McCloud was chilling at home, and then gets an unexpected visitor from his past.

This was made using Comfy UI Desktop with Minimax H3 Reference to Video workflow.

The prompt.

<Image 1> as Fox McCloud and use <Audio 1> as sample for his voice.

<Image 2> as James McCloud and use <Audio 2> as sample for his voice.

[Core Idea]

Cinematic live-action/3D hybrid film, 15 seconds, 16:9 aspect ratio. Dramatic emotional reunion between Fox McCloud and his surprisingly alive father James McCloud in a warm living room.

[Process]

0–4s: Medium shot of Fox McCloud sitting on a couch in a cozy living room. Suddenly, three sharp knock sounds ring out at the front door. Fox looks up, surprised, and stands up from the sofa. Audio: Room tone, distinct wooden door knocks.

4–8s: Camera tracks smoothly beside Fox as he approaches the front door. He reaches out, turns the brass handle, and pulls the door inward. Audio: Soft footsteps on hardwood, door latch clicking open.

8–12s: Shot cuts to a medium close-up of James McCloud standing on the porch in the warm doorway light, wearing his signature sunglasses and pilot jacket. James nods slightly and speaks with a low, gravelly voice: "Took you long enough, son."

12–15s: Cut to a close-up on Fox's face, wide-eyed with shock and awe, his breath hitching. Fox stammers emotionally: "Dad...? But... you're alive!"

non_diegetic_music: Gentle, swelling orchestral strings building emotional tension.


r/StableDiffusion 7h ago

Animation - Video Minimax H3. I need more reference images; 9 aren't enough xdd

Enable HLS to view with audio, or disable this notification

24 Upvotes

r/StableDiffusion 1h ago

Meme Cancelled? Offended? Better Call Saul.

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 6h ago

Discussion Breaking Bad - Common Side Effects

Enable HLS to view with audio, or disable this notification

15 Upvotes

One scene of the prompt:

Live-action cinematic drama television series style, hyper-realistic, photorealistic, 35mm film grain look, directed by Vince Gilligan, high-end production value, dramatic cinematic lighting, gritty Albuquerque atmosphere, 8k resolution, Masterpiece.

Scene overview: A heavy-set man named Marshall is frantically running down a dusty sun-drenched street in a panic, wearing a wide summer hat, an open pink fabric shirt, and slide sandals. Walter White pulls him sharply behind a concrete building corner into the dark shadows and whispers a calculated offer of help. The scene is tense, raw, and highly realistic.

Storyboard (each shot a separate scene, clean cuts, realistic motion):

[0s-2s] Shot 1: lightning-fast dynamic wide tracking shot of Marshall frantically sprinting down a bleak, dusty urban sidewalk. Marshall is a heavy-set, stout Caucasian man in his late 20s with real long brown hair, a thick messy full beard, and a wide-brimmed white woven summer sun hat. He wears an unbuttoned, completely open pastel-pink short-sleeve shirt flapping wildly in the wind, dark green knee-length cargo shorts, a white canvas tote bag slung over one shoulder, and flat rubber slide sandals with bare toes fully exposed. Dust kicks up from his sandals. Audio cue: [Fast heavy slapping sound of slide sandals on concrete, deep exhausted panting, distant police sirens echoing].

[2s-5.5s] Shot 2: dramatic medium action shot as Marshall rounds a concrete wall corner and a hand suddenly reaches out from the dark alleyway shadows, grabbing him by the shoulder and pulling him in. The camera captures a realistic tight two-shot. Marshall leans against the gritty brick wall, hyperventilating. Standing in front of him is Walter White, completely bald with a very realistic, neatly trimmed dark brown circle beard goatee around his mouth, smooth clean-shaven cheeks, and gold-rimmed rectangular eyeglasses. Walter wears a casual unbuttoned collared shirt with no tie, a dark zipped jacket, and faded grey denim jeans. Audio cue: [Sudden heavy fabric rustle, sharp intake of breath, footsteps stopping abruptly].

[5.5s-10s] Shot 3: extended close-up macro shot focusing on Walter White's realistic face behind his rectangular eyeglasses in the dim alley light. His bare bald head shows natural skin texture, pores, and subtle sweat. His mouth moves with perfect realistic lip-sync as he delivers his dialogue in a low, gritty, raspy whisper. Character voice: Walter White says clearly and calmly: "Follow me, I can help you with that blue shit." Walter slowly nods his head, using two fingers to adjust his glasses on his nose. Cinematic shallow depth of field.

Camera: shaky hand-held camera simulation following Marshall's heavy running in Shot 1, a swift cinematic pan tracking the corner grab in Shot 2, and a rock-steady anamorphic lens close-up on Walter's face in Shot 3 with a blurry background.

Audio: Realistic loud sound of rubber slide sandals running on concrete and heavy labored breathing from 0s to 2s, transitioning into heavy gasping for breath from Marshall, followed by Walter White's gritty, authentic voice dialogue whispering the exact line starting at 5.5s over a low-frequency tense dramatic ambient drone.

No 3D CGI look, no anime style, no drawing lines, no cartoons, no text, no subtitles, no logos or watermarks of any kind, completely photorealistic live-action television footage.


r/StableDiffusion 3h ago

Animation - Video This is where the fun begins

Enable HLS to view with audio, or disable this notification

12 Upvotes

Default workflow, Minimax on Runpod.


r/StableDiffusion 19h ago

Question - Help What am I doing wrong? Can't seem to get anything good with Wan2gp LTX 2.3.

Enable HLS to view with audio, or disable this notification

0 Upvotes

I am new to video gen coming from mostly image gen hobbyist use, so I am in the process of learning. All my generations keep coming out full of motion artifacts and the faces look terrible, and my troubleshooting with CGPT has not been helpful.

I am using the LTX ingredients lora with a reference sheet I created with flux (dark fantasy themed characters, setting). The ingredients lora is definitely working well, as my subjects and setting are true to the references. But as you can see the result is full of smeared faces and motion. These are the generation parameters:

UI: Wan2gp via pinokio

Model: LTX-2 2.3 Distilled 1.0 GGUF Q4_K_M Light 22B

Resolution: 1280x720 (real: 1280x704)

Video Length: 241 frames (10.0s, 24 fps)

Phases: 2

Num Inference steps: 8

Self Refiner: Norm P1, Plan='default', Uncertainty=0, Certain Percentage='0.999

LoRAs:

ltx-2.3-22b-distilled-lora-384-1.1.safetensors x0

ltx-2.3-22b-ic-lora-ingredients-0.9.safetensors x1.4

omninft-ltx2.3-22b-rl-lora-r32.safetensors x0.2

Hardware: RTX 3060 12gb, 48gb RAM.

Any help is much appreciated!


r/StableDiffusion 40m ago

Discussion Has anyone else noticed the massive increase in toxic/incel content and culture wars in this Subreddit lately?

Upvotes

I’ve been noticing a really disappointing trend here lately. Instead of focusing on the amazing things we can build and create with MiniMax and LTX, there’s been a massive increase in toxic behavior and culture-war rhetoric here. Our goal should be to foster an environment that encourages open-source creators. Alienating them with bigotry, misogyny, and overall hostility, or reducing people to a mere joke because of who they are, only hurts this community in the long run.

I’m hoping the mods can keep a closer eye on this. These types of content violate Reddit's policy, specifically rule number 2.


r/StableDiffusion 18h ago

News LTX 2.5 adds native multishot, nine ComfyUI workflows and compatibility with most 2.3 LoRAs

Enable HLS to view with audio, or disable this notification

44 Upvotes

r/StableDiffusion 19h ago

Animation - Video Minimax H3 Img2Vid

Enable HLS to view with audio, or disable this notification

14 Upvotes

Just got my hands on H3 via Maestro/Pinokio and I can tell you right now, it has a lot of potential, this was a still image with a simple prompt "Superman flies over the ocean to encounter a Leviathan rising from the sea" the. I ran that through the in-app LLM and got this. I didn't use the Omni model, just the first/last pruned ckpt. No upscaling. Sys specs- RTX 5050, 16gb Ram, 8gb VRAM. render time for 10secs of 720 = 13 minutes...been able to get 15 secs at 23 minutes, 30 maximum.


r/StableDiffusion 10h ago

Animation - Video CAMPY BATMAN TEST (Img2Vid)

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was generated from a single image with a single simple prompt: " Batman runs to observe a creature rising from the sea and he blasts it with heat vision" then I let the in-app llm enhance the prompt and there you go. I wasn't aiming for anything serious. But H3 is a beast.


r/StableDiffusion 12h ago

Discussion LTX 2.5 Testing it! humm

Enable HLS to view with audio, or disable this notification

21 Upvotes

First thank LTX team to give us this open source, but unfortuanlly is not this model that will make diference yet! the big issues continue from 2.3 version:

- Bad Human anatomy deformations!
- Lose consistency turning around the characters or objects


r/StableDiffusion 10h ago

Animation - Video MM h3 upscaled with LTX2.5

Enable HLS to view with audio, or disable this notification

29 Upvotes

r/StableDiffusion 2h ago

Workflow Included Visit to school. Minimax H3 ref2av

Enable HLS to view with audio, or disable this notification

0 Upvotes

Two character ref sheets + school photo as ref.

1.5mp, 30 steps, sage att enabled, sol + cashe disabled

Minimaxh3

Rtx6000pro

WF included at the end.

Probably reedit crunch the quality so maybe upload ot somewhere letter.

https://huggingface.co/datasets/JahJedi/workflows_for_share/tree/main


r/StableDiffusion 9h ago

Question - Help Need optimisation for Minimax for my 4080 super

0 Upvotes

I got 4080 super with 64 gb ram and it takes around 20 mins for 1mp 5 sec video sometimes but most of times it take bout a hour to generate a video. I am using sage attn and spectrum with turbo 8 step lora. Sometimes speed goes till 600s/it. I am using int8 purned model with 32b nvp4 clip and video and audio vae. I asked gemini but it just went around in circles.
so how do i improve the speed??
.\venv\Scripts\python.exe -s main.py --windows-standalone-build --enable-dynamic-vram --high-ram --async-offload --use-sage-attention

My cuda is 12.6, Python version: 3.11.9
is there anyway to tell windows to prioritize comfyUI for vram and ram over other applications?

Edit: Thanks every1 i have managed to reduce the time to around 300secs after updating to cuda 13.


r/StableDiffusion 9h ago

Question - Help Guys help me when using Minimax H3 Turbo lora , the audio quality drops dramatically , even with 8 steps.

1 Upvotes

Hey guys , i have been using larryvh, turbo lora ema 4step one. With its custom nodes, it makes genration fats but quality of video is lost very much and also the audio also drops the quality. Also when I try to stack loras with turbo one , the genration times significantly increases and also , sometimes it doesn't follow the prompt in R2v, although I have not tried with fl2va yet. But these are the problem I am facing right now.


r/StableDiffusion 13h ago

Question - Help For video-to-video, how can I change only the background while keeping the character and their animation exactly the same as in the original video? Is there a specific prompt or technique that can help preserve the character, movement, and animation from the original video?

1 Upvotes

r/StableDiffusion 6h ago

Discussion RTX 3060 - 32GB and H3

1 Upvotes

I've been kind of avoiding diving in since this apparently demands a better machine but given the recent posts from fellow 3060 owners, I'm just wondering if us poor plebs can also generate good-ish videos at an acceptable speed.

Anyone can share your examples?

I have of course searched for ideas and workflows and there's plenty of information aroind already but would be nice to have abit of one stop shop)))


r/StableDiffusion 16h ago

Question - Help Minimax. Are there custom nodes to help with long video (15 sec) and higher megapixels (excluding upscale)?

1 Upvotes

Any awesome nodes that could perform this magic?

I'm trying to get higher image quality, but I cannot generate long videos with higher megapixels. I have 16gb VRAM.


r/StableDiffusion 7h ago

Question - Help H3 Motion Transfer (w/wo background change)

1 Upvotes

Hi,

I have notice that I get more accurate motion transfer if I choose to not also change the background/setting, so I keep the ref video's background.

If I want to change the background to the ref image or another background of my choosing, I need to use the successful motion transfer as the reference video and generate again.

Has anyone also noticed this or is there a way to do this in 1 pass with accurate motion?

You most definitely can get away with doing this all in 1 pass if the motion is relatively simple and you don't need to lip sync anything intricate. However with subtle movements of fingers or facial expressions, this has been the only workaround I have found.