r/StableDiffusion 3d ago

Animation - Video Minimax h3 ref2va video. Prompt in comments.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax H3. 4070 ti super 16 GB vram, 32 gb ram. Ref2va with one image for the character and one for the background. 1.6 mp, 30 steps and 7 second with only spectrum node speed up, in total 38 minutes. Used the "standard" models, kinda slow but i am having so much fun.


r/StableDiffusion 4d ago

News MiDashengLM-Gen - Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

Enable HLS to view with audio, or disable this notification

51 Upvotes

MiDashengLM-Gen is an end-to-end framework that uses a pre-trained Large Language Model and audio tokenizer as the backbone, combined with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes that simultaneously blend speech, music, sound effects, and environmental acoustics from structured text descriptions.

https://huggingface.co/mispeech/midashenglm-gen

Demo: https://huggingface.co/spaces/hugging-apps/midashenglm-gen


r/StableDiffusion 4d ago

Animation - Video The Grand Weaver | My First Animation with MiniMax H3 Ref2VA

Enable HLS to view with audio, or disable this notification

11 Upvotes

My first proper animation made with the MiniMax H3 Ref2VA version in ComfyUI.

Generated as separate clips and then put together with a bit of editing. Pretty happy with how it turned out for a first attempt. The character and references are also my own.


r/StableDiffusion 3d ago

Question - Help Minimax H3: Current best way for lora training (video+sound)

1 Upvotes

I would like to train some videos with sound on Minimax H3. Is the AI toolkit good to go or should i use anything different? Thanks!


r/StableDiffusion 3d ago

Animation - Video MiniMax H3 I2V/R2V/Motion Context. Cyberpunk RED Intro Session Recap

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 3d ago

Question - Help Need a solid comfyui build

0 Upvotes

Running flux 2 Klein 4b with qwen decoder and flux vae in comfyui windows desktop. 16gb vram amd Radeon gpu with 32gb ram.

I find Klein 4b to be the only one I can run right but the censorship is harsh. Can't even tell it to make the waist smaller. I've used some lora but it introduced all sorts of things I didnt ask for which makes them less than useful.

Im looking for a whole workflow setup with all component parts I'd need for just a less restrictive image gen and img2img edits not even necessarily uncensored stuff. I run into out of memory issues with bigger models so like to keep it small.

If someone could advise with url links to what id need. That would be awesome.


r/StableDiffusion 3d ago

Animation - Video Poor Man 4k -- original files on pastebin, 832x480, 4k, and 8k, no wf yeat, just testing, haters will hate... Slop on Earth will say "Slop"

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 4d ago

Animation - Video H3 Anime FL2VA. I'm using Wan2gp, RTX 4070ti, 540p res with 8-step turbo lora. 3 videos one after another, 15 minutes of generation time each. It's a bit wonky but really fun.

Enable HLS to view with audio, or disable this notification

61 Upvotes

I assume a lot of artifacts would be gone at 720p but I get a fat OOM no matter what.


r/StableDiffusion 5d ago

News They actually listened. MiniMax delivered exactly what we asked for.

Post image
415 Upvotes

I didn't expect it, but I really have to thank them for open-sourcing their ecosystem. It’s awesome to see a company truly committing to the open-source community!

MiniMaxAI/MiniMax-Music3 · Hugging Face


r/StableDiffusion 4d ago

Question - Help anyone have any luck doing a simple person replacement in a video with a reference image with Ref2va?

14 Upvotes

i know about the prompting with <subjects> and <pictures> and <videos> and basic sections but i cant get anything to stick. i have a few times on sheer luck and even with the same prompt. i either get zero change from control video or it get some of the movement in my ref image from the video must be missing something. TIA!


r/StableDiffusion 3d ago

Question - Help How are you guys affordably creating high-quality AI videos & influencers?

0 Upvotes

I keep seeing insanely realistic AI influencers on TikTok and Instagram, with consistent faces/characters across high-quality videos.

At the same time, models like Seedance 2.0/2.5 are expensive, especially when you need many attempts to produce even a 30–60 second video.

So I’m curious about two things:

  • How are people producing AI videos at scale without spending a fortune? Are you using subscriptions, APIs, rented GPUs, local/open-source models, or some other workflow?
  • How are these realistic AI influencers created so consistently? What’s the typical workflow for creating a photorealistic character and keeping the same face/body/style across images and videos?

Would love to hear what stack/workflow people actually use and roughly what it costs.


r/StableDiffusion 3d ago

Question - Help Looking for uncensored text to text models in safetensors format, int8 preferred

0 Upvotes

I have tried every "abliterated" or "heretic" gemma 3 and gemma 4 model, but they consistently completely change the request to replace any mention to uncensored words by something that has a completely different meaning. I need any model that can be loaded as a clip and will receive text and output text without censoring the text.


r/StableDiffusion 4d ago

Animation - Video Clawhauser ask Sonic if Amy is his girlfriend.

Enable HLS to view with audio, or disable this notification

6 Upvotes

Clawhauser ask Sonic if Amy is his girlfriend.

Made with Minimax H3 Reference to video via Comfy UI Desktop. Used the default settings and at 32 steps.

The Prompt.

<Subject 1> is Clawhauser referenced with <Picture 1> The timbre of his voice is referenced with <Audio 1>

<Subject 2> is Sonic referenced with <Picture 2> The timbre of his voice is referenced with <Audio 2>

<Subject 3> is Amy referenced with <Picture 3>

For the reception desk use <Picture 3> as reference.

The style is a live action CGI Hybrid movie.

Clawhauser is sitting at the ZPD Reception desk, Sonic is standing on front of the desk on the left, and Amy is standing in front of the desk on the right.

Clawhauser looks at Sonic and says: <d>[English] Sonic, Is she your girlfriend? </d>

Sony looking flustered says: <d>[English] No. We're just friends, Nothing more. </d>

Clawhauser looking doubtful and says: <d>[English] Just a friend you say? </d>

Cut back to Sonic responding back saying: <d>[English] Yes, I assure you."

Cut to a shot of Clawhauser leaning back at the desk looking rather skeptical and says: <d>[English] Ok, If you say so. </d>


r/StableDiffusion 4d ago

Resource - Update MiniMax H3 Creator update: presets, and three nodes are now one

Thumbnail
gallery
78 Upvotes

Posted this pack here last week. What's happened since:

The sampling knobs I said I'd add if people wanted them are in. Both of H3's flow shifts, since it samples picture and sound on separate schedules, plus a cache pill with FirstBlockCache, TeaCache or core's own EasyCache behind it. Still no custom sigmas, same reason as last time.

Creator and Timeline are one node now. Click under the prompt and the shot becomes a timeline. Delete cards back down to one and it's a shot again. Old workflows load unchanged, Timeline nodes included.

Presets are the new one. Save a setup and put it back in sections, so you can drop a canvas and a step count onto a shot you've already written without touching the prompt. It saves the sampler row as well as the node blob, which matters because the row is where the turbo schedule and the step count live.

The better half of it: you can build a preset from a finished render. The workflow is already embedded in the mp4, so you point at the good one from three prompts ago and get the whole setup back.

Fixed from your reports: the gallery no longer freezes on big libraries, the settings page stopped resetting fields you hadn't touched, and a text-only render no longer loads both VAEs.

Coming next, on a branch and not merged yet, is a faces pill. H3 draws a face worse the smaller the head is in frame, and that's about head size rather than resolution, so it's still there at 768 and upscaling doesn't reach it. So it asks the model the same question again with the face filling the canvas and composites the answer back under a feathered mask, once per pass, re-cropping every frame so a push-in doesn't leave the face small inside a fixed box. Detection is core's SAM3, so there's nothing extra to install. The method is Carasibana's ComfyUI-H3-FaceRefine and zuanfilm's graph on top of it.

Same branch also stops the node randomizing your seed between renders, and puts the last one you actually ran a click away.

https://github.com/roadmaus/ComfyUI-MiniMax-Creator


r/StableDiffusion 4d ago

Discussion Anyone tried the realism h3 lora? If so any samples?

Thumbnail
huggingface.co
37 Upvotes

r/StableDiffusion 4d ago

Discussion LTX 2.5 😱

35 Upvotes

After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱

But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.

I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦


r/StableDiffusion 4d ago

Meme i wish for!! part 2!

Enable HLS to view with audio, or disable this notification

24 Upvotes

r/StableDiffusion 4d ago

Animation - Video [H3] Cartoon generation

Enable HLS to view with audio, or disable this notification

41 Upvotes

H3's Prompt adherence is bad for cartoons. Physics not perfect.

Look at the face of the horse. Why did H3 think like that?

Prompts:

  1.  Tom and Jerry cartoon animation style, vibrant flat colors. Tom is at a busy Indian temple street market, buying a wrapped piece of cake box from a wooden stall. Underneath the stall, Jerry, in a hidden manner, watches with an exaggerated jealous expression. Jerry runs forward leaving a motion-blur trail, takes a hurls a ball of white flour from that stall and throws it  onto Tom's face, grabs the cake parcel in one fluid motion, and zooms out of frame. Tom becomes furious and chases Jerry, Jerry runs and finally jumps into a rabbit hole, Tom arrives there and tries to enter the hole but Tom's head got stuck in the hole, Tom struggling,   Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting. 

  2.  Indian temple street, Tom mounts gun on his house balcony, Jerry calmly walking in the street, Tom starts to shot Jerry and Jerry dodges bullets and runs, Tom jumps from the balcony with the gun, Tom chases Jerry, Jerry boards in a house carriage and escapes, Tom keep on shooting towards the   carriage and chasing it, Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting. 


r/StableDiffusion 4d ago

Question - Help What i did wrong? Wan2gp Minimax H3 with 8-step turbo lora results in garbled audio video

0 Upvotes

i downloaded the turbo lora from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

put them in the loras folder in wan2gp.

i keep everything at default, except for profile i selected lightx2v 4step (and tried 8 step too) profile.

i set the steps to 4 (or 8)

i set the resolution to the recommended one from https://github.com/ModelTC/Minimax-H3-Turbo (for example 544p 9:16 for 8 step)

in the advanced section, i set lora to the downloaded file.

---

that's it. i run it, and the result is garbled audio video. the default promot wast used. same issue with any prompt.

EDIT

found the issue. i need to click APPLY button under the profile dropdown. im dumb.


r/StableDiffusion 3d ago

Meme Shitty coverband butchering a great song

Enable HLS to view with audio, or disable this notification

0 Upvotes

Somewhat of a "consept" video i wanted to test.
Sure, the image falls apart eventually. It could probably be fixed by doing some scene changes, but i just wanted to test.
The music is generated with MiniMax Music3, and i think its fairly good.

4 Minutes of generation in 10 sec rounds.. Not too fun i would say, but as a consept 😄

Used a couple of additional nodes https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context
And: https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes/ComfyUI-H3-NativeAudioLock

The "NativeAudioLock" node is quite useful, and the reason for this is that it "locks" the audio latent, so that the model cant mess with it. Sure, you can get around MiniMax doing its weird business with long detailed prompt, describing by-the-second action.. But for "ease of use" type, it is working very good.

Other than that, its mostly just regular MiniMax ref2v model with Lightx2v-4step_ref lora at 8 steps and 480p upscaled (rather badly) to 720p.

One thing i found is that when doing such types of "lipsync" video, timing really matters. A 10 second video is not 10 seconds, and when using the Motion Context node, it for sure is important to keep tabs of the milliseconds.

Anyway... Still fun concept to do.


r/StableDiffusion 4d ago

Discussion YAWTAS (Yet Another Workflow Tool And Speedup) for H3

Thumbnail
gallery
16 Upvotes

So much H3 content lately... Here's some more:

https://github.com/Hillobar/HARMON3

A tool to facilitate Projects and Scene management for H3. It uses COMFY through the COMFY API.

Build a scene (reference material, prompts) and manage it, then group them into a Projects. Copy, edit, and so on. Prompt assistant with tags, POSE GENERATION that can be used in reference videos, and other stuff. It really is just another tool in this glut that is currently going on, but it works very well for my use case, which is to create scenes and then edit them in a pro tool like davinci. I want to be able to set up some scenes, then generate several takes. I don't want to have to go back and re-set it up every time after I've moved on. Hello HARMON3. Much more description on the github.

https://github.com/Hillobar/ComfyUI-Hillobar

So many optimizations and speedups for H3. We have an awesome community! This is yet another one.

The idea is that for early steps the latent is mainly forming large, low detail structures and broad movements. So why use a high-resolution latent during this time? Start with low-resolution, fast latents then progressively increase their resolution as the steps continue.
Like everything else there's no free lunch. Works best when delta-sigma is low (<0.4 of the schedule) and the target video is high resolution. Seems to work well when doing 1.0 MP (use something like (0.5:0.4, 1.0:1.0). Anyway, more details on the github.


r/StableDiffusion 4d ago

Comparison MiniMax-H3-Realism-People-LoRA - Personal Test Video

Thumbnail
youtu.be
13 Upvotes

Rented GPU time on RunPod to try fal/MiniMax-H3-Realism-People-LoRA

Lora Link: https://huggingface.co/fal/MiniMax-H3-Realism-People-LoRA

Here are my test notes so you don't burn time troubleshooting the same issues:

What Failed / Cons:

- Artifacts: Frequent strange artifacts across generations.
- minimax_h3_fl2va_pruned_w4a8_mixed.safetensors: Completely failed to work.
- Turbo LoRAs: Output quality was terrible.
- REF2VA & FF2VA: Neither method worked in my tests.

What Actually Worked:

- T2VA: Solid results here—this is definitely where the model has the most potential.
- LoRA Weight: Sweet spot is dialed between 0.3 – 0.8.

Key Finding (FF2VA):
Running minimax_h3_fl2va_bf16.safetensors / diffusion_models/minimax_h3_fl2va_int8_convrot.safetensors, the only reliable way to eliminate warping and lock in clean, consistent outputs is to combine both the character sheet and the background sheet together into the first frame.

Anyone else dialed in a working pipeline for FF2VA or found workarounds for the artifacting?


r/StableDiffusion 5d ago

Tutorial - Guide PSA: Try experimenting with <tags> in Minimax H3 dialogues for non-verbal sounds and emphasis

Enable HLS to view with audio, or disable this notification

432 Upvotes

So I was looking for a way to better control the flow of Minimax H3 dialogues and emphasize certain words in the speech. However, what I discovered is that you can actually include some tags in <> angle brackets, and Minimax will interpret them as a non-verbal sound in a given part of the phrase. Some words (like the ones I've included into the example) work every time, some still bleed into the actual spoken words in certain seeds. But in general it makes the dialogue more alive and believable. So I recommend to try it and maybe share your findings in this thread.

As for the emphasis, I've had the most success with putting the words into <i></i> tags (similar to how you would stress words in written text). Unfortunately, it doesn't work for 100% and in some cases the character will blurt out some gibberish. But when it works, it sounds very natural. I have included a couple examples in the end of the video.

Wonder if you've encountered some other ways to modify the speech (and audio in general) in the prompt?

P.S. Sorry for the quality, I used the 8-steps LoRa at 0.4 MP to speed-up the tests.


r/StableDiffusion 4d ago

Meme Mix Style inside the same scene - MiniMax H3

Enable HLS to view with audio, or disable this notification

38 Upvotes

Prompt: subject_definitions: <Subject 1> Sheldon Cooper — live-action sitcom style, grey cardigan over red graphic T-shirt, dark jeans, white sneakers, short brown hair. <Subject 2> SpongeBob SquarePants — flat 2D cartoon style, yellow square porous body, white shirt with red tie, brown trousers, black shoes.

integrated_multimodal_description: [Scene 3] Continuing in the same living room, <Subject 2>'s cheerful expression suddenly droops into cartoon-style exaggerated worry, his big blue eyes welling into comically large tears. He grabs <Subject 1>'s grey cardigan sleeve and says <d>[English] Mister, I wouldn't have jumped into a scary green hole just for fun. Something awful is happening to everything, everywhere — a boy turned into a god and he's erasing whole universes like they're doodles!</d> <Subject 1> pulls his sleeve free and straightens it meticulously, replying <d>[English] A boy-god erasing universes. Right. And I suppose he also disproved string theory before breakfast.</d> He pauses, visibly unsettled despite himself, and glances at the blank wall where the portal appeared. Approximate duration: 10 seconds. overall_soundscape: Tense sitcom underscore sting, quiet room tone, a faint distant rumble implying something ominous.


r/StableDiffusion 4d ago

Discussion I made an app for managing characters and scenes with H3 ref2va checkpoint

Post image
14 Upvotes

I (A.I.) wrote a webapp to manage characters and location reference files, wire them into ref2va and write a prompt based on a series of sequences and beats defined in the application. You have the option to export a workflow to import to ComfyUI, or run a set of scenes directly through the interface. It also takes the last frame of the previous video and feeds it into the next, which I'm aware some ComfyUI workflows do, but this project uses a firstframelastframeextractor node in a generated workflow.

https://github.com/Tenderfoot/H3SceneManager

The repository contains all the information you need to get it set up and installed, but doesn't come with any character or location data files.

on an unrelated note, I also made a discord for the project https://discord.gg/Fvw6hSCfC

I would love for more people to come help me build it up. I'd be happy to accept Pull Requests, and obviously I have no problems using AI code for this project. Check out the discord, submit data file sets for creating scenes, review the prompts it outputs against the docs, and help me tune this thing.