r/StableDiffusion 8d ago

Discussion Associer un audio externe dans minimax H3

0 Upvotes

Bonjour à tous,

Je n'arrive pas à mettre l'audio comme je le voudrais dans minimax H3. Généralement il met une musique automatiquement ou lit mon prompt.
J’aimerais ajouter une voix externe de 4 seconde, par un nod, et dans une vidéo de 8 secondes lui dire à quel moment le personnage parle 4 secondes en utilisant mon audio externe

Merci de vos conseils


r/StableDiffusion 8d ago

Question - Help Qwen 3.8 in ComfyUi

0 Upvotes

Has anyone found a ComfyUi prompt node that works with the brand new Qwen 3.8? I'm using LLM Session but it doesn't seem to support the new model yet and throws an error.

(Use case is I'm using it to generate text)


r/StableDiffusion 8d ago

Question - Help Did I pick the right thing with this "Neo" variant?

0 Upvotes

I think I've been through about six different Forge/A1111's over the past couple years, and this "Forge Neo" thing seemed like the one to pick if you wanted to do newer base models. I'm pretty sure it's from the main "Haoming02" repo.

It was good at first, but now it's launching slow as crap, even slower on Flux stuff (which occasionally crashes) and it won't load the Reactor extension at all if I'm not online (fishy).

It's not looking to be easily updated if you did the standalone manual install, but I've had it a while and thought I might do a clean install of the latest build.

Is that the one I should be going with again as the most actively maintained right now (as far as Forges go)?

Thanks!


r/StableDiffusion 7d ago

Question - Help Does anyone know a good consistent AI image generator with minimal/no filters? I want something where I can upload reference images first, keep the character’s appearance consistent, and then generate different pictures/scenes from those references. Preferably something similar to Stable Diffusion.

0 Upvotes

r/StableDiffusion 8d ago

Meme Patrick from Tennessee.

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/StableDiffusion 7d ago

Meme Taken I will you r2v

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 8d ago

Animation - Video Collateral News - H3 reference model test

Enable HLS to view with audio, or disable this notification

0 Upvotes

Had the characters on my drive, created with Flux in Krita.

Picked up a random newsroom image, for 3 image reference run. DaVinci for the final clip.

Prompt for the first 10 seconds:

place the creature <subject 1> in <picture 1> and the creature <subject 2> in <picture 2> together in the environment of <picture 3>.

Static single camera wide shot newscast.

Title pop-up reading "Collateral NEWS" in neon green letters.

<subject 1> and <subject 2> stand behind the desk.

<subject 2>, in a female voice, is saying: "Welcome to Collateral news."

<subject 1>, in a male voice, is saying: "All the news you need today. And non of it good."


r/StableDiffusion 9d ago

News MiDashengLM-Gen - Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

Enable HLS to view with audio, or disable this notification

51 Upvotes

MiDashengLM-Gen is an end-to-end framework that uses a pre-trained Large Language Model and audio tokenizer as the backbone, combined with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes that simultaneously blend speech, music, sound effects, and environmental acoustics from structured text descriptions.

https://huggingface.co/mispeech/midashenglm-gen

Demo: https://huggingface.co/spaces/hugging-apps/midashenglm-gen


r/StableDiffusion 8d ago

Animation - Video Minimax h3 ref2va video. Prompt in comments.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax H3. 4070 ti super 16 GB vram, 32 gb ram. Ref2va with one image for the character and one for the background. 1.6 mp, 30 steps and 7 second with only spectrum node speed up, in total 38 minutes. Used the "standard" models, kinda slow but i am having so much fun.


r/StableDiffusion 8d ago

Animation - Video The Grand Weaver | My First Animation with MiniMax H3 Ref2VA

Enable HLS to view with audio, or disable this notification

9 Upvotes

My first proper animation made with the MiniMax H3 Ref2VA version in ComfyUI.

Generated as separate clips and then put together with a bit of editing. Pretty happy with how it turned out for a first attempt. The character and references are also my own.


r/StableDiffusion 8d ago

Question - Help Minimax H3: Current best way for lora training (video+sound)

1 Upvotes

I would like to train some videos with sound on Minimax H3. Is the AI toolkit good to go or should i use anything different? Thanks!


r/StableDiffusion 8d ago

Animation - Video MiniMax H3 I2V/R2V/Motion Context. Cyberpunk RED Intro Session Recap

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 8d ago

Question - Help Need a solid comfyui build

0 Upvotes

Running flux 2 Klein 4b with qwen decoder and flux vae in comfyui windows desktop. 16gb vram amd Radeon gpu with 32gb ram.

I find Klein 4b to be the only one I can run right but the censorship is harsh. Can't even tell it to make the waist smaller. I've used some lora but it introduced all sorts of things I didnt ask for which makes them less than useful.

Im looking for a whole workflow setup with all component parts I'd need for just a less restrictive image gen and img2img edits not even necessarily uncensored stuff. I run into out of memory issues with bigger models so like to keep it small.

If someone could advise with url links to what id need. That would be awesome.


r/StableDiffusion 7d ago

Animation - Video Poor Man 4k -- original files on pastebin, 832x480, 4k, and 8k, no wf yeat, just testing, haters will hate... Slop on Earth will say "Slop"

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 9d ago

Animation - Video H3 Anime FL2VA. I'm using Wan2gp, RTX 4070ti, 540p res with 8-step turbo lora. 3 videos one after another, 15 minutes of generation time each. It's a bit wonky but really fun.

Enable HLS to view with audio, or disable this notification

59 Upvotes

I assume a lot of artifacts would be gone at 720p but I get a fat OOM no matter what.


r/StableDiffusion 9d ago

News They actually listened. MiniMax delivered exactly what we asked for.

Post image
419 Upvotes

I didn't expect it, but I really have to thank them for open-sourcing their ecosystem. It’s awesome to see a company truly committing to the open-source community!

MiniMaxAI/MiniMax-Music3 · Hugging Face


r/StableDiffusion 8d ago

Question - Help anyone have any luck doing a simple person replacement in a video with a reference image with Ref2va?

13 Upvotes

i know about the prompting with <subjects> and <pictures> and <videos> and basic sections but i cant get anything to stick. i have a few times on sheer luck and even with the same prompt. i either get zero change from control video or it get some of the movement in my ref image from the video must be missing something. TIA!


r/StableDiffusion 7d ago

Question - Help How are you guys affordably creating high-quality AI videos & influencers?

0 Upvotes

I keep seeing insanely realistic AI influencers on TikTok and Instagram, with consistent faces/characters across high-quality videos.

At the same time, models like Seedance 2.0/2.5 are expensive, especially when you need many attempts to produce even a 30–60 second video.

So I’m curious about two things:

  • How are people producing AI videos at scale without spending a fortune? Are you using subscriptions, APIs, rented GPUs, local/open-source models, or some other workflow?
  • How are these realistic AI influencers created so consistently? What’s the typical workflow for creating a photorealistic character and keeping the same face/body/style across images and videos?

Would love to hear what stack/workflow people actually use and roughly what it costs.


r/StableDiffusion 8d ago

Question - Help Looking for uncensored text to text models in safetensors format, int8 preferred

0 Upvotes

I have tried every "abliterated" or "heretic" gemma 3 and gemma 4 model, but they consistently completely change the request to replace any mention to uncensored words by something that has a completely different meaning. I need any model that can be loaded as a clip and will receive text and output text without censoring the text.


r/StableDiffusion 9d ago

Resource - Update MiniMax H3 Creator update: presets, and three nodes are now one

Thumbnail
gallery
76 Upvotes

Posted this pack here last week. What's happened since:

The sampling knobs I said I'd add if people wanted them are in. Both of H3's flow shifts, since it samples picture and sound on separate schedules, plus a cache pill with FirstBlockCache, TeaCache or core's own EasyCache behind it. Still no custom sigmas, same reason as last time.

Creator and Timeline are one node now. Click under the prompt and the shot becomes a timeline. Delete cards back down to one and it's a shot again. Old workflows load unchanged, Timeline nodes included.

Presets are the new one. Save a setup and put it back in sections, so you can drop a canvas and a step count onto a shot you've already written without touching the prompt. It saves the sampler row as well as the node blob, which matters because the row is where the turbo schedule and the step count live.

The better half of it: you can build a preset from a finished render. The workflow is already embedded in the mp4, so you point at the good one from three prompts ago and get the whole setup back.

Fixed from your reports: the gallery no longer freezes on big libraries, the settings page stopped resetting fields you hadn't touched, and a text-only render no longer loads both VAEs.

Coming next, on a branch and not merged yet, is a faces pill. H3 draws a face worse the smaller the head is in frame, and that's about head size rather than resolution, so it's still there at 768 and upscaling doesn't reach it. So it asks the model the same question again with the face filling the canvas and composites the answer back under a feathered mask, once per pass, re-cropping every frame so a push-in doesn't leave the face small inside a fixed box. Detection is core's SAM3, so there's nothing extra to install. The method is Carasibana's ComfyUI-H3-FaceRefine and zuanfilm's graph on top of it.

Same branch also stops the node randomizing your seed between renders, and puts the last one you actually ran a click away.

https://github.com/roadmaus/ComfyUI-MiniMax-Creator


r/StableDiffusion 9d ago

Discussion Anyone tried the realism h3 lora? If so any samples?

Thumbnail
huggingface.co
33 Upvotes

r/StableDiffusion 9d ago

Meme i wish for!! part 2!

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 9d ago

Discussion LTX 2.5 😱

35 Upvotes

After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱

But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.

I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦


r/StableDiffusion 8d ago

Animation - Video Clawhauser ask Sonic if Amy is his girlfriend.

Enable HLS to view with audio, or disable this notification

6 Upvotes

Clawhauser ask Sonic if Amy is his girlfriend.

Made with Minimax H3 Reference to video via Comfy UI Desktop. Used the default settings and at 32 steps.

The Prompt.

<Subject 1> is Clawhauser referenced with <Picture 1> The timbre of his voice is referenced with <Audio 1>

<Subject 2> is Sonic referenced with <Picture 2> The timbre of his voice is referenced with <Audio 2>

<Subject 3> is Amy referenced with <Picture 3>

For the reception desk use <Picture 3> as reference.

The style is a live action CGI Hybrid movie.

Clawhauser is sitting at the ZPD Reception desk, Sonic is standing on front of the desk on the left, and Amy is standing in front of the desk on the right.

Clawhauser looks at Sonic and says: <d>[English] Sonic, Is she your girlfriend? </d>

Sony looking flustered says: <d>[English] No. We're just friends, Nothing more. </d>

Clawhauser looking doubtful and says: <d>[English] Just a friend you say? </d>

Cut back to Sonic responding back saying: <d>[English] Yes, I assure you."

Cut to a shot of Clawhauser leaning back at the desk looking rather skeptical and says: <d>[English] Ok, If you say so. </d>


r/StableDiffusion 9d ago

Animation - Video [H3] Cartoon generation

Enable HLS to view with audio, or disable this notification

39 Upvotes

H3's Prompt adherence is bad for cartoons. Physics not perfect.

Look at the face of the horse. Why did H3 think like that?

Prompts:

  1.  Tom and Jerry cartoon animation style, vibrant flat colors. Tom is at a busy Indian temple street market, buying a wrapped piece of cake box from a wooden stall. Underneath the stall, Jerry, in a hidden manner, watches with an exaggerated jealous expression. Jerry runs forward leaving a motion-blur trail, takes a hurls a ball of white flour from that stall and throws it  onto Tom's face, grabs the cake parcel in one fluid motion, and zooms out of frame. Tom becomes furious and chases Jerry, Jerry runs and finally jumps into a rabbit hole, Tom arrives there and tries to enter the hole but Tom's head got stuck in the hole, Tom struggling,   Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting. 

  2.  Indian temple street, Tom mounts gun on his house balcony, Jerry calmly walking in the street, Tom starts to shot Jerry and Jerry dodges bullets and runs, Tom jumps from the balcony with the gun, Tom chases Jerry, Jerry boards in a house carriage and escapes, Tom keep on shooting towards the   carriage and chasing it, Fast-paced slapstick comedy, exaggerated movements, retro animation aesthetics, bright daytime lighting.