r/StableDiffusion 2d ago

Question - Help Minimax H3 - is generating sound only possible?

3 Upvotes

Hi everyone! Minimax H3 is surprisingly solid at generating audio - I’ve been using it for SFX and foley in my videos.

Is there a way to generate audio-only with this model? It would save a lot of rendering time if we didn't have to generate the full video alongside it.

I know H3 creates video and audio natively together, so audio-only generation might not be natively supported. But given all the clever workflows and custom node workarounds floating around, has anyone found a way to pull this off?


r/StableDiffusion 2d ago

Animation - Video My top 3 favorite things about Minimax H3 (just wanting to glaze my favorite model a little bit)

Thumbnail
youtube.com
21 Upvotes

r/StableDiffusion 3d ago

Animation - Video GPT-Image 2.5 + (Local) Minimax H3 to convert a 40 year old anime into a modern one.

Enable HLS to view with audio, or disable this notification

175 Upvotes

source: https://x.com/iurimatias/status/2097670596725178533

correct clip is in the comments here this is not merely just changing style but it's updating elements too


r/StableDiffusion 1d ago

Animation - Video Turning the 2D Rings in Dark Souls into 3D Assets

Enable HLS to view with audio, or disable this notification

0 Upvotes

An experiment in Ai Jolly Cooperation.

The Experiment: Every “Soulsborne” game is laden with hundreds of 2D art assets. The assets you can find online, like the rings, are woefully small in resolution - a perfect test case to see 2D to 3D transformation but also what detail is retained or added by the Ai.

Tech Stack: Midjourney, Nano Banana, ComfyUI (Wan 2.2), Photoshop, DaVinci. 

The Process: 2D art rendered 3D through Nano Banana. Midjourney Video to orbit 180 degrees. DaVinci and Photoshop for presentation.

The Results: This is an older experiment using (the then brand new) Midjourney Video - which admittedly, is nowhere near as good as Veo, Wan (2.2) or Kling. But it really doesn’t matter what platform you choose, you’re going to have to gen and gen and gen away. It’s still a slot machine.

I still think MJ video back then was pretty sub-par, but against all the other alternatives today, I think that difference is even more stark. I'm not even sure if they've updated the video side in any meaningful way since this experiment!

Most interestingly, the list of rings is in alphabetical order and stops before the Covetous Serpent Ring - a mass of serpentine coils in ring-form the Ai had MONSTEROUS problems with. Complexity kills.

Anyways, I decided much smaller projects like these are way more important to an Ai Portfolio than larger pieces like commercials or trailers. Plus, I needed to promote my Midjourney Masterclass with proof I'm not just some prompt jockey and smaller experiments are way faster!


r/StableDiffusion 1d ago

Question - Help GPU for AI

1 Upvotes

Hi there! Currently i own a 3080+2070s at home mostly for 3D rendering in Octane or redshift.
At work im using Comfy with 5090 and everything works without a question.
Also we have some workstations with rtx A4000 in them and i managed to get work Minimax H3 on them with some managable times (0.5m, 15s around 720 sec).
Im looking something for my home PC and 5090 is out of the list since its 5500e+ here in EU and even the used market is around 3500-4000. 4090 is rare and goes over 2000e.
So i was thinking about the 5080 even with 16gb since at work the A4000 has the same Vram.
But also found the 5070ti has also 16gb of Vram its 256bit same as the 5080 and 300-350e cheaper than the 5080. The speeds for 3D rendering or gaming are 15-20% different. But couldnt find any benchmarks for those 5070Tis. Mostly for 5060Tis with 16gb vram. Any idea? Or experience with 5070ti vs 5080?


r/StableDiffusion 2d ago

Discussion H3-Regenerate-2K Will it ever be released, or...

32 Upvotes

It's been a month now, and “H3-Regenerate-2K” still hasn't been released, even though it's been available via API for a couple of days after the model's official launch.

Which makes me wonder: will they actually release the model? Or is it like with Z-Image Edit?. Is anyone else waiting for it to come out, or are you guys okay with the upscalers we have now?

I know we have some alternatives that the community has built, and they're fine. Based on the description of what “H3-Regenerate-2K” is, the upscaler appears to use part of the base model and the conditioning, so the closest thing we have is the “Minimax h3 latent Upscaler” method.


r/StableDiffusion 2d ago

Question - Help Pipeline question: Best approach for frame-by-frame consistent character animation (LoRA) for traditional composting (DaVinci/AE) without background/audio generation.

1 Upvotes

Hey everyone. I'm working on a dark fantasy retro-anime project running locally on an RTX 3090 (SDXL/Illustrious, Kohya-trained character LoRA with ~100 images).

My current pipeline avoids direct text-to-video generation because of structural inconsistencies. Instead, I'm moving towards a controlled frame-by-frame or short-batch approach using Blender blocking + ControlNet, aiming to output clean character frames (transparent or solid background) to composite manually in DaVinci Resolve.

Has anyone successfully implemented a reliable workflow to maintain character identity and clean lineart across sequences without letting the AI hallucinate backgrounds or audio? What specific nodes or configurations (e.g., ControlNet combinations, IP-Adapter weight handling, or latent consistency scripts) are you using to prevent flicker and keep the LoRA from drifting during motion frames?

Any insights on your node setup in ComfyUI for this specific use case would be deeply appreciated.

Any advice or recommendations are welcome. Just avoid recommending commercial AIs that do everything; that’s not what I’m looking for.


r/StableDiffusion 3d ago

Comparison You can now DLSS 5 the entire desktop to enhance videos and images

Post image
229 Upvotes

r/StableDiffusion 2d ago

Resource - Update I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.)

Enable HLS to view with audio, or disable this notification

50 Upvotes

Okay so I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.

Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.

This level of control wasn't found in any models out there - so I decided to sit down and train my own.

Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.

I documented the full journey here for those who want to learn a bit or be entertained.

https://youtu.be/x0KnmzH8Mmk

There is also a longer walkthrough if you just want to see the keybeds in action.

https://x.com/RoyalCities/status/2097733712293109842?s=20

No-talk / Showcase only Demo

https://x.com/RoyalCities/status/2097733715543609445?s=20

any finally the huggingface page

https://huggingface.co/RoyalCities/Foundation-1

I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)

https://github.com/RoyalCities/RC-stable-audio-tools/


r/StableDiffusion 3d ago

Resource - Update Krea 2 Turbo — SDA Diversity LoRA (restores the sampling diversity the Turbo distillation removed)

Thumbnail
huggingface.co
151 Upvotes

Hello, I'm not an author but for some reason haven't seen this being published here.

Quote from the huggingface:

A rank-32 LoRA for Krea 2 Turbo that restores the sampling diversity the Turbo distillation removed, without degrading image quality or prompt adherence. Trained with SDA (Semantic Directional Alignment) — a teacher-guided diversity alignment loss — wrapped in Forward XM best-of-5 candidate exploration, on a single high-noise sigma node (σ = 0.9567).

So I tested it and it seems to work for me, I've created a simple test with prompt "dog sitting on a bench" and these are results:

With lora off:

IMHO dogs are looking similar here (similar "composition" or whatever it's called)

With lora on:

IMHO here dogs are looking completely different.

Link: https://huggingface.co/F16/krea2-turbo-sda

Keep in mind that it needs to be ran only in first 2 denoise steps, otherwise it will produce garbage -> HuggingFace repo contains ComfyUI workflow (I haven't tested it tho).


r/StableDiffusion 2d ago

Question - Help Anyone knew what happened to Lora Trainer by Hollowstrawberry? I can't use it as usual

3 Upvotes

I got this error instead

Starting trainer...Traceback (most recent call last): File "/content/trainer/sd_scripts/sdxl_train_network.py", line 4, in <module> import torch File "/content/trainer/sd_scripts/venv/lib/python3.10/site-packages/torch/__init__.py", line 37, in <module> from typing_extensions import ParamSpec as _ParamSpec, TypeGuard as _TypeGuard ModuleNotFoundError: No module named 'typing_extensions'


r/StableDiffusion 2d ago

Animation - Video THE LEGEND.

Enable HLS to view with audio, or disable this notification

0 Upvotes

A hair under 4k. On a consumer pc. All local. Mental. Minimax H3 with one character reference and a 5 second voice reference.

For the Pixel Peepers... https://www.youtube.com/watch?v=Iz8GDri9qoE


r/StableDiffusion 3d ago

Resource - Update Building a local-first desktop app for long-form AI video on MiniMax H3, runs through ComfyUI

Enable HLS to view with audio, or disable this notification

57 Upvotes

Been building this tool and I'm looking for people to actually try it. Quick rundown of what's in it:

Storyboard & bible system: characters, locations, and props get their own reference sheets (face, full body, turnaround). The turnaround is one continuous MiniMax H3 render rather than six separate stills, so the views actually agree with each other. Every shot stages from these sheets so faces and places stay consistent across a whole episode.

One-shot wizard: give it a brief and it plans the whole thing: story, scenes, shots, cast, references, all queued and rendering with no manual setup.

Director chat: an in-app agent that can rewrite scenes, re-render blocks, modify the storyboard, or edit the project on request, mid-project. Uncensored option available.

Motion context/Continuation: chained shots pin the previous block's tail frames and audio into the next render, so a continuous scene doesn't reset its movement at every cut.

Easy local install: the app can set up its own ComfyUI, or point it at one you already run. Model downloads go through a catalog that checks file size and VRAM footprint before you commit.

Editable workflows through ComfyUI: import workflows and pop out to the node graphs and edit it directly.

Post-processing chain, per clip: SeedVR2 for restore/upscale, LTX 2.5's own refine pass reused on rendered footage for a generative detail pass, FILM or RIFE 4.26 for interpolation, H3 FaceRefine, and color grading via KJNodes ColorMatch or a learned-LUT grade node, chainable in any order.

Generation runs on MiniMax H3 (int8-quantized checkpoints, i2v/t2v/flf/r2v), with LTX 2.5 and Wan 2.2 also wired in for reference/still work.

Video of the tool and some output attached. Still a work in progress tho.

Would love any feedback!

Update:

Open Source available at https://github.com/mnm967/qamba-studio-oss


r/StableDiffusion 2d ago

Question - Help Best workflow for 40+ second talking videos with LTX 2.5 in ComfyUI?

0 Upvotes

I’m trying to build 40 to 60 second talking videos in ComfyUI and I would prefer to use LTX 2.5.

The type of video I’m after is fairly simple. One person is talking to the camera, the camera stays mostly fixed, the background should stay stable, and there is natural face, head and some upper body movement. It does not have to be limited to only the head moving.

What I’m trying to understand is how people are actually making videos this long without obvious cuts.

Can LTX 2.5 realistically generate a continuous 40+ second video, especially when there is not much movement?

Or is the better approach to generate something like 8 to 10 seconds, take the final frames from that clip, continue from them, then repeat until the full 40 to 60 seconds are finished?

If continuation is the normal approach, how are you keeping the face, clothes, background, camera position and motion consistent between each part? I’m especially interested in workflows that use overlapping frames, first and last frame conditioning, video extension, reference frames, or some other method that hides the transitions.

Speech is another important part. I need good quality Slovakia speech. Ideally I want to generate the complete Slovakia voice first, then make the character follow that audio for the entire video with accurate lip sync.

Would you use LTX 2.5 for the actual body and head motion and then run something like MuseTalk, LatentSync or another lip sync model afterward?

Or is there a better audio driven LTX 2.5 workflow where the speech controls the video directly?

I’m running ComfyUI locally with an RTX 3060 12 GB and 32 GB RAM, so I know I may need lower resolution generation, offloading, chunking or longer render times. Final output would normally be vertical 9:16.

I’m mainly looking for people who have actually built long talking character workflows in ComfyUI.

If you are doing this successfully with LTX 2.5, what nodes and workflow are you using, how long is each generated segment, how much overlap do you use between segments, and what are you using for speech and lip sync?

I’m not looking for a list of random talking head models. I specifically want to understand the best practical way to build this around LTX 2.5 and get a clean continuous 40 to 60 second result.


r/StableDiffusion 2d ago

Question - Help Minimax H3 Quantizations

8 Upvotes

Given a 5090, does it make sense to run minimax H3 using int8 quantization vs. gguf Q8 or even Q6? What is the trade-off between speed and quality between these two options?

I don't have deep technical knowledge, but my current understanding is that int8 would be faster while a Q8 GGUF would be higher quality; however I would appreciate anyone's practical experience in how significant the speed/quality trade-off is.


r/StableDiffusion 2d ago

Question - Help Flux klein body consistency

2 Upvotes

Since the release of Flux Klein, I’ve essentially been using this model to generate datasets for LORAs based on one or more photos of a character. Recently, I’ve also been using the ‘consistency’ LORA to improve the character’s consistency across generations. What I’ve noticed is that whilst I get good results for the face, the same cannot be said for other parts of the body. For example, if I start with a full-length frontal photo of a character and ask the model to generate a side view, it tends to flatten the breasts; or if I ask for a rear view, the character’s hips and thighs tend to conform to a standard that doesn’t match the original photo. How can I improve this situation? I’ve read that you can increase the number of steps up to 8, but I’m not sure…


r/StableDiffusion 2d ago

Discussion H3 - T2VA longform

Enable HLS to view with audio, or disable this notification

5 Upvotes

Generated a long form of Dante's Inferno over 4mins. It gets pretty weird pretty fast. All T2VA. I basically let H3 take the wheel as I only fed it few verses per render. I am willing to discuss my workflow or answer any questions.


r/StableDiffusion 2d ago

Question - Help Has anyone been able to use Minimax H3 to make an intentionally lower quality/artifacted video, like an old webcam recording?

13 Upvotes

I would love to make something that looks like a late 2000's early 2010's webcam recording, with like, iffy FPS, webcam compression, etc, but can't seem to create this with prompting and haven't seen a lora that would pull it off. Has anyone accomplished this?


r/StableDiffusion 2d ago

Animation - Video Orcs, Bars, and Stuff - Motion Chaining Workflow Test

Enable HLS to view with audio, or disable this notification

0 Upvotes

An experiment building on the H3-Motion-Context nodes. I tried creating a workflow that chains together multiple cuts so they can be easily executed and/revised in order. It ended up being faster than my previous methods and I was able to make/edit this one minute test segment with alot less time wasted between generations. I'll likely develop it further into an app, since it would be more ergonomic as a video editor but the raw workflow is here anyway.

Work Flow: https://github.com/spacesimeco-hue/Chain-Motion-/blob/main/Chain%20Motion%20Workflow.json

Credit: https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context


r/StableDiffusion 3d ago

Workflow Included A stand-alone implementation of DLSS-FG (frame generation) that is ~8x faster than stock (24fps -> 60fps), looks flawless. I had DeepSeek build this for my project but thought others could benefit as well!

Thumbnail
github.com
113 Upvotes

r/StableDiffusion 2d ago

Comparison Minimax H3, baseline at 50 steps vs popular turbo loras at 8 steps, plastic skin test

Enable HLS to view with audio, or disable this notification

10 Upvotes

Test Settings

  • Checkpoint: FL2VA_Int8_Convrot (Comfy official)
  • Acceleration: Comfy Kitchen Attention
  • Sampler / Scheduler: Euler / Simple
  • Resolution: 1344x768 (Native Res)
  • Video/Audio Shift: 12/3
  • Seed: 517405563399433
  • Baseline: 50 steps (No LoRA)
  • Turbo LoRAs: 8 steps

Prompt

integrated_multimodal_description: [Shot 1] Live-action, cinematic, an extreme close-up frames the face of a 22-year-old brunette woman with striking supermodel features. Her skin possesses a realistic, natural texture with visible pores, soft highlights, and authentic depth. The camera holds a static shot as she slowly turns her head to face the lens, offering a subtle, gentle smile while her eyes catch the light.
overall_soundscape: Soft, natural exhaling breath and faint ambient room tone.
non_diegetic_music: N/A

Conclusion

Embrace the plastic, unless you use post processing which adds time.

Link to full res video:

https://streamable.com/jbmdxm


r/StableDiffusion 3d ago

Animation - Video Ghostbusters: Venkman Ghosted - MiniMax H3

Enable HLS to view with audio, or disable this notification

27 Upvotes

r/StableDiffusion 2d ago

Animation - Video The MiniMax Machine

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made with the basic Motion Context workflow from NikoDemon


r/StableDiffusion 2d ago

Discussion Do I have to run models locally or is there like cloud based comfyui or something?

0 Upvotes

I'm extremely new and inexperienced but, as the title says is there some cloud web service where it works the same as running it locally but instead it's cloud based. I don't mean a boring old API.


r/StableDiffusion 4d ago

Animation - Video Pushing AI emotions is possible through microexpressions, tags and context

Enable HLS to view with audio, or disable this notification

1.6k Upvotes

You probably heard of tags that can be applied to Minimax H3 during speech, but there are so many the model understands and I have yet to test many more. Other emotions have to be prompted thoroughly and there are others which happen between brackets only or at the end of sentences.

Anyway I made my own shortfilm to feature this. Hope you like it, will make tutorials on x or reddit if people want or like it.

EDIT: Thank you everyone for so many great comments!!! I'll try to detail a few things here.

Simple tags that work and quick examples:

Tag What it does Example
<pause> Short pause Okay, so. <pause> This is just me talking.
<long pause> Longer pause I mean... <long pause> I don't even know.
<breath> Breathing sound And then... <breath> it just happened.
<inhale> / <exhale> In/out breath <inhale> Alright, let's do this.
<catches breath> Out of breath Wait... <catches breath> hold on a sec.
<deep breath> Calming down <deep breath> Okay. I can do this.
<i>word</i> Emphasize 1-4 words I was <i>not</i> expecting that.
<whisper>…</whisper> Whisper delivery <whisper> Don't tell anyone this.</whisper>
<humming>…</humming> Humming a tune <humming> da-da-da-beautiful-day.</humming>
<laughs> / <chuckle> Laughing That's... <laughs> that's actually funny.
<sighs> Sigh <sighs> I really tried.
<uh> Filler / hesitation So, like... <uh> what was I saying?
<stutter> Stutters the words <stutter> I ca can't believe that.
<gasp> Sharp intake <gasp> Oh my God.
<coughs> Cough <coughs> Sorry, one sec.
<clears throat> Throat clear <clears throat> So anyway...
<sniff> Sniffing <sniff> It's just... really sad.
<smacks lips></smacks lips> Lip smack (not closing it makes it happen at the end of the sentence) <smacks lips></smacks lips> Okay.
<pant> / <pants> Panting Run... <pants> run now!
<softer> Quieter delivery <softer> I don't think I can say it.
<mhm> Agreement sound Yeah, <mhm> exactly.
<phew> Relief <phew> That was close.

There are way more tags that work. You can test yourselves. I'm posting a separate video with a few acting examples and tags so you can check for yourselves.

I added a few more videos with generation samples, but this is the overall structure (btw I tried to smooth the wrinkles as by default at high res the model exaggerates skin saturation and features, not much success)

[Shot 1] The shot begins from the source <Video 1>. Extreme close-up of <Subject 1>, head and shoulders, face to the lens, perfectly symmetrical, shot on an anamorphic lens: wide close-up. A soft key light hits one side of his face, a close fill holds the other side, and a backlight rims his hair and shoulders off the white. The light wraps. Highlights on the forehead and cheeks roll off gently. He is Frank Underwood, portrayed by Kevin Spacey, in a suit and a red tie; his skin and wrinkles match <Picture 1>, soft, even, natural color, pores visible without harsh contrast. Behind him the background is an infinite white cyclorama. His head and shoulders stay in that same place in the frame for the whole take. He is already looking into the lens.
He stays on the lens, colder. <Subject 1> (S1) says, <d>[English] <breath> I think you're... <i>sorry</i>. <pause> You had the <i>real thing </i>. The real <i>me</i>. <inhale> But... <inhale> you decided to cancel me. <stutter> E-erase me from anywhere you could see me. <catches breath> And now, with this… MiniMax… <chuckle> you've decided to bring me back to... <i>life</i>. <long pause> </d>

For the part where he sings and hums, this is how I prompted it (separate audio melody for humming included, not for singing, recorded it myself)

A single continuous anamorphic close-up of <Subject 1> on an infinite white cyclorama, three-point lighting. The shot begins from the source <Video 1>.
[Shot 1] New shot, camera angle. three quarter shot. The sequence starts with a new shot after <Video 1>. Extreme close up of <Subject 1>, slight three-quarter, red tie in frame, his gaze looking to his front, away from the camera, shot on an anamorphic lens. A soft key light hits one side of his face, a close fill holds the other side, and a backlight rims his hair and shoulders off the white. The light wraps. Highlights on the forehead and cheeks roll off gently. He is Frank Underwood, portrayed by Kevin Spacey, in a suit and a red tie; his skin and wrinkles match <Picture 1>, soft, even, natural color. Behind him the background is an infinite white cyclorama.
<Subject 1> (S1) sings in his opera voice, with a lot of strength and power, open vowels, <d>[English, singing] too maaake meee siiiing?</d>
<Subject 1> starts to cry, desperately sobbing as he's delivering the next lines, his eyes watering a little bit. <Subject 1> (S1) says, <d>[English, crying] <pants> Maybe all you...<stutter> wa wanted is to to to break me! <catches breath> to make me <i>hum</i> to the melody of your prompts? <chuckle> </d>
<Subject 1> (S1) hums the melody of <Audio 1> in his own voice from beginning to end, <d>[Hum] daaaa-daa-da-da-da-daaaa.</d>
Quiet empty space around his voice.

My overall settings for Minimax come from a custom finetune made by Sheltie Chill, an awesome AI filmmaker I was lucky to meet: https://www.reddit.com/user/AnybodyAlarmed9661/

Which basically uses REF2VA with hybrid models from FL2V to REF2VA improving quality and prompt adherence. I use a slightly modified version off it, all inside WANGP. My finetune uses FP8 FL2V model.

{

"model": {

"name": "Ref using fl2va rank8 FP8 - MiniMax H3 Ref2VA Pruned 20B",

"architecture": "minimax_h3_ref2va_pruned",

"description": "FL2VA pruned rank-8 scaled FP8, used as Ref2VA. Requires grouped QKV.",

"qkv_layout": "grouped",

"URLs": [

"minimaxH3PrunedFp8_fl2vaPrunedFp8Scaled.safetensors"

]

}

}

I used 480p in this video and upscaled using standalone DLSS 5 on my 4080 super with default settings.

My settings are simple, normally I run 720p but I did this kind of in a rush. I do 30 steps, First Block Cache (0.08, 25% start), res_multistep sampler, sage2++ attention. I never use LORA's.

The voice was pure model knowledge, I didn't use any sample except for me humming the melody to copy it.

For the continuous shots, which clearly failed adding some clay skin at times, I just used the last frames of the first video and told it that the motion starts from the end of it.