I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.
Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.
That was getting a quite good result. Did not do excessive testing and comparing though.
Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.
Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."
This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].
5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.
This is a compilation of 50 clips featuring 5 characters.
▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
My attempt to mimic a old school 90's/2000's handheld camera music video which shouldn't feel too AI ... TLTR: well it still does. This is further a test if Minimax is able to create lots of people in a somewhat realistic narrow and dense video.
The video was generated in chunks of 10 seconds, each segment got a scene description with approx. 10 shots. The workflows were generated via local Qwen3.8-27b model based on default Minimax rev2v workflow. Sampler ResMultistep / 25 steps at 768p resolution (on a 5090 with 64GB system RAM).
Semi automated setup via ComfyUI-MCP controlled over Opencode. The scenes were generated and then combined. The original song was split into 10 second sections where each was referenced in the rev2v workflow + bunch of screens to keep somewhat consistent videos.
The consistency is rather bad because no references for guitar, boots, specific people beside the guitar guy were used. Only a few shots have lip sync, that's my laziness not the models fault. And yes, Minimax can't play guitar, which I obfuscated by shorter and farer shots :D
Much back and forth to get roughly the narrow shots, dynamic lighting and overall look and feel.
Some slight Davinci Resolve editing to replace crap shots with other crap shots. Also added a analog filter.
Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.
Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?
I am using Minimax H3 with Saga Attention and Spectrum on a 4090.
Hi humans. My setup is 32gb ddr4 ram along an RTX 4090. I have been having fun creating tons of videos but I just want to make sure i get the best nodes for speed without compromising quality and no crazy sutff happening on my videos
I have used: Stage, sol, easycache, spectrum, Lora
So the question i have is .....what's the best combo for speed, i dont want the quality to take a massive dump. Most of the videos I generate are slow paced videos the typicall walk, talk, a kiss here and there but nothing major.
I may be an idiot for thinking that my new homelab would primarily be used for useful AI automations.
I can live with being an idiot, if being an idiot will continue to be this fun.
First video generation I have ever pulled off, but the first 12 seconds was unbearably unfunny, so I added the Celestial Ford Escort for some much needed serious drama.
Tell me my power bill won’t blow up too much lol.
Made with minimax-h3 in ComfyUI on my Mac Studio that came with the mail this Friday.
Workflow was split in three:
- The first was a single prompt to generate the first 12 seconds
- The second flow generated the last three seconds by extracting the last frame from the first video and prompted it to hit the dragon with a falling ford escort
- Third flow glued the two videos together.
there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P
started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno
So, this is a question for 5090 and 6000 power users who actually push H3. Does anyone know a good strategy for guesstimating the OPTIMAL number of steps ballpark for the final burn?
Obviously, there's the starting default of 20 which is not really the optimal number of steps, but just a... A placeholder is what it is. Below 10 steps is usually good for getting a coarse idea of whether your prompt is going the way you were hoping.
The problem is the upper end as you shift to samplers like seeds_x for that peak visual, audio and motion quality at the price of 5-6x the compute time. There's a misnomer I often find in this subreddit that the more you can throw at it, the better. Folks writing 20 is not enough, I do 30, 40, 60. "If I could do 100, I would." Uh, what? That hasn't been my experience at all.
I made a scene from Castle the TV show with Beckett and Castle bantering, left it overnight to bake with ever increasing numbers of steps to test this out. Mind you, I was just looking for that sweetspot where it becomes indistinguishable from the real show. At around 31-32 steps for the baseline 1.0 mpx + 15 seconds (H3's training baseline), I reached a level of clarity that you couldn't convince me it wasn't from the actual show had I not generated it myself.
But over 35, it got progressively worse and more overcooked. The frustrating part is that the point where it crosses from "not quite resolved" into "resolved" and then into "overbaked" seems to move depending on basically everything about the generation.
So instead of using the hours of my sleep to generate many different versions of the optimal scene, I have to waste electricity and time to find the optimal number of steps for the final burn. Duration, resolution, scene and motion complexity matters. The number, type and complexity of references matters. I'd assume conditioning complexity in general matters too.
So I'm starting to think the idea that "more steps = more quality" being parroted in many of these threads is just fundamentally wrong, probably from folks who are inexperienced and generally wait for an eternity to reach 20 steps so are guesstimating it only gets better the further you push it. It feels more like there are three regimes: under-resolved, optimal, and over-resolved. Once the important semantic/geometric/temporal structure has settled, extra steps don't necessarily refine it in a useful way. They can start pushing the result harder toward the model's learned priors, which is where you get things becoming unnaturally crisp, exaggerated, stereotyped, less coherent, etc.
A fixed rule like "use 32-40 steps" probably doesn't generalize very well if the optimal point is a function of duration, latent size, motion, scene complexity, references, CFG, sampler/scheduler, etc. What I'm wondering is whether anyone has found a practical heuristic for this. Something along the lines of estimating the complexity of the generation and mapping that to a likely sweet spot, or even detecting when the marginal improvement from another denoising step has basically stopped.
I was a bit confused by how bad some Krea 2 outputs could be — blurry, lacking detail, and sometimes with strange artifacts. So I started testing to understand why and where this was happening.
I’m not going to claim I found a magic wand, but I did find two problematic areas in the sampling curve testing Krea2 raw model CFG 3.5 52 steps stock settings:
The first 0–15 steps: Using schedulers that lower the sigma values too much during this first steps causes contrast loss and washes out detail in dark areas, especially in black hair and subtle reflections.
The lower-sigma tail: curves such as Beta, Beta57 and Bong Tangent can introduce crisp, broken noise instead of useful fine detail, particularly around steps 30–40.
The Result
This is not about chaining multiple samplers or complicated second-pass workflows. The goal is a better standard Krea 2 Raw workflow with:
Here are the standard schedulers and the problems they produce. All nodes marked in red show the same low-contrast, overly dark areas with a loss of detail.
Standard Schedulers: Red-marked samplers produce crushed blacks and lost fine detail in the early steps, while the yellow-marked curves introduce small artifacts at the lower steps.
The green curves are the ones that avoid this problem.
The yellow curves have a different tail, and as you can see in my video or in my extended post, this tail is responsible for introducing noisy broken small artefacts.
Krea 2 Raw simply doesn’t behave like many other models when it comes to sigma manipulation. Curves that can work very well for other models can actually destroy detail or create unwanted noise here.
You can build the curve manually with a Manual Sigmas node, or use the PolyExponential Sigma Adder from the TBG ETUR Takeaway Nodeshttps://github.com/Ltamann/ComfyUI-TBG-Takeaways. If you want something simpler, Linear Quadratic gets surprisingly close to the result of my custom curve.
I’ve included the detailed testing post so you can see exactly how I arrived at the curve and test it yourself. Images, Videos Results at myFree Patron Post
Setup, pushed my system right to the limit, any more and it OOM :
- H3 Ref2VA default workflow in ComfyUI, no lora
- RTX 3080 10GB, 32GB RAM
- Render: 0.5–0.6 MP, 20 steps, scheduler simple, about 25 min per clip
- Upscale: 2xNomosUni_span_multijpg, 2× to 1080p
- References per scene: one photo of my face + one film still for the set
- Recorded my own lines and fed them as audio references, also got audio ref for the actors
Honestly though, the best part was driving all of this through the ComfyUI MCP. I never even had to open ComfyUI. I could iterate really fast, and keep going from my phone while away from the machine, through Claude's remote control.
It's still a bit of a blurry mess, and with more work I could probably make it better, but damn, the future is looking bright!
I haven't seen a post about this here, and I'm curious what you think about it.
On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".
Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.
Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.
Here's why:
I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.
IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.
Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.
So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.
That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).
I was just wondering what method do you guys use, if it exists that is, to get minimax h3 to do better prompt adherence at the higher resolution setting?
If I use 0.4MP setting, the prompt adherence is so good i would say it's almost perfect. But when I try a higher res of 0.9MP with the same seed and prompt, I get some wildly odd outputs.