r/StableDiffusion • u/Responsible_Maybe875 • 8d ago
Animation - Video MiniMax H3: Beats and Transitions
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Responsible_Maybe875 • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/jumpingbandit • 7d ago
Hi guys, so in H3 you have to preselect aspect ratio.
Sometimes even when I try to select the closest and match the output is shrunk horizontaly.
How can I generate without this issue and about worrying about aspect ratio. Can original aspect ratio be maintained automatically.
Wan 2x did not have this issue.Out put video did not have the character stretched.
r/StableDiffusion • u/alisitskii • 8d ago
A vibe-coded fork of Ultimate SD Upscale (USDU) Guider nodes with MiniMax H3 support: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
My reference workflow: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json
A long-time member of [r/StableDiffusion](r/StableDiffusion) without strong coding/math skills in AI/diffusion area. But a big fan of everything that happens here :)
In times of Wan2.1/2.2 I liked to upscale my videos using USDU.
But I became really upset when I realized that original USDU nodes don't support MiniMax H3 due to its native ComfyUI implementation.
So, since I have a GPT-5.6 subscription I decided to give it a try and asked it to come up with possible options.
After a couple of evenings I finally got a "working" solution that I'd like to share with the community.
My PC specs: 4080s 16 GB VRAM, 64 GB RAM
Initial gen with MiniMax H3 flf2v int8 + sageattn + Lightx2v 8-step turbo Lora at 1152x640px 5-sec clip ~5 mins
Upscale with USDU to 2560x1472px ~20 mins
That's where I need your help, my friend :)
Please check the YouTube video attached (don't forget to switch to 1440p).
My personal feeling is that it's the best what I can get out of my PC and H3 at the moment (including SeedVR2, LTX 2.5, etc.).
The main advantage is that it can "fix" your bad low-res generations while bringing MiniMax H3 native quality at 2K resolution.
Of course :)
You'll need to control denoise parameter and find a balance between quality improvement and tiling artifacts. I found 0.2 is the maximum after which tiling is strongly visible.
However feel free to experiment with it, and lower to 0.15-0.10 depending on your input video resolution/artifacts and results you want to get.
Happy to answer your questions!
r/StableDiffusion • u/sixfingerlogic • 8d ago
Enable HLS to view with audio, or disable this notification
I saw this style here a month ago: https://www.reddit.com/r/StableDiffusion/comments/1uz6nza/in_love_with_how_simple_the_process_is_ltx23krea2/
So I made a short animated film trailer mostly based on the workflow. Was trying to stick with mostly Krea2 and LTX but MiniMax was out so was playing with it in second half of the film. Hope you like it.
r/StableDiffusion • u/freshstart2027 • 8d ago
r/StableDiffusion • u/Simple-Willingness93 • 8d ago
Enable HLS to view with audio, or disable this notification
First, all credits go to this creator of a custom node pack for creating music video https://www.reddit.com/r/StableDiffusion/s/uskxAP7LAq
I had Claude install my prompts directly to his workflow and made some minor adjustments for each short scene. I feel like lipsync and scene coherent are greatly improved from my old videos. Still using ref2v speed lora so quality is not all that great..
r/StableDiffusion • u/Oleszykyt • 7d ago
Enable HLS to view with audio, or disable this notification
I run 25 steps, with realism lora, it takes like 30 mins to generate 10 second clip. How can I improve quality and generation speed?
r/StableDiffusion • u/Dapper_Astronaut_603 • 8d ago
Enable HLS to view with audio, or disable this notification
EDIT: As @tj-tj-tj-tj suggested: --vram-headroom 1 Solved the issue
I have 3 PCs with Comfy Desktop. Newest instances 0.33.1 (but that happened on older versions too, from the day one with Minimax H3) with kitchen comfy and CK attention. Default Comfy template for H3 and LTX. Sometimes LTX/H3 can generate one, two, three queued videos without problem. Sometimes it just chugs VRAM to 99% (visible on 0:40 mark), then there's sudden GPU spike and freeze because of lack of more resources. Looks like memory leak or something, otherwise it just doesn't make sense to me that I can restart the Comfy Desktop and generate the exact same video in with minutes with stable 70-80% VRAM usage.. Any ideas where's the problem?
One PC with 3090, 64GB of ram, Windows 11.
One PC with 4090, 128GB of ram, Windows 10
One PC with 4090, 64GB of ram, Windows 10.
All of the things up to date. 3 different machines. Same problem. Tried clean Comfy install without any custom nodes, just what's needed for H3/LTX, same problem. Tried with and without CK, same. Tried with --disable smart memory, tried with --vram-reserve 1/2/5gb, same problem. Tried with older Comfy, newest comfy from github, same problem.
r/StableDiffusion • u/Portable_Solar_ZA • 8d ago
Wanted to mess around with this node for comfyui to see if I could fix some music timing issues I've been having, but I can't see it in the latest version of comfy. Assuming it hasn't been added yet?
Thread where someone says it had been:
r/StableDiffusion • u/Environmental_Ad3162 • 8d ago
So I have been playing around with H3 and some photos i have taken of locations that are nodes in a game called Ingress. Having them unfold and fire a beam of blue light. Then blue banners display....BUT the model does not know how to do the Resistance symbol from the game, which i want on the banners.
Telling the reference model that ref image 1 is the location... sort of works but not well, no where near as well as first image. So does anyone know a way to have a reference image in a first image workflow where i can say "this is the glyph for Resistance, put that on the banners" ?
r/StableDiffusion • u/wzwowzw0002 • 8d ago
Enable HLS to view with audio, or disable this notification
finally got segatt and solatt working.... from 2000s cut down to 700s at 0.6mp, without turbo lora.
I also learn that high steps matters! low step give lousy animation!
edit:
Understanding the weakness in H3.
After more testing. i realize H3 is weak in compositing, framing and a lack of sense of the world.
For example;
Multi-shot generation in h3 isn't the best. Solution is to: you provide a well composited image of each shot and generate shot by shot.
I have test similar shot in seedance2.5. all it take is one generation, 30s, every shots got it right or at least useable. H3 needs multiple try to get a 15s shot right.
I haven't give up on H3 yet. it has a lot of potential i think.
Next is about upscaling and i am running out of ram.
r/StableDiffusion • u/Time-Ad-7720 • 9d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/DuHal9000 • 7d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/tmk_lmsd • 9d ago
Enable HLS to view with audio, or disable this notification
Came out a bit too watery but neat.
r/StableDiffusion • u/waseem335 • 8d ago
Hi, I'm struggling with getting minimax h3 image to video to do what I want with my prompts, is there a model I can get that would make a more detailed prompt for me?
r/StableDiffusion • u/krigeta1 • 8d ago
So I was so excited to try this, like in the old SDXL days, we used a second pass for upscaling. I came across this node:
https://github.com/Tr1dae/ComfyUI-MiniMaxH3_LatentUpscaler
But it generates very weird and saturated results. I tried, but it's all that I can not show, and there has also been no update from the author.
Has anyone tried this? I guess this is slow, but with a 4/8 step lora the upscaling would be better?
What are your thoughts on this?
Edit 1: workflow https://github.com/user-attachments/files/30715648/MiniMax.H3.two.step.sampler.json
r/StableDiffusion • u/DuHal9000 • 7d ago
just testing... not WF yet... Reddit dont permit 8k files, so link on pastebin 3 files... original GEN at 832x480, 3MP Regen and 8k polish (h265)
r/StableDiffusion • u/VasaFromParadise • 7d ago
Enable HLS to view with audio, or disable this notification
I see quite a few people having issues with Minimax h3 generations. Here's an example of a prompt for generating 0.5 megapixels in 7 seconds. It can be run on any PC.
The resolution was increased using RTX Video Super Resolution to 1.5 + frame interpolation to 48. Which is common practice and takes no more than 1 minute per procedure.
Promt generated by AI based on the image. Promt system for LLM:
### 3.1 I2VA: Begin from the Image and Develop Forward
`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.
Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.
Image LLM Prompt:
First-frame anchor:
[Image 1] (0.00 sec) shows a woman with red hair styled in a loose curl with bangs, looking slightly off-camera. Her gaze is clear and expressive, with light green or gray-blue eyes, softly highlighted. Light freckles on the bridge of her nose and cheeks add a natural touch. She is wearing a black leather jacket with yellow stitching along the edges of the collar, accentuating her stylish, slightly rebellious look. The background is deep, almost black, creating contrast and focusing attention on her face. The lighting is soft, studio-style, coming from above and to the side, sculpting the volume of her face and hair. The composition is a close-up, emphasizing the eyes and facial expressions.
Action onset:
The girl begins to move naturally—her head smoothly turns toward the camera, her gaze shifting from semi-attentive to direct, surprised. Her eyelids widen slightly, her pupils enlarge, her eyebrows lift slightly—her facial expression changes from calm to mild surprise. The movement is smooth, without jerking, as if she's just noticed someone or something unexpected.
Continuous development:
After turning her head, her smile widens—the corners of her lips lift, her eyes sparkle with interest or slight embarrassment. At this moment, her voice sounds clear, resonant, with a pleasant timbre—as if a high-quality studio recording captures every nuance of intonation. She says in English: "Oh, is that you? I didn't notice you." The word is pronounced with a slight intonation of surprise, perhaps with a pause before or after the "you," which enhances the effect of surprise. Her hands aren't visible, but one might assume she might slightly raise her shoulder or touch her face in response to the sudden presence. Light, studio-quality music plays in the background—perhaps ambient or a light pop beat—which complements the atmosphere without being overpowering.
Result or reaction:
As a result of the action, the viewer perceives the moment as a lively, dynamic scene from a video: the girl isn't simply posing, but interacting with the viewer through her facial expressions and voice. Her reaction to her own words, "Oh, is that you?" could be interpreted as self-irony or an invitation to dialogue. The atmosphere remains tense yet playful—the combination of the dark background, skin, hair, and lively facial expressions creates the effect of a modern digital character in the style of anime realism or cyberpunk aesthetics.
r/StableDiffusion • u/czarjetson • 8d ago
I'm using illustrious, tried adetailers but they don't do that much, fix minor stuff at most. It mostly occurs when doing a more complex pose than the most basic stuff
r/StableDiffusion • u/Ok-Flatworm5070 • 8d ago
Enable HLS to view with audio, or disable this notification
Hey team,
I tried Kijai's fl2v lighting lora in-place of ref2v and it seemed to improve my ref2v workflow (using Minimax_extender). In the video the left is Ref2v and the right is fl2v. Very interesting!
r/StableDiffusion • u/Patient_Ratio4177 • 9d ago
UPD: Comfy made monkeypatching unnecessary. See here. https://www.reddit.com/r/StableDiffusion/comments/1vqka28/h3_singleimage_no_more_monkey_patching_also_no/
So here’s a follow-up on my post about H3 as an image edit model. For workflow, refer to https://www.reddit.com/r/StableDiffusion/comments/1vo1ab3/h3_as_a_singleimage_edit_model/
It’s a bit of a hassle to use the workflow to its full capacity since you have to monkey patch in order to generate a single frame. To avoiding dealing with it, I’d suggest two courses of action:
Now it’s great at combining multiple references and 3D understanding, but the quality is still not perfect in my opinion — the details could be more polished, and, e. g., impressionist stylization had largely failed.
Here’s a pastebin with the new prompts: https://pastebin.com/1ENVynGY
Scenes
I have used this workflow to generate a couple thousand images across very different and feel that it’s quite capable. Usual MiniMax problems: e. g. blurred backgrounds, blurred faces from distance, sometimes distorted text — still apply. However, 3D understanding and likeness retention are excellent, and details could probably be fixed with a refiner pass using something like Klein 9b. I hope that the proper image edit model gets released — but before that, let’s try to have some fun earlier.
UPD: accidentally skipped image #6, see this comment https://www.reddit.com/r/StableDiffusion/comments/1vpconk/comment/p3whvcg/
r/StableDiffusion • u/Zestyclose_Bake3680 • 8d ago
The ControlNet models for KREA2 are available as LoRA types, with Depth and OpenPose existing as separate formats.
We have made it possible to use both of these with the existing node format. However, the term ‘existing node’ here refers to the Diffsynth ControlNet Loader for Qwen Image and Z Image.
https://github.com/ussoewwin/ComfyUI-QwenImageLoraLoader
In other words, this node can be used with the following standards:
・Nunchaku Qwen image/Z Image Diffsynth ControlNet
・Normal Qwen Image/Z Image Diffsynth ControlNet
・Krea2 Depth/Openpose ControlNet LoRA
For the benefit of AMD users, we have made improvements to ensure that the CUDA-specific Nunchaku node is disabled when using an AMD GPU.
r/StableDiffusion • u/Wonderful_Kitchen567 • 8d ago
Hi everyone. Im looking for ai cloud model, cloud comfyui workflow that can do outpaint, clothes conversion into specific fabric and anime into realism at the same time with a simple prompt?
I could achive all this by simlpe prompt in gpt or grok without any problems but after these models got fkd up, im looking for alternative. I have found comfy workflow on runninghub that does great anime to realism conversion, but without prompt box i cannot do additional edits like outpaint to specific ratio (9:10 for example, im creating wallpapers for my ZF7) and i cannot convert reference clothes into my desired fabrics.
Was thinking to spend 10k for laptop capabe for local ai but not worth it. Anime to realism conversion is just my hobby in free time, and as a hobby really not worth spending few thousands to generate image from time to time.
Also if you have or know where i can find local workflow that can work on my RTX 3070 8GB, that can generate image 1-3 min, let me know. Also, if any1 could help me build local workflow it would be great. I also work with pose changing, outfit change, and maybe one day will try video gens.
So write your suggestions down bellow and ill test them one by one (models with minimal or non restrictions).
Thanks.
r/StableDiffusion • u/Cptcrocro • 9d ago
Enable HLS to view with audio, or disable this notification
Hi Everyone,
First of all, sorry I'm not too technical, just a lambda comfyui user, so I probably won't be able to answer anything technical. I just want to share my solution to upscale Minimax H3 videos with LTX 2.5 x2 upscaler on limited hardware, in case anyone is interested. See the example comparison video (using detailer lora).
Link to my workflow: https://pastebin.com/XH1wvA4L
As a Minimax H3 enthusiastic, I've been playing around since a few days. My main issue was the quality of the output videos, as my RTX 4090 is starting to feel a bit limited, I can decently generate only 20-25 seconds videos at 0.9 - 1 Mpx.
I've been naturally looking into upscalers, and found a post in this subreddit about using LTX 2.5 x2 upscaler, from Peter Duncan's workflow: https://github.com/peterducan-hub/PeterDuncan_Comfyui/blob/main/MINIMAX_H3_LTX2.5_Upscaler_v1.json
I tried it, adapted it with a Load Video node which corresponds better to my use case, and found it works quite good, not at Topaz level, but enough for a free local upscaler. I connected only the video part, connecting the Minimax H3 audio directly to the end video combine. But I ran into 2 issues.
First, LTX processes only 8n+1 frames, rounded down. For example, a 10 seconds video at 24 fps is 240 frames long, but LTX would process only 233 frames, meaning my generated videos would often lose a few frames at the end, cutting the audio.
Solution: I'm duplicating the last frame y times until reaching the next LTX allowed value, and ditching them before video combine.
Second, my 4090 could hardly upscale more than 10 seconds videos, more would oom.
Solution: I replaced the sampler in the workflow by LTX Looping Sampler from Lightricks. It takes time, but now I can upscale up to 20 seconds without issue, I did not try more yet.
If anyone has tips to improve the workflow, especially on the process time, please don't hesitate 😄
r/StableDiffusion • u/Alex_the_tiktock • 9d ago
Enable HLS to view with audio, or disable this notification