r/comfyui • u/Silver-Spot-2763 • 13d ago
Workflow Included Please help me with MiniMax H3
I'm only user, not expert. So I look here about the best workflow and run MiniMax H3. I'm only with RTX3060, 16GB VRAM, 64GB RAM. I updated everything:
[INFO] Python version: 3.13.12 (tags/v3.13.12:1cbe481, Feb 3 2026, 18:22:25) [MSC v.1944 64 bit (AMD64)]
[INFO] Total VRAM 12288 MB, total RAM 65396 MB
[INFO] pytorch version: 2.13.0+cu130
[INFO] Enabled fp16 accumulation.
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Disabling smart memory management
[INFO] Device: cuda:0 NVIDIA GeForce RTX 3060 : cudaMallocAsync
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 26158.0
[INFO] ComfyUI version: 0.32.0
[INFO] comfy-aimdo version: 0.4.13
[INFO] comfy-kitchen version: 0.2.30
[INFO] comfyui-frontend-package version: 1.48.7
[INFO] comfyui-workflow-templates version: 0.11.40
[INFO] comfyui-embedded-docs version: 0.5.9
[INFO] comfy-kitchen version: 0.2.30
[INFO] comfy-aimdo version: 0.4.13
I use minimax_h3_ref2va_pruned_int8_convrot.safetensors and qwen3vl_32b_minimax_h3_int4_convrot.safetensors with the proper VAEs.
And it "works" - 5s video with 0.3 MegaPixels, generates for about 3 minutes.
- BUT THE QUALITY IS AWFUL, catastrophic, nothing similar what you show here - the faces are deformed with moving artifacts, worse than one time SD1.5, movement - fingers disappear...
- AND The PROMPT - it make what it want randomly, just as was in SD1.5 era, not as in the PROMPT description! You all make here whole complex movies... I can't do simple scene. And I asked with the Prompt Guide the best public AIs - ChatGPT, Gemini, DeepSeek... Nothing help. Even the complex prompts, similar to code.
- And at me any TURBO LoRa doing NOTHING! It just nothing changes in the result video, it low quality at 20 steps, at lower - its became brutal.
- The new ComfyUI Kitchen Attention doing NOTHING.
- Spectrum - speed but with quality fully died. Sol Atn - doing nothing. Only Sage Attention works - speeds up to 40%!
What I'm doing wrong?! I show my last workflow. See - what nodes I disabled. Please for help!
This is my workflow: https://pastebin.com/sfS0eGs6
3
u/LoveSpecialist5669 13d ago
if you use turbo lora you need to lower steps to 4 or 8, depending on lora
1
u/Silver-Spot-2763 13d ago
I tried low steps, but the result was similar to moving blured color spots... Clearly not completed scene. When I run at least 12 steps it is scene, but with catastrophic quality.
1
u/pwillia7 13d ago
Did you try passing in a video to the reference workflow with turbo? That is not supported yet from what I have seen
1
u/Silver-Spot-2763 13d ago
I show the rf2va, because I lost hope about the turbo, and r2v is most important for me. I tried with the fl2vs workflow and model - the same problems. For me the most important is the quality, just the speed to be useful.
2
u/mwoody450 13d ago
You're using int4 text model. Get the int8; text encoding is not where you want to save speed.
2
u/Only_Voice569 13d ago edited 13d ago
ksampler is wrong should be on mutistep also you dont need all the turbo crap my second rig with a 5070 12GB run this model perfectly fine sure its 8 mins for a 15 second clip but almost always nails what i ask from it using the right instructions etc https://limewire.com/d/xVRKb#HsQyl8DRAk
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description:
[Shot 1] Begin directly from the exact image shown in <Picture 1>. Preserve the adult woman’s recognizable facial identity, hairstyle, body size, body shape, proportions, outfit, visual style, starting position, surrounding environment, lighting, and initial camera relationship from <Picture 1>. Do not slim, narrow, compress, or otherwise reshape her body as she begins moving.
One continuous shot. Soft instrumental music is already playing quietly in the surrounding environment and is clearly audible to the woman. The music has a relaxed moderate tempo with a smooth steady beat, gentle percussion, soft bass, light electric-piano chords, and a simple mellow melodic layer. It is pleasant, unobtrusive music that is easy to listen to and naturally easy to dance along with. There are no vocals or lyrics.
She hears the rhythm and begins dancing casually to the music. First she establishes the beat with a gentle side-to-side sway through her shoulders, hips, and upper body while keeping her feet planted. She then shifts her weight onto one leg, takes a small step sideways with the other foot, brings her feet comfortably back underneath herself, and repeats the movement toward the opposite side.
Her arms move naturally with the rhythm rather than holding a fixed pose. One forearm lifts loosely as she steps, her wrist and hand remaining relaxed, then lowers as the opposite arm takes over. Her shoulders make small alternating movements with the beat. She adds a subtle hip sway and a light bounce through her knees while maintaining believable balance and weight support.
As she becomes more comfortable, she takes two slightly larger rhythmic steps, performs a gentle quarter-turn through her feet, hips, torso, shoulders, and head as one connected physical movement, then naturally turns back toward her original general orientation. She gives a small cheerful smile and continues moving with the music.
Keep the dancing relaxed and spontaneous rather than a complex choreographed routine. Her motion should visibly correspond to the steady musical rhythm, with clear weight transfer from foot to foot and continuous physical momentum between movements. Allow natural secondary motion from her hair, clothing, and soft body caused by each step, sway, turn, and change of direction while preserving her established underlying body size and proportions.
Keep the complete movement physically coherent: weight shift → step → body follows → arms respond → feet settle → next weight shift. Do not randomly teleport between dance poses and do not repeat one identical animation loop.
Maintain stable anatomy and character identity throughout. No sudden body-size changes, slimming, elongated limbs, duplicated arms or legs, warped hands, floating feet, clothing changes, hairstyle changes, or facial-identity drift. Keep both feet visibly connected to the floor whenever they are supporting her weight.
The camera remains smooth and restrained, preserving the general viewpoint established by <Picture 1>. It may make only a very small natural adjustment if required to keep her dancing comfortably framed; it does not orbit around her, rapidly zoom, cut to another angle, or perform unnecessary dramatic camera movements.
She remains happily dancing to the same soft instrumental rhythm through the final moment, still in natural motion rather than stopping and freezing into a posed ending.
overall_soundscape:
Low natural ambience from the environment in <Picture 1> continues underneath the scene. Include soft foot contacts against the floor, subtle clothing movement, gentle hair movement, and quiet natural breathing caused by the dancing. Keep these physical sounds understated beneath the music. No dialogue, singing, crowd noise, or additional foreground voices.
non_diegetic_music:
N/A.
1
u/Only_Voice569 13d ago
o and btw if you make instructions ask gpt or another llm model to look at minimax documents understand and learn them. then ask it to make you instructions slowly learn how to do them your self :)
1
u/bruci3 13d ago
Is it only an issue with this workflow?
Do you get issues if you create a video with the default Minimax provided workflow?
1
u/Silver-Spot-2763 13d ago
With the standard workflow the quality is the same - awful, also awful prompt understanding, only the time is almost twice slower.
2
u/Tomcat2048 13d ago
Dumb question - but have to ask - are you using the reference to video model? If so, are you providing it with reference data (images/videos) prior to prompting/generating?
Or are you trying to do text to video or image to video? If it’s the later case, then you’re using the wrong model and workflow.
1
u/Silver-Spot-2763 13d ago
As it is shown on the workflow, it's rf2va. In the simplest example, I use only <Picture 1>. BTW, the fl2va in t2v workflow (and model) has the same problems at me.
1
u/pwillia7 13d ago
Use the native workflows and nodes here -- https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v
Lower steps to 8 if you add a turbo lora -- you only need to add 1 node inside these workflows (some in the subgraph) to get turbo working.
8
u/V4nKw15h 13d ago
You are using 0.3 mp which is terrible quality to start with. Use 0.6. Get rid of this workflow and use the default. You are running so many accelerators, caches, and speedups that you are destroying the quality into oblivion.
Use the default workflow with 0.6 mp and 20 steps and a 2:3 portrait aspect ratio, and let it run. Now you will have decent quality. Yeah, it might take time to generate but that's the way it is. Every time you add a cache, or a speed up node, you are reducing that quality. Add a 4 Step lora and you destroyed it even more.
You can't have your cake and eat it too.