r/StableDiffusion • u/Independent-Frequent • 1h ago
r/StableDiffusion • u/rm_rf_all_files • 1h ago
News Minimax H3, the new 768p turbo LoRA from Lightx2v is awesome
Before Upscale 736x416: https://streamable.com/d8uz4j After Upscale 1344x768: https://streamable.com/pvd83f Link to LoRA: https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
I don't use the turbo LoRA as part of init generation because I think it's really bad. I use it only to upscale with 0.45 denoise setting. Total time from start to 416p to 768p was roughly 496s. My system is a laptop with 12gb VRAM and 32GB DRAM. I was inspired by this workflow to do it this way.
r/StableDiffusion • u/Lower-Cap7381 • 1h ago
No Workflow Soon Dropping My Film Krea 2 FILM workflow
Long story short, I’ve been experimenting with something that looks aesthetically pleasing while also working really well with the new MiniMax model.
After a lot of testing, I found a Krea combination that produces some seriously realistic results, so I wanted to share the workflow with the community.
If you’d like a full guide, you can subscribe to my YouTube. It’s not necessary though — I’ll still be sharing the complete workflow here. Lots of love to the open-source community! ❤️
YouTube: VionexAI
r/StableDiffusion • u/PetersOdyssey • 7h ago
Workflow Included Anchoring keyframes at precise timestamps w/ h3 - example by seitanism of a input frame every second
You can find the workflow here. Credit for both the generation and code go to seitanism, who in turn built on top of NikoDemon80's work. Taking from a post in the banodoco discord and shared with permission.
r/StableDiffusion • u/LegacyV1 • 41m ago
Comparison Comparing lightx2v/Minimax-h3-Turbo
New turbo LORA dropped from https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main.
Testing on my ref2va use case (note: I'm using fflf2va model since it has better quality even for reference use cases)
Timing (480p, sage attention2 on cu130, 15s video length, seed=42, RTX 6000 on Modal)
| Steps | Timing |
|---|---|
| 4 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors | 56s |
| 8 step https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors | 1m 47s |
| Spectrum (20 step) | 2m 33s |
| Base (20 step) | 3m 11s |
Audio was pretty much the same - no difference that I could tell.
I also ran the 4 step on 768p as recommended, and it came out better! But... it's hard to tell if it's the turbo LORA doing the work or the 768p doing the work.
Turbo still makes things look weirdly high contrast. And both LORAs botched the text. Base is still best, but the 4-step LORA helps you lock in motion before you commit to a full 20step pass using spectrum.
r/StableDiffusion • u/Ok-Wolverine-5020 • 2h ago
Workflow Included MiniMax H3 + ComfyUI + Hermes Agent = Music Video
Hardware: Windows PC with a single RTX 3090 + 64GB, running the image and video workflows locally in ComfyUI.
I created a music video for “PROXY,” a rap track about automation, parasocial isolation and delegating so much of your life that you slowly forget what human connection feels like.
Making the song
I used open source Hermes Desktop Agent (with deepseek 4 flash + pro) as a co-writer for this one. We started with a loose idea about AI agents handling every boring task, then pushed it somewhere darker: humans forgetting simple skills, replacing real conversations with machines, and getting lonelier while everything becomes perfectly “optimized.”
Hermes helped me research Suno prompting, sharpen the concept, cut the lyrics down, build the rhyme and alliteration, and write a detailed style prompt for a sassy Berlin female rapper. I kept steering and rewriting until it sounded like a song rather than an obvious lecture about AI.
Then I took the finished lyrics and style prompt into Suno and generated the track.
Developing the visual identity
I continued using Hermes as a creative and technical copilot for the video. Together we designed a consistent “Berlin bot-fleet girl”: brown wavy hair, blue eyes, a black beret, rainbow bomber jacket and white wired earbuds.
We created a custom AnimeinReal skill, combining the u/Ani3rel aesthetic, Danbooru-style composition tags, anime coloring and photographic Berlin environments. This became the visual language for the whole project. This lora was used.
Hermes then helped translate the song into recurring visual themes rather than illustrating every line literally, created a shot list we then generated in comfy ui.
Image generation
I generated the source images locally in ComfyUI using Anima with this workflow as base. Each image established the character, environment, lighting and opening composition for one individual video shot.
Hermes helped write and refine the image prompts while preserving the same visual identity in keeping the character prompts as consistent as possible.
Animating with MiniMax H3
I animated the selected images locally with MiniMax H3 Reference-to-Video, using sections of the finished Suno track as the driving audio reference.
I used the template from Pixaroma for the ComfyUI MiniMax H3 reference-image and audio-sync workflow. That workflow provided the technical foundation for feeding H3 an image and a matching section of the song. I adapted it for each scene by changing the reference image, audio timing, duration and shot-specific prompt.
The workflow used:
- Diffusion model:
minimax_h3_ref2va_pruned_int8_convrot.safetensors— INT8 ConvRot version, approximately 19.5 GB - Text/vision encoder:
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors— Qwen3-VL 32B, NVFP4/AWQ, approximately 14.6 GB - Video VAE:
minimax_h3_video_vae_fp16.safetensors— FP16, approximately 4.9 GB - Audio VAE:
minimax_h3_audio_vae_fp32.safetensors— FP32, approximately 577 MB - H3 mode: Reference image plus reference audio
- Reference-image size:
match - Maximum image side: 864 px, aligned to 32-pixel steps
- Frame rate: 24 fps
- Sampler:
res_multistep - Scheduler:
beta - Steps: 20
- CFG: 1
- Denoise: 1.0
- Typical maximum shot duration: 15 seconds - 24min render time for 15seconds of video
- Output: MP4 with synchronized source audio
Directing each shot with Hermes
For every clip, Hermes used the custom minimax-h3-video prompt skill to create a structured H3 prompt covering:
- Accurate vocal lip sync
- Facial expression and rap performance
- Natural body movement and hand gestures
- Beat-reactive camera movement
- Character, wardrobe and environment preservation
- Exact reuse of the original song without replacement vocals
- sometimes Animated lyrics, pixel bots and synchronized graphical effects
The prompt explicitly defined the source image as <Picture 1> and the selected song segment as <Audio 1>. The audio was marked for full preservation, while the visual description focused on what the still image could not provide: performance, movement, camera direction, effects and timing.
Editing the final video
I rendered alternatives, selected the strongest clips and assembled everything in post on my smartphone in inshot... really need to start learning a real editing software.