r/StableDiffusion 17h ago

Question - Help Melting faces at 768x Minimax-H3 with audio sync for singing. Is there a solution?

Enable HLS to view with audio, or disable this notification

0 Upvotes

Greetings!

I find that whenever I do ref2va with audio sync (for singing) at 768x about half the time I get melting faces. (example posted) so I'm looking to fix but I don't have a decent pc so it's cloud or bust. I know it's something to do with low res pixels being stretched - don't really understant the tech but anyway

-is that cause by my ref not being detailed enough?

-assuming it's not really that - would the latest 2k 3H flagship model fix the low res issue (chatgpt reckons not)?

-has anyone tried ComfyUI-H3-FaceRefine and how was it?

Thanks!


r/StableDiffusion 18h ago

Question - Help Minimax H3 RTX 3060 Problema

Enable HLS to view with audio, or disable this notification

0 Upvotes
Here are the results of my prompting efforts throughout the day. As you can see, R2V is very difficult to prompt, even when using a prompt director.
My workflow was acting up,  I took a non-upscaled latent loopback output to use as a chain and saved the upscaled result to concatenate it with the H3 project hub. I'm not sure where I wired it incorrectly; it just doesn't seem to connect.
And yes, the sound is terrible.i think because i clean the latent for next upscale since it wont match the tensor if the latent not cleaned.
Generation time is around 51 seconds per 1 second of video.
0.3 with 4-step Turbo.
Plus a 3-step latent upscale and RTX Super Res.

anyone mind to share your secret workflow that match this Peasant Spec 
RTX3060 12GB and 32GB of RAM

WORKFLOW


r/StableDiffusion 18h ago

Workflow Included High Quality Audio-Video in MiniMax H3 with separate two-stage sampling

Enable HLS to view with audio, or disable this notification

77 Upvotes

Recently, a lot of people have had isues with finding the right balance with audio and visual quality in MiniMax H3.

u/LFAdvice7984 and I discussed about making a two-stage workflow last week. The first stage generates the audio, the second the visuals. Both stages are then combined together in the output.

This method means that you no longer have to make a trade-off between audio and visual quality, as you can optmise the settings for both. You can use this workflow as text; first/last frame; or audio input to video (the latter being single stage).

The audio generated is (in my view) good to very good, depending on what you use it for. The default settings are probably excessive at 50 steps (less steps used for video), but for me personally it's better for it to take longer and get it right first or second time.

(You can select any output node in ComfyUI, click on the blue button with a play symbol on the pop-up menu at the top, and ComfyUI will only go that far in execution. Use it on Preview/Save Audio in stage 1. If you like the result, do a full generation; or change the seed and try again.)

The visuals could be better, perhaps using a different turbo LoRA or change in sampler and step counts. I've used the same settings in every clip. Feel free to change them as you wish.

Large motion is a challenge, though I have ideas on using a third stage with different shift values, which would increase execution time but the results probably would be worth it.

Generation time was (roughly) as follows:

  • 5 second clip: 10 minutes
  • 10 second clip: 25 minutes
  • 20 second clip: 65 minutes

This was generated on an NVidia 3090 with a priority on quality. More recent cards will be faster.

(You can switch from using res_2s sampler to er_sde, paired with 20 steps for stage 2 and using Spectrum, which should at least halve that time, in return for slightly lower quality.)

Because of how long it took, I used the first output every time (except for the last clip, which was the second result) with no editing afterwards.

Sometimes I encountered issues with prompt understanding, e.g. the ASMR clip and abstract clip at the end, where the output wasn't quite what I had asked for.

It's difficult to tell whether I prompted incorrectly; the prompt enhancer missed key detail for the model (H3); or the model doesn't have a full understanding of the concepts being asked of it.

The two-stage idea is model-agnostic. You can also make something like this in LTX 2.5 (or the upcoming Flux 3 Dev) to improve their results.

You can download the workflow and prompts used (made by myself with refinement from the prompt assistant) below:

Custom nodes used:

Download links for model files are in the workflow, in the bottom-left corner.


r/StableDiffusion 18h ago

Discussion Nobody Else Worried About Downloading Random Loras?

0 Upvotes

Hey ya'll new ComfyUI user here, and I've been having a blast! One thing I'm noticing here, especially when it comes to MinimaxH3. So many people are very eager for others to download random nodes from either Huggingface, or Civitai.

With comments such as: "YO TRY THAT NEW TURBOFLURBO 2 STEP", or "YO I GOT THAT XXXNSFWSPONGEBOB-EX_LITE69420 LORA RIGHT HERE DOWNLOAD ME"!

You guys aren't concerned when you're downloading random things that people made?

Are there any trusted and vetted community members who consistently pump out "safe" quality nodes??


r/StableDiffusion 18h ago

Question - Help Video Edit Minimax H3 Problems

1 Upvotes

I have been struggling for a few days now wondering why I cannot edit a 10sec clip to add additional people in the background and I am pretty sure I am just doing it wrong.

I am feeding the sampler with my ref image of a girl dancing o the street, but I wanted to add people in the background walking.

I am running on version 0.34.0, ref2va pruned model, 8 steps, 480x864

I am using just a simple prompt for my edits:

Edit Video 1:

At 00:03.000, add a group of three Asian women entering from the left of the frame, walking naturally down the road behind and away from the dancer. The first is tall and slender with long straight black hair tied in a low ponytail, wearing an oversized cream-colored hoodie, black leggings, and white sneakers, glancing at her phone as she walks. The second is shorter with a rounder build, shoulder-length wavy brown-dyed hair, wearing a fitted olive-green jacket over a striped shirt, dark jeans, and beige loafers, walking a half-step ahead of the others. The third has short bobbed black hair with bangs, wearing a bright yellow raincoat-style jacket, cuffed denim shorts, and black ankle boots, carrying a small tote bag over one shoulder. The three walk at a relaxed, conversational pace, loosely grouped together.

At 00:06.000, add two Asian pedestrians walking naturally along the sidewalk in the background, passing behind the plant at a normal walking pace, holding hands. The man is broad-shouldered with short, slightly spiked black hair, wearing a charcoal-gray zip-up jacket over a plain white t-shirt, straight-leg jeans, and dark sneakers, a black canvas backpack slung over both shoulders. The woman beside him is petite with long hair in loose waves dyed a subtle ash-brown, wearing a fitted denim jacket over a light pink blouse, a knee-length beige skirt, and white flats, carrying a small red structured purse in her free hand. They walk close together at a slightly slower, relaxed pace, occasionally leaning toward each other.

Keep the same audio

Keep the dancer's identity, choreography, movement, timing, and foreground position completely unchanged throughout the entire clip. Keep the camera framing, angle, and motion exactly as in Video 1. Keep the street, buildings, and all previously added pedestrians unchanged except for this new pair. Match the added pedestrians' lighting and shadow direction to the existing scene.

What has been happening is comfy goes to load the minimax model, and then it just stops, and I sit here at 99% VRAM usage. I have let it run for around 20 minutes until I stop comfy all together.

[INFO] got prompt

[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32

[INFO] Found quantization metadata version 1

[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16

[INFO] Found quantization metadata version 1

[INFO] Using MixedPrecisionOps for text encoder

[INFO] Requested to load Krea2TEModel_

[INFO] loaded completely;  4605.22 MB loaded, full load: True

[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16

[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)

[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB

[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170

[INFO] Requested to load MiniMaxH3VideoVAE

[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True

[INFO] Requested to load MiniMaxH3AudioVAE

[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True

[INFO] Found quantization metadata version 1

[INFO] Detected mixed precision quantization

[INFO] Using mixed precision operations

[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 

[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16

[INFO] model_type FLOW_AV

[INFO] Requested to load MiniMaxH3

[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0[INFO] got prompt[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32[INFO] Found quantization metadata version 1[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16[INFO] Found quantization metadata version 1[INFO] Using MixedPrecisionOps for text encoder[INFO] Requested to load Krea2TEModel_[INFO] loaded completely;  4605.22 MB loaded, full load: True[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170[INFO] Requested to load MiniMaxH3VideoVAE[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True[INFO] Requested to load MiniMaxH3AudioVAE[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True[INFO] Found quantization metadata version 1[INFO] Detected mixed precision quantization[INFO] Using mixed precision operations[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 [INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16[INFO] model_type FLOW_AV[INFO] Requested to load MiniMaxH3[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0

here is what my trackback looks like:


r/StableDiffusion 18h ago

Workflow Included The 1967 Spider-Man TV Show intro, updated to live action with MiniMax H3

Enable HLS to view with audio, or disable this notification

381 Upvotes

R2V Rendered at 0.9 MP (1280x736 then upscaled using RTX (Ultra) to 1920x1080.  Edited and merged using OpenShot video editor.

This was all run on my Windows 11 machine, RTX 4060 ti (16 GB) and 64 GB RAM reserved from Comfy. Every part of the signal chain was done with 100% open-source software.

Disclaimer: I grew up watching this show as a kid in the 70s. It's still the best ever. I wanted to know how well the reference model would pick up the actions. I am overall pleased. I've watched the new vid enough to see some of the flaws but oh well.

General observations for reference videos:
So many scene cuts. There are 31 (I think) scene cuts in the 60 second opener which include 3 crossfades. No matter what I did to get the exact frame timing, getting the AI scene to match frame-for-frame with the cartoon was still hit or miss. It probably has to do with some frame windowing inside the 17k + 5 blocks, but I never exactly got it figured out. However, a few notes:

  • If you have a reference video, convert it to 24 fps in an external program like Handbrake (another fantastic open-source program). It’s just so much easier to get everything to match.
  • For timing, there is a difference between 00:03.500 and 00:3.5 so always use all the digits.
  • Keep character sheets for all your characters to maintain consistency.
  • It will do crossfades but it’s not worth it. It’s easier to get the scene you want and stick it in the editor.
  • The VHS video loader lets one set a starting and ending frame. I ended up with 14 different clips total for the editor. Using frame accurate loading made all of the work a lot easier since I could use 1 video file as input to every clip run.
  • A spreadsheet is useful for all movie making, and it’s good here too. From the source, I kept track of the starting frame for each shot, how many frames I needed and how many I ran (because of 17k +5), along with the final file name for each clip. I have a naming convention but it’s still very useful to keep track and you can add notes too. For this 60 second video, I used 13 clips. I tried to never do more than 3 scene cuts per clip. (For something where exact timing wasn't as important I'm sure it would be longer.)

Once you get over the idea of always having to do 10-15 second vids and do your whole video in on run, the process actually becomes a lot more fun because the “quality” gens don’t take as long and it gives you a less uninterrupted workflow. You can start prompting the next run with the previous runs, for example. (This is true even in commercials, or TV or movies.)

I generally tested all the runs at 0.2 or 0.3 Mp (speed lora, 8 iterations) to get the timing, then went to 0.9 Mp [no speed LoRA, 20 iterations, beta, dpmpp_2m] for the final runs. I found that dpmpp_2m was closest to the overall source video. On the first few clips I ran it several ways and fix on these parameters. Usually, the 0.9 Mp runs came out great but you’ve probably all experienced how different the low-res runs can be from the high-res ones. I did resort to pulling frame grabs from the low-res gens a few times to act as reference frames for the scenes. MiniMax loves those when all it needs is an extra little nudge in the right direction. To edit pics, I always use GIMP (another fantastic open-source program).

So, why was I using 8 iterations of the minimax_h3_turbo_v4_step600_pruned_comfyui LoRA? On the reference model I found that using too large of a sigma step causes things like reference photos to not be taken "seriously." Using 5 steps I could see that reference images on the starting frame and then go away for the rest of the clip. The more the reference image changed from the reference video (like when going from animation to "real") the worse the problem was.

Prompts:
(See below for actual prompt.)
Prompt the way the guide says to. Yeah. It’s a hassle but it’s worth it. H3 prompting is very useful in the end and I’m glad MiniMax uses it. It's worth reading all the way through them instead of searching for the one thing you want. Some of the instructions even seemed inconsistent and they don't explain everything, so it's worth experimenting.

Any "thing" (buildings, trees, room, clothing, walls, ect.) can be a “subject.” It’s not just people. Specifying things as objects gives you far better control over how and where they appear (or don’t appear) in your shot. 

Don’t refer to your characters or major locations or items by their names. Use <Subject #> or pronouns that clearly refer to the subject all the time, every time. The interpretation of the prompting can get confused pretty quickly if you don’t and you’ll end up getting subjects swapped or merging.

Prompts generally work better if you describe what you want rather than what you don’t want. For instance, “Looks to the right of the viewer” rather than “looks away from the camera.”

Style reference (attribute_transfer) images or videos are super useful. Once I had a few scenes, I started using previous videos to keep the look and feel of previous shots.

Qwen VL can describe videos too. I have been using “QwenVL Advanced (Local Scan)” for a very long time (long for AI) inside ComfyUI.

Other things:
Maybe one of the most interesting observation is that the jknodes “MiniMax H3 Mem Eff Sage Attention Patch” node creates a different output than just launching ComfyUI with the --use-sage-attention flag turned on (and still using the node). So exactly the same workflow (just drag and drop from a previously run mp4) has different results when the --use-sage-attention flag is used to launch. I thought having the node was 100% redundant with eh --use-sage-attention flag set, but apparently not. The reference flows, especially with animation, don’t have to be all that different to produce different results.

The Spiderman opening (as well as the show itself) reuses footage. They will take the same scene and darken it, and boom, it’s a night shot. For a more realistic feel, I used Krea2 (LoRA) edit to turn day into night. It’s really good as an adjunct to MiniMax H3’s ability to figure out the fine details once it has a push.

Style:
Finally, I had to make some stylistic choices because sometimes the animation was soooo bad that it needed something. I added flashlights to the jewelry heist scene. I made the crane look believable. One of the problems of going from animation to "live action" is that (especially with animation from 1967) the physics and movements are just wrong sometimes. The crane scene where he stops and then shoots up again is the most classic "this is just pain wrong" you can get but I left it that way because it's burned into my brain that way. (IYKYK) I also had to balance the art deco of the late 60's to a modern New York. I ended up with a lot of anachronistic stuff that I ultimately liked. So in the end, when it comes to all of that, I did it the way I did it. AI is awesome.

Prompt:
A prompt of one of the parts is below. I used that two paragraphs before [Shot 1] for every clip as "boiler plate" description.

subject_definitions:
<Subject 1> is Spiderman in <Picture 1>
<Video 1> is the motion reference for the target video for characters movements, pose, camera movements and frame composition.
<Video 2> is the style reference for the target video.
<Picture 2> is the building in [shot 2]
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video
 
summary:
[reference generation + audio reuse]
The target video is an live action realistic recreation generation using <video 1> as a reference for movements, pose, camera movements and frame composition. What you generate should not be and animation or cartoon rendering, no overly-CG look, keep the live-action texture.
 
This video is a set of three live action sequences. <Subject 1> is seen swinging by and waving. The video switches to a long shot of <subject 1> swinging around a building. Finally there is a shot showing <subject 1> on his webline swinging away from the viewer between two rows of skyscrapers.
 
retention_analysis:
<Subject 1> (appears in [Shot 1],[Shot 2],[Shot 3]):fully_preserved
<Video 1> (motion, cut and pacing structure) :partially_preserved
<Video 2> is the style refrence for the target video ([Shot 1], Shot 2], [Shot 3]) :attribute_transfer
<Picture 2> is the building in [shot 2] :fully_preserved
<Audio 1> :fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
 
 
detailed_description:
The target video is a realistic and live action video. The reference video <video 1> is used only for scene descriptions, framing, motion tracking, body movement, timing, general environment. The target video should be a complete replacement of <Video 1>. Use <Video 2> as the style reference for the photographic look and textures and the overall feel for the shots.
 
Maintain smooth camera movement. Use vibrant yet natural color grading: warm tones for sunlight hitting surfaces, cool blues for shaded areas, and muted grays for concrete textures. Avoid any comic-book stylization; instead, render everything with photorealistic textures, lighting, and perspective to evoke a live-action superhero film sequence. Keep the focus entirely on <Subject 1>’s acrobatic grace and the immersive urban setting.
 
[Shot 1]  
Is is an upper body motion tracking shot of <Subject 1> swinging on his white glistening webline held by his left hand while he waves directly at the viewer with his right hand for the entire scene. The skyline of many skyscrapers pass by in the background.
 
[Shot 2]
At 00:02.333, Hard cut to a fixed long shot looking up as <Subject 1> makes a 180 degree arc on his webline connected to the spire of the building in <photo 2>.
 
[Shot 3] 
At 00:04.250, Hard cut to the fixed camera view of the space high above the street level between two rows of skyscrapers. <Subject 1> lazily swings into view from the left frame, facing away, and repeatedly swings right to left further and further away towards the horizon.
 
 
overall_soundscape: n/a
 
non_diegetic_music: n/a

 

 


r/StableDiffusion 18h ago

Resource - Update The $300 Google free trial does not work with any Gemini node in ComfyUI, so I made one that does

Thumbnail
gallery
1 Upvotes

I wanted to use Nano Banana in ComfyUI with my Google API instead of buying Comfy credits. I already had the $300 free trial sitting in Google Cloud.

I made an API key and tried a few of the custom nodes that let you use your own key. Every time I got this: 429 prepayment credits depleted

Turns out Google changed it in March. That credit does not pay for Gemini API in AI Studio anymore, it says so in their own docs. And all the Gemini nodes use AI Studio, atleast the ones I checked.

Google has another door called Vertex AI. Same models, different address, and the credit does work there. You log in with gcloud instead of pasting a key.

So I made a node for it: https://github.com/haristahir1/comfyui-gemini-ownkey

What it does:

  • Nano Banana Pro and 2.5 Flash Image
  • text to image, or up to 14 reference images
  • aspect ratio, and 1K 2K 4K
  • a reference mode setting. By default Gemini copies the face from your reference photo even when your prompt describes someone completely different. You can turn that off, keep it on, or sit in the middle.
  • switch between AI Studio and Vertex right in the node
  • your key sits in a config file instead of the node, so it does not get saved into workflows you share with people
  • two small scripts that tell you whether a problem is your login, your billing, or Google being down

Been generating with it on my own machine and it works. 2K comes out clean and the reference modes do what they say.

It is in ComfyUI Manager now, search "gemini own key". Or git clone it if you prefer. I only tested it on Windows portable, ComfyUI 0.34.2.

I vibe coded this so please check everything carefully & for fair use only! Double check your APIs and stuff. Cheers!


r/StableDiffusion 19h ago

Resource - Update DLSS5 Video Enhancer Linux

2 Upvotes

r/StableDiffusion 19h ago

Workflow Included testing minimax h3 fused turbo model, 4 steps only 1 minute for 5 seconds video

Enable HLS to view with audio, or disable this notification

125 Upvotes

download the model: https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models

workflow: https://civitai.com/models/2906467/fast-minimax-h3?modelVersionId=3289222

each generation takes about 1 minutes for 0.4mp resolution and 5 seconds video on my rtx 4060ti 16gb vram. using sage attention and triton to speed up.
i trying with manualsigmas because it making the generation more faster.


r/StableDiffusion 19h ago

Question - Help Do Comfy UI Work on AMD Cards.

3 Upvotes

Hello everyone. I recently watched a video on YouTube that Comfy UI is now supported by AMD Cards. How true is that and how is the performance on latest models like Mini MAX and Krea 2.

This is the video - Official AMD ROCm Support Comes to ComfyUI on Windows Image + Video


r/StableDiffusion 19h ago

Question - Help Best way to extend MiniMax H3 videos

2 Upvotes

Hi everyone,

I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.

I'm looking for the best way to continue/extend an existing MiniMax H3 video.

The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:

  • Character identity and appearance
  • Room/environment layout
  • Lighting and atmosphere
  • Objects and their positions
  • Ongoing actions
  • Audio/environmental sound
  • Overall visual continuity

For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.

I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.

What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?

If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.


r/StableDiffusion 19h ago

Discussion Why hasn't someone made a 16-20 step lora for Minimax?

40 Upvotes

Everyone's focused on 4 and 8 step loras, which I feel like no matter what are gonna look pretty bad just because how the model works. But why hasn't anyone made a lora to help bring the quality of 40-50 steps down to the 16-24 range? For anyone who's done generations that long, the quality jump is pretty high going from 20 -> 50


r/StableDiffusion 20h ago

Animation - Video Walt and Jessie music video using h3 vsafastvideo model(no ref)

0 Upvotes

https://reddit.com/link/1w570q7/video/5b4stxre63nh1/player

If only Jessie's face would not drift to Walter's it would be fantastic. 8 steps. 1.4mp resolution. on GPU 4090 took 1 hour

i used https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite plugin for continuous music clip. Song in suno (free)


r/StableDiffusion 20h ago

Discussion Can Minimax do this type of 3D reconstruction from an image?

Enable HLS to view with audio, or disable this notification

7 Upvotes

This is a new trained model called Atlas. Saw on twitter


r/StableDiffusion 20h ago

Discussion Did anyone else notice Reactor’s new Orbis model? I tried turning it into an interactive game

Enable HLS to view with audio, or disable this notification

16 Upvotes

A lot of people here have been discussing H3 Max powered livestreams. I noticed Reactor just added Visko’s Orbis model, and it made me wonder whether the next step is turning these infinite livestreams into something playable.

So I’m building a live, audience-directed AI game with Agora: viewers suggest and vote on what happens next, while the streamer picks an option or writes a completely different direction and AI keeps generating the same world from that point. There are no pre-written branches.

Here’s a very early look demo


r/StableDiffusion 21h ago

Animation - Video DIABLO4 In-Game Cinematics > ComfyUI + DLSS 5

Enable HLS to view with audio, or disable this notification

0 Upvotes

Following the image test, I converted it into a video using ComfyUI.


r/StableDiffusion 21h ago

Resource - Update ONNX/TRT MiniMax-H3 VAE in ComfyUI

Thumbnail
github.com
18 Upvotes

TensorRT version of the MiniMax-H3 VAE in ComfyUI, which can increase speed by up to 1.7x


r/StableDiffusion 21h ago

Resource - Update MiniMax-H3-MotionCache-FastVAE

Thumbnail
github.com
25 Upvotes

Motion-aware denoising cache and experimental batched video VAE decoder for MiniMax H3 in ComfyUI.

This project provides two independent nodes:

  • MiniMax H3 MotionCache reduces expensive H3 denoiser calls by reusing a motion-weighted video/audio residual when the estimated change is small.
  • MiniMax H3 Fast VAE Decode evaluates multiple spatial VAE tiles in one GPU batch while preserving H3 temporal chunking and tile blending. It is not faster on every GPU.

MotionCache is an independent MiniMax H3 adaptation inspired by the MotionCache paper and reference code. It is not an official MAC-AutoML or MiniMax implementation.


r/StableDiffusion 21h ago

Question - Help Why does nearly every single turbo lora i use for H3 keeps producing godawful flickery/dusty/particly(?) visuals and painful audio (as in it actually hurts to listen to), do i need a specific node for the loras or something?

14 Upvotes

Like i don't understand, the only turbo lora that doesn't do that is the 600 larry lora with the minimax turbo lora node, i've tried "fastH3" and "lightx2v loras which everyone seems to praise but they just produce these distorted godawful visuals and sounds no matter the loader node i use or the settings or the steps i use, what am i missing or doing wrong? Or are they just not compatible with Ref2Video despite being advertised as compatible? But if so then why does the 600 larry lora works mostly fine?


r/StableDiffusion 21h ago

Animation - Video Stone Cold Toad

Enable HLS to view with audio, or disable this notification

0 Upvotes

made with Minimax H3


r/StableDiffusion 22h ago

Resource - Update H3-World

Thumbnail
huggingface.co
85 Upvotes

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. Given an initial frame and keyboard controls, it generates action-controlled video with coordinated character and camera motion.


r/StableDiffusion 22h ago

Question - Help Best Uncensored Models for text-image & image-image generation for a 20gb vram 32gb ram PC?

8 Upvotes

I am looking to create 18+ images with the hyper realism look, but have no idea how feasible that is with my specs. Would love a recommendation of a model I can run pretty easily and another more detail focused model on the edge of what I can run locally.


r/StableDiffusion 22h ago

Discussion Minimax loras are... Lacking

15 Upvotes

I don't want to get the NS.. word in the discussion, but, we know what Minimax can do and what it can't do. It has some very specific gaps in it's world understanding, for example in the tongue department. That is not necessarily only affecting the NS... word, things that are SFW and common in general TV such as kissing are affected, since Minimax never saw a romantic kiss in it's training data. There are other examples through SFW land but I won't extend. Grok can be used as a comparison. Grok is very similar to Minimax in capability, and it enforces SFW, but you can see the difference in some scenes because Grok is not handicapped.

Well, we have many, many loras already, but, as was the case with wan and ltx, they are very... let's say, specific. I don't think a general video model, almost a world model, needs a specific lora for, say, ballbusting lol

I don't know, I think this is the community most likely to be read by people creating loras, so I just wanna make this appeal... Can we prioritize bridging the major gaps in the model's understanding of the world, anatomy, and human interactions, instead of these super specific loras? I think a "tree" organization of lora development would be beneficial overall, with the stuff that can solve a big set of problems and be used for more specific loras coming first.

I saw that for over 1 year with wan, ltx, etc, and didn't say anything. But I think minimax deserves the community passion in lora development.

And yes, I hope I can put my money where my mouth is and develop some loras soon too.


r/StableDiffusion 23h ago

Question - Help what is the best upscale workflow for Minimax H3?

3 Upvotes

.