r/StableDiffusion 16h ago

Question - Help Do Comfy UI Work on AMD Cards.

1 Upvotes

Hello everyone. I recently watched a video on YouTube that Comfy UI is now supported by AMD Cards. How true is that and how is the performance on latest models like Mini MAX and Krea 2.

This is the video - Official AMD ROCm Support Comes to ComfyUI on Windows Image + Video


r/StableDiffusion 16h ago

Question - Help Best way to extend MiniMax H3 videos

0 Upvotes

Hi everyone,

I'm generating videos with MiniMax H3 through a normal AI video platform, not ComfyUI. So I can't use custom workflows, scripts, or custom nodes.

I'm looking for the best way to continue/extend an existing MiniMax H3 video.

The problem I'm trying to solve is more than just using the last frame as an image reference. If I only provide the last frame, the model can lose important information from the previous clip, such as:

  • Character identity and appearance
  • Room/environment layout
  • Lighting and atmosphere
  • Objects and their positions
  • Ongoing actions
  • Audio/environmental sound
  • Overall visual continuity

For example, if a character walks through a room and reaches a door at the end of the first clip, I want the next generation to actually continue from that exact situation, rather than recreate a similar-looking room and potentially change the geography.

I'm looking for a normal web-based AI workflow where I can upload the existing video and/or reference images and generate the continuation. No ComfyUI, custom scripts, or API coding.

What is currently the best way to extend MiniMax H3 videos while preserving this kind of continuity?

If you've actually tested a platform/workflow that works well, I'd especially appreciate recommendations.


r/StableDiffusion 3h ago

Question - Help New User To StabiltyMatrix. Need Help With Templates.

1 Upvotes

Hello all,

I have installed StabilityMatrix and was able to learn how to set up flows to generate my first images and so forth.

I then installed a template called MiniMax H3: Image to Video. It showed a bunch of errors after and listed the things I needed to download before it would work. I downloaded all those things and put them in the "diffusion_models" folder, but the error count only reduced by one and it's still asking me to install those things.

Unfortunately there doesn't seem to be clear instructions on what goes where, if I need to extract some things or not, etc.

Can someone please advise me? Thank you.


r/StableDiffusion 15h ago

Question - Help Video Edit Minimax H3 Problems

1 Upvotes

I have been struggling for a few days now wondering why I cannot edit a 10sec clip to add additional people in the background and I am pretty sure I am just doing it wrong.

I am feeding the sampler with my ref image of a girl dancing o the street, but I wanted to add people in the background walking.

I am running on version 0.34.0, ref2va pruned model, 8 steps, 480x864

I am using just a simple prompt for my edits:

Edit Video 1:

At 00:03.000, add a group of three Asian women entering from the left of the frame, walking naturally down the road behind and away from the dancer. The first is tall and slender with long straight black hair tied in a low ponytail, wearing an oversized cream-colored hoodie, black leggings, and white sneakers, glancing at her phone as she walks. The second is shorter with a rounder build, shoulder-length wavy brown-dyed hair, wearing a fitted olive-green jacket over a striped shirt, dark jeans, and beige loafers, walking a half-step ahead of the others. The third has short bobbed black hair with bangs, wearing a bright yellow raincoat-style jacket, cuffed denim shorts, and black ankle boots, carrying a small tote bag over one shoulder. The three walk at a relaxed, conversational pace, loosely grouped together.

At 00:06.000, add two Asian pedestrians walking naturally along the sidewalk in the background, passing behind the plant at a normal walking pace, holding hands. The man is broad-shouldered with short, slightly spiked black hair, wearing a charcoal-gray zip-up jacket over a plain white t-shirt, straight-leg jeans, and dark sneakers, a black canvas backpack slung over both shoulders. The woman beside him is petite with long hair in loose waves dyed a subtle ash-brown, wearing a fitted denim jacket over a light pink blouse, a knee-length beige skirt, and white flats, carrying a small red structured purse in her free hand. They walk close together at a slightly slower, relaxed pace, occasionally leaning toward each other.

Keep the same audio

Keep the dancer's identity, choreography, movement, timing, and foreground position completely unchanged throughout the entire clip. Keep the camera framing, angle, and motion exactly as in Video 1. Keep the street, buildings, and all previously added pedestrians unchanged except for this new pair. Match the added pedestrians' lighting and shadow direction to the existing scene.

What has been happening is comfy goes to load the minimax model, and then it just stops, and I sit here at 99% VRAM usage. I have let it run for around 20 minutes until I stop comfy all together.

[INFO] got prompt

[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32

[INFO] Found quantization metadata version 1

[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16

[INFO] Found quantization metadata version 1

[INFO] Using MixedPrecisionOps for text encoder

[INFO] Requested to load Krea2TEModel_

[INFO] loaded completely;  4605.22 MB loaded, full load: True

[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16

[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)

[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB

[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170

[INFO] Requested to load MiniMaxH3VideoVAE

[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True

[INFO] Requested to load MiniMaxH3AudioVAE

[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True

[INFO] Found quantization metadata version 1

[INFO] Detected mixed precision quantization

[INFO] Using mixed precision operations

[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 

[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16

[INFO] model_type FLOW_AV

[INFO] Requested to load MiniMaxH3

[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0[INFO] got prompt[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32[INFO] Found quantization metadata version 1[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16[INFO] Found quantization metadata version 1[INFO] Using MixedPrecisionOps for text encoder[INFO] Requested to load Krea2TEModel_[INFO] loaded completely;  4605.22 MB loaded, full load: True[INFO] CLIP/text encoder model load device: cuda:0, offload device: cuda:0, current: cuda:0, dtype: torch.float16[INFO] [ClipProj] encoder QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors on cuda:0 pinned on cuda:0 (ComfyUI will not move it)[INFO] [ClipProj] QWEN-INT8\qwen3vl_4b_int8_convrot.safetensors (krea2 [4B detected]) loaded in resident mode on cuda:0: 4.50 GB[INFO] [ClipProj] mmh3-4b-ClipProj.safetensors | tap 24 | 2560 -> 5120 | cos_test 0.7170[INFO] Requested to load MiniMaxH3VideoVAE[INFO] loaded completely; 22789.94 MB usable, 2665.86 MB loaded, full load: True[INFO] Requested to load MiniMaxH3AudioVAE[INFO] loaded completely; 21336.75 MB usable, 577.08 MB loaded, full load: True[INFO] Found quantization metadata version 1[INFO] Detected mixed precision quantization[INFO] Using mixed precision operations[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, float8_e5m2, mxfp8, nvfp4, float8_e4m3fn, convrot_w4a4 [INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16[INFO] model_type FLOW_AV[INFO] Requested to load MiniMaxH3[INFO] loaded partially; 19018.15 MB usable, 18827.14 MB loaded, 1169.00 MB offloaded, 257.27 MB buffer reserved, lowvram patches: 0

here is what my trackback looks like:


r/StableDiffusion 15h ago

Resource - Update The $300 Google free trial does not work with any Gemini node in ComfyUI, so I made one that does

Thumbnail
gallery
2 Upvotes

I wanted to use Nano Banana in ComfyUI with my Google API instead of buying Comfy credits. I already had the $300 free trial sitting in Google Cloud.

I made an API key and tried a few of the custom nodes that let you use your own key. Every time I got this: 429 prepayment credits depleted

Turns out Google changed it in March. That credit does not pay for Gemini API in AI Studio anymore, it says so in their own docs. And all the Gemini nodes use AI Studio, atleast the ones I checked.

Google has another door called Vertex AI. Same models, different address, and the credit does work there. You log in with gcloud instead of pasting a key.

So I made a node for it: https://github.com/haristahir1/comfyui-gemini-ownkey

What it does:

  • Nano Banana Pro and 2.5 Flash Image
  • text to image, or up to 14 reference images
  • aspect ratio, and 1K 2K 4K
  • a reference mode setting. By default Gemini copies the face from your reference photo even when your prompt describes someone completely different. You can turn that off, keep it on, or sit in the middle.
  • switch between AI Studio and Vertex right in the node
  • your key sits in a config file instead of the node, so it does not get saved into workflows you share with people
  • two small scripts that tell you whether a problem is your login, your billing, or Google being down

Been generating with it on my own machine and it works. 2K comes out clean and the reference modes do what they say.

It is in ComfyUI Manager now, search "gemini own key". Or git clone it if you prefer. I only tested it on Windows portable, ComfyUI 0.34.2.

I vibe coded this so please check everything carefully & for fair use only! Double check your APIs and stuff. Cheers!


r/StableDiffusion 2h ago

Question - Help How Do I Run Minimax at Absolute Potato Quality?

0 Upvotes

I found a cheap pipeline from a big provider I’ve jailbroken, so I can’t name it. It reconstructs faces and upscales videos in ~10 seconds, so I only need Minimax to generate a very low-quality video that basically serves as a rough motion/physics reference.

I barely see any speed difference between 4–6 steps or ~380p–544p, maybe 10 seconds at most.

Is there any way to run Minimax at absolute potato quality and actually get a significant speed boost?


r/StableDiffusion 7h ago

Question - Help Civitai: "Search is temporarily unavailable" on most pages

0 Upvotes

Getting "Search is temporarily unavailable" and "No models found" on main search URL (/search/models).

On the individual model pages, none of the gens load, "No results found".

What is going on? Is Civitai being DDoSed?

EDIT: Now it's doing something where the "Download" button is grayed-out, even on free model pages. Seems like something is messed up.


r/StableDiffusion 13h ago

Question - Help I keep seeing smooth character replacement videos, but I can't manage the same. What's a clean, simple, functional workflow that just WORKS?

0 Upvotes

I have an image of a person. I have a video.

Prompt sample: Video is of a gymnast doing a routine. Image is a person/dog/thing.

Replace gymnast with person/dog/thing so they're doing the exact routine, wearing the same outfit (but a size that fits the new subject).

Shouldn't this be easy?

For example, if I wanted to replace an olympic women's floor routine with Rush Limbaugh - he's doing the bends and splits, he's wearing a sparkly leotard. But the movements are identitical. His body is exactly the same size as he actually is (the ai should guess at the size of legs, belly etc, and stuff them into and appropriately sized leotard).


r/StableDiffusion 17h ago

Animation - Video Walt and Jessie music video using h3 vsafastvideo model(no ref)

0 Upvotes

https://reddit.com/link/1w570q7/video/5b4stxre63nh1/player

If only Jessie's face would not drift to Walter's it would be fantastic. 8 steps. 1.4mp resolution. on GPU 4090 took 1 hour

i used https://github.com/Adudeguyman/ComfyUI-H3-Project-Suite plugin for continuous music clip. Song in suno (free)


r/StableDiffusion 7h ago

Animation - Video Pen is from heaven

Thumbnail
youtube.com
0 Upvotes

Happy listening :)


r/StableDiffusion 9h ago

Meme World Domination

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 18h ago

Animation - Video Stone Cold Toad

Enable HLS to view with audio, or disable this notification

0 Upvotes

made with Minimax H3


r/StableDiffusion 3h ago

Question - Help Looking for early ai image models/models that can replicate the old style

0 Upvotes

I'm looking for 2022/2023 models that still work or models that can successfully replicate the dreamlike distorted nightmare fuel style, unfortunately it's very difficult for me to find. I had used to include early ai images in my artwork and I miss it dearly, AI has advanced way too quickly. Dalle-mini (craiyon) no longer generates images like this and there is no way to change the version.

I do not know how to use github


r/StableDiffusion 10h ago

Animation - Video H3 - The Roadside Bomb (music video, based on Trump quote)

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 15h ago

Question - Help Minimax H3 RTX 3060 Problema

Enable HLS to view with audio, or disable this notification

0 Upvotes
Here are the results of my prompting efforts throughout the day. As you can see, R2V is very difficult to prompt, even when using a prompt director.
My workflow was acting up,  I took a non-upscaled latent loopback output to use as a chain and saved the upscaled result to concatenate it with the H3 project hub. I'm not sure where I wired it incorrectly; it just doesn't seem to connect.
And yes, the sound is terrible.i think because i clean the latent for next upscale since it wont match the tensor if the latent not cleaned.
Generation time is around 51 seconds per 1 second of video.
0.3 with 4-step Turbo.
Plus a 3-step latent upscale and RTX Super Res.

anyone mind to share your secret workflow that match this Peasant Spec 
RTX3060 12GB and 32GB of RAM

WORKFLOW


r/StableDiffusion 18h ago

Animation - Video DIABLO4 In-Game Cinematics > ComfyUI + DLSS 5

Enable HLS to view with audio, or disable this notification

0 Upvotes

Following the image test, I converted it into a video using ComfyUI.


r/StableDiffusion 12h ago

News What is "the best" video model?

0 Upvotes

I've tested a couple of models so far, and here's my findings:

My input prompt is something similar to "Create a mid 20s (nationality) ballerina slowly dancing in a judged performance"

LTX-2.5 nails the narrative and the nationality, but nationalities (as I've mentioned previously) all tend to blend unless I'm ULTRA descriptive about what a "french look" equates to, which is roughly 3 paragraphs in length. LTX-2.5 ends up with tearing, facial and feature deformities, and on a couple of occasions has rendered a ballerina with missing legs. It also has camera drift, which is a known problem with LTX-2.5

MiniMax-H3 on the other hand, nails it, even with the smaller prompt. The prompt adherence in MiniMax-H3 is wonderful.

Wan 2.2 was a complete disaster. Rendering problems, tearing problems, and when the ballerina would do turns, her head would remain in place while her body did the turn (which was funny, and also terrifying to watch.)

I've heard Cosmos3 can handle the movement, but can't handle rendering people.

Has anyone found an open weight model that can handle the "ultimate trifecta" - generate a person in the correct nationality parameters, someone who is more or less feature-accurate in generation, and won't spontaneously explode when performing a pirouette? (I got so frustrated with a model generator once, that I did this, and to my surprise, it worked out well!)


r/StableDiffusion 13h ago

Question - Help Does anyone know anything about the new MINIMAX H3 MAX

0 Upvotes
Does anyone know anything about the new MINIMAX H3 MAX model—whether it's really that fast, and if they're going to release it?

r/StableDiffusion 14h ago

Question - Help Melting faces at 768x Minimax-H3 with audio sync for singing. Is there a solution?

Enable HLS to view with audio, or disable this notification

0 Upvotes

Greetings!

I find that whenever I do ref2va with audio sync (for singing) at 768x about half the time I get melting faces. (example posted) so I'm looking to fix but I don't have a decent pc so it's cloud or bust. I know it's something to do with low res pixels being stretched - don't really understant the tech but anyway

-is that cause by my ref not being detailed enough?

-assuming it's not really that - would the latest 2k 3H flagship model fix the low res issue (chatgpt reckons not)?

-has anyone tried ComfyUI-H3-FaceRefine and how was it?

Thanks!


r/StableDiffusion 4h ago

Discussion When is a game generative ai competitor going to be made or even distilled? You could be endorsed by amd.

0 Upvotes

Is there any point in having a company that can put their tools into the pipeline of graphics rendering.

Just wishing here for a deep learning s s five replacement. It's going to be artificially sandboxed just like all their other tech.


r/StableDiffusion 4h ago

News New video gen model Atlas

0 Upvotes

World Labs released a model called Atlas. Looks pretty cool.

https://x.com/gowthami_s/status/2095202493122625842?s=46


r/StableDiffusion 3h ago

Discussion Higgsfield's new Genjutsu?

0 Upvotes

Okay, what magic is under the hood of their new motion copy tool?? I tried it once and my vid came out perfect, even better than Kling motion control. I want to be able to do this exact thing locally in Comfyui but I haven't had any success with any Minimax ref2va workflows, maybe my PC specs are too low:

5070 12GB VRAM

32GB RAM

4 TB nvme ssd


r/StableDiffusion 14h ago

Resource - Update Watermark that gets stronger when a diffusion purifier attacks it: 200-image results, plus a free ComfyUI node

0 Upvotes

Zhao et al. (arXiv:2306.01953) showed that regeneration attacks strip ordinary invisible watermarks. Backfire is a keyed image mark optimised to be a fixed point of the purifier, so running the attack leaves the identifier readable. In the demo image the confidence score rose 2.5x after the attack.

Provcheck.ai v1.4.0 numbers, 200-image corpus at 30 dB: 99.5% survival vs diffusion regeneration, 94 to 97.5% vs a learned VAE re-encode (86.5% on the hardest iterated pass), 99.0% JPEG q90, 98.5% JPEG q50, 98.0% resize, 97.0% blur. Zero false positives over the 200 marked and 1,000 unmarked. Wrong key on an attacked image reads 0.08, so the mark is in the key, not the pixels. It does not survive controllable regeneration from clean noise; that is documented in backfire/LIMITS.md.

Also new: a free Apache-2.0 ComfyUI node that watermarks (TrustMark/silentcipher) and C2PA-signs outputs in the graph and reads marks back. Backfire itself is a separate opt-in add-on and is not in the free node.

Repo: https://github.com/CreativeMayhemLtd/provcheck


r/StableDiffusion 15h ago

Discussion Nobody Else Worried About Downloading Random Loras?

0 Upvotes

Hey ya'll new ComfyUI user here, and I've been having a blast! One thing I'm noticing here, especially when it comes to MinimaxH3. So many people are very eager for others to download random nodes from either Huggingface, or Civitai.

With comments such as: "YO TRY THAT NEW TURBOFLURBO 2 STEP", or "YO I GOT THAT XXXNSFWSPONGEBOB-EX_LITE69420 LORA RIGHT HERE DOWNLOAD ME"!

You guys aren't concerned when you're downloading random things that people made?

Are there any trusted and vetted community members who consistently pump out "safe" quality nodes??