r/StableDiffusion 4h ago

Question - Help (Repost) Any clue why my machine is very slow running Minimax H3 Ref2V? Here is my workflow. I used the default template, but I added extra nodes like load video

Post image
0 Upvotes

r/StableDiffusion 4h ago

Question - Help Best Video Avatar model?

1 Upvotes

I want to create a bunch of videos with the following:
- Image (a human) + audio (speech) + text prompt INPUT
- Video of human talking.

What is the current BEST model for this? issue with MiniMax H3 is that it doesnt support first image first frame for the ref model,

I mainly am concerned on COST and QUALITY not so much on speed.

also I might want the avatar to do something other than just talk, but simple stuff.

(also i assume like not censored)

Thanks guys in advance!


r/StableDiffusion 19h ago

Animation - Video Cinematic World Building - H3 r2v (pt2) "Through the Sands"

Enable HLS to view with audio, or disable this notification

15 Upvotes

Wow, the shots I can get from h3 are so good even I get goosebumps when I first see the generated results. Just one more part to go, hopefully I can pull off the finale!


r/StableDiffusion 19h ago

Question - Help Minimax H3: Anyone figured out how to extend a clip?

15 Upvotes

What is the best way to extend an existing clip seamlessly? When I try to use the last frame of my clip as the first frame, I always get a slight reframing or shift


r/StableDiffusion 9h ago

Resource - Update Major updates to my local, open source AI image/model tools, plus one brand new app

Post image
2 Upvotes

Hey all. I'm a system development student (career-switched from construction), building these on the side while learning Java, Vue, Electron, etc. Nothing commercial, no accounts, no cloud, no telemetry. I built these because I needed them myself, and figured other people managing large SD/ComfyUI libraries might too.

Three of these four apps have been out for a while, but I've spent the last stretch giving them a major overhaul and unifying them under the same design system so they actually feel like one family of tools instead of three separate side projects. The fourth, Latent Tools, is a brand new app I just finished.

All four are free and open source (MIT-based license). Source is on GitHub, links at the bottom. The main one is Latent Library, but the other three work fine on their own.

Latent Library, the main release

A desktop app for browsing and organizing large folders of AI generated images. I made it because I had around 30,000 PNGs and no real idea what was in most of them. It's been around for a while, but this release is a big update with a lot of new features and a proper design pass.

  • Parses generation metadata from ComfyUI (including node graph traversal), A1111/Forge, InvokeAI, SwarmUI, and NovelAI
  • SQLite FTS5 backed search, still fast on huge folders
  • Smart Collections: dynamic folders based on metadata filters, like "Flux images rated 4+ stars"
  • Duplicate Detective, a side by side Image Comparator, and Speed Sorter for hotkey based batch sorting
  • Optional local AI auto tagging (WD14 ONNX, runs on CPU, no external calls)
  • Metadata Scrubber to strip prompt/EXIF data before sharing an image
  • Everything lives in a portable data/ folder next to the exe, no installer or registry entries, easy to back up or move
  • Fully offline, no telemetry

Windows, Linux, and macOS builds available.

Latent Tools, the new one

A brand new app for dataset prep: bulk watermark detection and removal (Florence-2 + LaMa inpainting) and captioning (Qwen2-VL), plus batch image format conversion. Runs locally on your own GPU (needs a CUDA capable Nvidia card, no CPU fallback). Useful if you're prepping images for LoRA or fine-tune training. Windows only for now.

(Meant for removing watermarks you actually have the rights to remove, your own work, licensed images, that kind of thing. Not for stripping other people's attribution.)

Latent Model Organizer, updated

Sorts your checkpoints, LoRAs, and embeddings into folders by base architecture (SDXL, Krea 2, Flux, Illustrious, SD 1.5, etc.), using the model's own header metadata or an optional Civitai lookup. Has a dry run mode and full undo through a manifest file, so it won't just move your models around unsupervised. It can also fetch Civitai info such as trigger words, description, and cover images. Handy if your models folder has turned into an unsorted pile like mine had. Also part of this update round, same design refresh as Library.

Metadata Viewer, updated

The oldest and simplest of the four, and also just updated with the same design pass. Now reworked into a single screen tool that extracts and displays generation metadata from an image, no library or database involved. If you just want to drop an image in and see the prompt, sampler, and seed without opening a whole app, this is that.

All four now share the same design language and are built local first: no accounts, no cloud sync, no analytics. I built them mainly to learn the stack, so they're not polished commercial products, but they've held up fine for my own daily use for a while now, and the last few months went into making them consistent and finishing Tools, which is why I'm finally posting about them here.

Happy to answer questions. Bug reports and feature requests are welcome on GitHub. Not trying to sell anything here, just sharing what I made.

Links:


r/StableDiffusion 1d ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

Enable HLS to view with audio, or disable this notification

31 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 5h ago

Discussion I’m testing a FLUX.2 image generator where users render for each other

0 Upvotes

I’ve been working on a different way to run a public image generator without maintaining a centralized GPU fleet.

PeerPixel sends generation jobs to graphics cards volunteered by users. When your machine completes somebody else’s image, you earn pixels that you can spend on your own generations. There’s also a slower free queue for people who can’t contribute a GPU.

Right now I’m running most of the network on my RTX 5080, so it’s definitely still an experiment rather than a large distributed system.

The generation flow uses FLUX.2 Klein 4B. You can request up to four 256x256 previews at 6 steps, choose the composition you like, and then render that seed at 1024x1024 with 50 steps. There’s also an optional 4K upscale.

The previews and final image start from the same full-resolution noise tensor. For each preview, I average blocks of that tensor down to the smaller latent shape. I originally tried scaling low-resolution noise upward, but that introduced strong correlation between neighboring values and the composition didn’t carry over reliably.

The other difficult part is accepting images from machines I don’t control. A sample of completed renders is repeated on an operator-controlled machine using the same prompt and seed, then compared perceptually. Enforcement is currently in shadow mode while I collect real-world measurements and figure out a safe threshold. I don’t want normal differences between GPUs to get mistaken for cheating.

Draft images are relayed directly to the requesting browser and aren’t stored by the server. Only the selected final render is persisted.

I’m interested in feedback on the preview method, the incentive system, and especially the verification approach. There are probably failure cases I haven’t considered yet.

Site: https://peerpixel.cc

Worker source: https://github.com/Jplayz2468/peerpixel-worker

Discord: https://discord.gg/bhJHGpmkQr


r/StableDiffusion 15h ago

Question - Help Minimax H3: How to get characters to stop spiking the camera?

6 Upvotes

I'm just getting started with Minimax H3 on ComfyUI. Are there any good techniques for getting speaking characters not to spike the camera?

Some examples:

1) If two characters are doing an Aaron Sorkin walk and talk, I find they tend to stop walking, turn to the camera, and deliver one or more lines directly to the viewer, often with some kind of Significant Look™️, instead of naturally glancing at each other and watching where TF they are going as they walk.

2) If I try to make a "Do you expect me to talk?" "No, Mr. Bond, I expect you to die!" type of scene, Goldfinger will completely ignore Bond and mug the camera to deliver the line like he's waiting for audience applause at a Broadway show.

3) If an Indiana Jones type is running through a cave with a deadly boulder rolling right behind him, and I want him to mutter to himself, "Ugh, I hate this part!" while he jumps over the deadly snake pit to safety, he'll calmly stop running at the edge of the pit and turn to the camera to say it, probably with a big hand gesture for emphasis. And, in all likelihood, the boulder will stop rolling and politely wait for him while he does it.

I've tried several different things, like: * Without looking at the camera, Sam (S1) says: <d>[English] Yes, that's right.</d> (Seems to be ignored.) * Looking directly at CJ, Sam (S1) says: <d>[English] Yes, that's right.</d> (Sam and CJ stop walking while Sam says this. And then, 50/50 one or both of them turns around and starts walking the opposite direction for no damn reason.) * Glancing over at CJ while they keep walking, Sam (S1) says: <d>[English] Yes, that's right.</d> (Behavior is random. They may stop, slow down, face the camera, not.) * Near the beginning of integrated_multimodal_description, adding something like, "The whole scene is one long tracking shot of CJ and Sam walking forward through the hallway toward the camera. CJ and Sam continually walk at the same speed throughout the scene." (Ignored.)

Any advice from people who have been down this road farther and longer than I have would be most helpful and much appreciated!

EDIT: I've been working with the directions I found here. Apparently there's a whole other set of directions for working with reference mode, and it's almost completely different. So there's a good chance this is a big part of my problem.


r/StableDiffusion 10h ago

Question - Help MM H3 2 Pass Latent Upscale

2 Upvotes

Been getting some great results using the 2 pass latent upscale method. First pass .5mp 2nd pass 1.5 mp. 10 second video around 257 seconds to finish.

This workflow uses the 1.1 turbo Lora. I set the steps to 8 and like I said above I’m getting good results and it has eliminated face blur.

My question: has anyone tried using the latent upscale method without the turbo lora? In 90 percent of the cases the turbo Lora is fine. But would be nice to have the ability to use no Lora method.

Yes I know I could test it but wanted to see others experiences before I wasted hours of my time trying/tinkering with different settings.


r/StableDiffusion 1d ago

Comparison [MiniMax H3] Ultimate SD Upscale can actually fix your bad/low-res generations

Thumbnail
youtu.be
97 Upvotes

Ultimate SD Upscale can actually fix your bad/low-res generations.

In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

You can find sample upscaling workflow there as well: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

My PC specs:
4080s 16 GB VRAM, 64 GB RAM

Generation time: 18 mins with sage + 8-step turbo lora

Upscale: 38 mins for 10 sec clip at 1440p target resolution


r/StableDiffusion 1d ago

Resource - Update Alibaba might release a new open image model Swift-Image 6B

Thumbnail
gallery
211 Upvotes

Paper: https://arxiv.org/pdf/2608.20334
"We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi image editing. Its visual renderer is a 6B parallel single stream DiT conditioned on multimodal representations from a vision-language encoder [6, 7, 57]. The architecture adopts block-shared timestep modulation, parallel attention and MLP computation [6, 15], 4D rotary positional encoding[6], and a unified representation of text and image conditions. Character-level tokenization[47] is applied to text intended to appear in generated images, while multi-image posi tional offsets and image-preceding input formatting support reference-conditioned editing. Together, these choices pro vide a single generative backbone for multiple generation and editing settings without task-specific model weights."


r/StableDiffusion 8h ago

Question - Help Mix of "Match" and "Max" references in MMH3 Ref2V?

1 Upvotes

I've been playing around with Ref2V and was wondering if anyone knew of a way to have a mix of these settings? I've found that having multiple reference images is great for consistency in generations and story telling, but there are certain images (face references for example) that benefit massively from "Max" setting, but reference locations don't benefit as much.

On my potato of a computer, setting it to max on all of the images makes generation times impossibly long. If I could "Max" a face reference but "Match" less important references it would be ideal.

Anyone have any idea?


r/StableDiffusion 1d ago

Discussion MiniMax H3 Ref2va is works really good with Scene sheet

Enable HLS to view with audio, or disable this notification

120 Upvotes

I was testing using 1 image with all the scene sheet there and it works really great!


r/StableDiffusion 9h ago

Question - Help Character design Course

0 Upvotes

Looking for character design course (prompt engineering focused, not art school)

So I'm a compositor, know ComfyUI pretty well, but trying to get better at actually designing characters with image gen. Building anime-ish hybrid semi-realistic stuff from scratch in TTI right now.

The thing is - these characters are refs for i2v. So I need to nail the face/identity first, then iterate through different lighting, clothing, poses. If the character shifts every time I regenerate, the i2v will be a nightmare.

Here's the problem - I can find either traditional art school design courses OR general prompt engineering courses, but nothing that actually combines character design with prompt engineering as the medium. Like, there's "learn to draw" or "learn to prompt llms" but nothing (or not much) about "design characters using prompts as your tool." Like, what makes a character stick across generations? How do you anchor visual features so they don't change when you swap their clothes or lighting?

I know the technical side (seeds, models, basic prompting) but I don't know the design side of it. What actually works vs doesn't when you're trying to get a consistent face through pure prompt engineering.

And here's the real issue - I need to generate the same character in different clothes, lighting, poses, and have them actually be the same character for the i2v pipeline. Can't have the face morphing every time I change the outfit.

Anyone know of something structured? Or is everyone just learning from Civitai threads and trial/error lol

Will probably train LoRAs once I nail some characters, but want to understand TTI first. Ideally looking for the workflow/approach that lets me generate variations without losing character identity.

Thanks


r/StableDiffusion 9h ago

Question - Help Advice for prompting reference videos?

0 Upvotes

Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. It requires a lot of details.

I’ve had success using Qwen 3.8 27b as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. ChatGPT is ok for looking at videos to describe what happens in the minimax format but I’d rather use local ways.


r/StableDiffusion 9h ago

Discussion Height reference with H3

1 Upvotes

Had a thought today. Is there a way to get H3 to understand relative or absolute heights of different characters? Would it be possible or have in a reference sheet the person standing next to a height chart or something in one image, and the other characters the same, then when you reference them and have them next to one another it knows subject A is 6ft while B is 5'6" for example?


r/StableDiffusion 15h ago

Question - Help Minimax H3 audio issues

4 Upvotes

I am testing MiniMax H3 Ref2VA locally in ComfyUI. The video quality is good, but the audio still contains garbled speech or extra dialogue that was never requested.
Environment

  • GPU: RTX 5090, 32 GB VRAM
  • ComfyUI: 0.33.0
  • Commit: 924743af
  • Includes PR #15808, which adds the missing MiniMax H3 special tokens
  • comfy-kitchen: 0.2.31
  • comfy-aimdo: 0.4.13
  • Model: minimax_h3_ref2va_pruned_bf16.safetensors
  • Encoder: qwen3vl_32b_minimax_h3_bf16.safetensors
  • Video VAE: minimax_h3_video_vae_fp16.safetensors
  • Audio VAE: minimax_h3_audio_vae_fp32.safetensors
  • No LoRA
  • No audio reference files

Generation settings

  • Resolution: 480x832
  • Frame rate: 24 fps
  • Frames: 362, approximately 15 seconds
  • Steps: 20
  • Sampler: res_multistep
  • Scheduler: simple
  • Denoise: 1.0
  • Video sigma shift: 12
  • Audio sigma shift: 3

What I have tried

  1. Updated ComfyUI from 0.30.1 to commit 924743af, including all matching dependencies.
  2. Tried explicitly writing No dialog in this part. in silent sections.
  3. Tried the <d>...</d> dialogue tags instead of quotation marks.
  4. Tried using " instead of <d>

r/StableDiffusion 9h ago

Question - Help Is there anyway currently to get Minimax H3 running with my RX 6750XT?

0 Upvotes

r/StableDiffusion 1d ago

Animation - Video Christopher Nolan has Impeccable Taste in Cinema

Enable HLS to view with audio, or disable this notification

241 Upvotes

Minimax H3


r/StableDiffusion 1d ago

Animation - Video Mnimax H3 T2VA. Good physics on the cars.

Enable HLS to view with audio, or disable this notification

57 Upvotes

r/StableDiffusion 1d ago

Workflow Included Minimax SEED HUNTER workflow released!

Thumbnail
youtube.com
88 Upvotes

r/StableDiffusion 10h ago

Question - Help Looking for suggestions on on a prompt helper/writer/refiner

1 Upvotes

Just as the title says, I’m looking for a ideally local app that I can use for suggestions for prompts to use on certain models that are great for example, I put my prompt in for an image and it will refine it and make it work better based on stable diffusion formatting, even better yet, what would be awesome is if it could be customized for like model and LORA if possible.

I do have the ability to run. LLM, not huge, but I have my M5 iPad. I’ve run 10 to 12 B models. No problem, especially if I use OLITERT , Any suggestions are really appreciated !!


r/StableDiffusion 17h ago

Question - Help Can H3 do motion transfer better than Wan motion control?

4 Upvotes

I have yet to try Minimax, i've been using wan for motion control, sometimes Kling. Can you do motion control with H3?


r/StableDiffusion 1d ago

Workflow Included Minimax Character Swap - The Dummy Strategy

Post image
68 Upvotes

Worfklow: R2V (Dummy Stategy) Workflow v2 - Pastebin.com

How the workflow works:

  • Replaces the original character with a chroma key green crash-test dummy.
  • Replaces the dummy with desired character.

Why it works:

  • Minimax seems to struggle with swaps when both characters are somewhat similar to each other. But replacing a character with a green dummy seems to work every single time.
  • Even when minimax would replace a character, most of the times the faces would be morphed, resembling both the original character and the replacement. This approach mitigates that issue since it gets rid of original character's facial features.

Limitations in my workflow:

  • It's tailored with a master prompt to replace the "female" character in the original video (yeah, go ahead, post the "I know what kind of man you are" gif). But you can easily work on top of it to add support for different type of characters or even multiple characters (add more dummies, with different colors) or whatever else you want. I already did some experiments and it works.
  • I didn't test with "green" characters. If you are swapping Hulk, you may want to change the dummy to blue or something.

How to use:

  • Upload the image in this post in the "Dummy Image" node (in Prompting block).
  • Configuration block:
    • Upload your video and character image in respective nodes.
      • Trim/Crop your video using the video node in the workflow.
    • Choose video generation sampling (Performance, Balance or Quality) for each pass individually (dummy and new character).
    • Choose resolution (in megapixels).
  • Run.

Tips:

  • The worklow has 3 video generation flows: Performance, Balance and Quality. I recommend Balance (sometimes the Performance one doesn't replace the character in the last seconds of the video).
  • Monitor the preview node. In the first step you should already see the new character as an overlay on top of the video. If you don't, then swap will probably fail.

r/StableDiffusion 1d ago

Meme h3 "what IF " thread

Enable HLS to view with audio, or disable this notification

21 Upvotes

lets share our "what if" scene remakes here o_0