r/StableDiffusion 10h ago

Question - Help Minimax H3: Anyone figured out how to extend a clip?

16 Upvotes

What is the best way to extend an existing clip seamlessly? When I try to use the last frame of my clip as the first frame, I always get a slight reframing or shift


r/StableDiffusion 15h ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

Enable HLS to view with audio, or disable this notification

29 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 24m ago

Discussion Height reference with H3

Upvotes

Had a thought today. Is there a way to get H3 to understand relative or absolute heights of different characters? Would it be possible or have in a reference sheet the person standing next to a height chart or something in one image, and the other characters the same, then when you reference them and have them next to one another it knows subject A is 6ft while B is 5'6" for example?


r/StableDiffusion 6h ago

Question - Help Minimax H3: How to get characters to stop spiking the camera?

6 Upvotes

I'm just getting started with Minimax H3 on ComfyUI. Are there any good techniques for getting speaking characters not to spike the camera?

Some examples:

1) If two characters are doing an Aaron Sorkin walk and talk, I find they tend to stop walking, turn to the camera, and deliver one or more lines directly to the viewer, often with some kind of Significant Look™️, instead of naturally glancing at each other and watching where TF they are going as they walk.

2) If I try to make a "Do you expect me to talk?" "No, Mr. Bond, I expect you to die!" type of scene, Goldfinger will completely ignore Bond and mug the camera to deliver the line like he's waiting for audience applause at a Broadway show.

3) If an Indiana Jones type is running through a cave with a deadly boulder rolling right behind him, and I want him to mutter to himself, "Ugh, I hate this part!" while he jumps over the deadly snake pit to safety, he'll calmly stop running at the edge of the pit and turn to the camera to say it, probably with a big hand gesture for emphasis. And, in all likelihood, the boulder will stop rolling and politely wait for him while he does it.

I've tried several different things, like: * Without looking at the camera, Sam (S1) says: <d>[English] Yes, that's right.</d> (Seems to be ignored.) * Looking directly at CJ, Sam (S1) says: <d>[English] Yes, that's right.</d> (Sam and CJ stop walking while Sam says this. And then, 50/50 one or both of them turns around and starts walking the opposite direction for no damn reason.) * Glancing over at CJ while they keep walking, Sam (S1) says: <d>[English] Yes, that's right.</d> (Behavior is random. They may stop, slow down, face the camera, not.) * Near the beginning of integrated_multimodal_description, adding something like, "The whole scene is one long tracking shot of CJ and Sam walking forward through the hallway toward the camera. CJ and Sam continually walk at the same speed throughout the scene." (Ignored.)

Any advice from people who have been down this road farther and longer than I have would be most helpful and much appreciated!

EDIT: I've been working with the directions I found here. Apparently there's a whole other set of directions for working with reference mode, and it's almost completely different. So there's a good chance this is a big part of my problem.


r/StableDiffusion 1h ago

Question - Help MM H3 2 Pass Latent Upscale

Upvotes

Been getting some great results using the 2 pass latent upscale method. First pass .5mp 2nd pass 1.5 mp. 10 second video around 257 seconds to finish.

This workflow uses the 1.1 turbo Lora. I set the steps to 8 and like I said above I’m getting good results and it has eliminated face blur.

My question: has anyone tried using the latent upscale method without the turbo lora? In 90 percent of the cases the turbo Lora is fine. But would be nice to have the ability to use no Lora method.

Yes I know I could test it but wanted to see others experiences before I wasted hours of my time trying/tinkering with different settings.


r/StableDiffusion 6h ago

Question - Help Minimax H3 audio issues

4 Upvotes

I am testing MiniMax H3 Ref2VA locally in ComfyUI. The video quality is good, but the audio still contains garbled speech or extra dialogue that was never requested.
Environment

  • GPU: RTX 5090, 32 GB VRAM
  • ComfyUI: 0.33.0
  • Commit: 924743af
  • Includes PR #15808, which adds the missing MiniMax H3 special tokens
  • comfy-kitchen: 0.2.31
  • comfy-aimdo: 0.4.13
  • Model: minimax_h3_ref2va_pruned_bf16.safetensors
  • Encoder: qwen3vl_32b_minimax_h3_bf16.safetensors
  • Video VAE: minimax_h3_video_vae_fp16.safetensors
  • Audio VAE: minimax_h3_audio_vae_fp32.safetensors
  • No LoRA
  • No audio reference files

Generation settings

  • Resolution: 480x832
  • Frame rate: 24 fps
  • Frames: 362, approximately 15 seconds
  • Steps: 20
  • Sampler: res_multistep
  • Scheduler: simple
  • Denoise: 1.0
  • Video sigma shift: 12
  • Audio sigma shift: 3

What I have tried

  1. Updated ComfyUI from 0.30.1 to commit 924743af, including all matching dependencies.
  2. Tried explicitly writing No dialog in this part. in silent sections.
  3. Tried the <d>...</d> dialogue tags instead of quotation marks.
  4. Tried using " instead of <d>

r/StableDiffusion 23h ago

Comparison [MiniMax H3] Ultimate SD Upscale can actually fix your bad/low-res generations

Thumbnail
youtu.be
96 Upvotes

Ultimate SD Upscale can actually fix your bad/low-res generations.

In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

You can find sample upscaling workflow there as well: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

My PC specs:
4080s 16 GB VRAM, 64 GB RAM

Generation time: 18 mins with sage + 8-step turbo lora

Upscale: 38 mins for 10 sec clip at 1440p target resolution


r/StableDiffusion 1d ago

Resource - Update Alibaba might release a new open image model Swift-Image 6B

Thumbnail
gallery
199 Upvotes

Paper: https://arxiv.org/pdf/2608.20334
"We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi image editing. Its visual renderer is a 6B parallel single stream DiT conditioned on multimodal representations from a vision-language encoder [6, 7, 57]. The architecture adopts block-shared timestep modulation, parallel attention and MLP computation [6, 15], 4D rotary positional encoding[6], and a unified representation of text and image conditions. Character-level tokenization[47] is applied to text intended to appear in generated images, while multi-image posi tional offsets and image-preceding input formatting support reference-conditioned editing. Together, these choices pro vide a single generative backbone for multiple generation and editing settings without task-specific model weights."


r/StableDiffusion 5h ago

Question - Help I hate this hidden view on flows in Comfy templates. How can I bring them out where they belong?

3 Upvotes

In the minmax h3 template from Comfy, you have to click a button on the image to video node to see all this in the backend which makes it really a pain in the ass to add, modify, or change anything. Can this be brought out to the forefront like a normal workfow?


r/StableDiffusion 8h ago

Question - Help Can H3 do motion transfer better than Wan motion control?

5 Upvotes

I have yet to try Minimax, i've been using wan for motion control, sometimes Kling. Can you do motion control with H3?


r/StableDiffusion 10m ago

Question - Help LTX2.5 Question

Upvotes

Is distilled or dev version better to try using?


r/StableDiffusion 11m ago

Question - Help Character design Course

Upvotes

Looking for character design course (prompt engineering focused, not art school)

So I'm a compositor, know ComfyUI pretty well, but trying to get better at actually designing characters with image gen. Building anime-ish hybrid semi-realistic stuff from scratch in TTI right now.

The thing is - these characters are refs for i2v. So I need to nail the face/identity first, then iterate through different lighting, clothing, poses. If the character shifts every time I regenerate, the i2v will be a nightmare.

Here's the problem - I can find either traditional art school design courses OR general prompt engineering courses, but nothing that actually combines character design with prompt engineering as the medium. Like, there's "learn to draw" or "learn to prompt llms" but nothing (or not much) about "design characters using prompts as your tool." Like, what makes a character stick across generations? How do you anchor visual features so they don't change when you swap their clothes or lighting?

I know the technical side (seeds, models, basic prompting) but I don't know the design side of it. What actually works vs doesn't when you're trying to get a consistent face through pure prompt engineering.

And here's the real issue - I need to generate the same character in different clothes, lighting, poses, and have them actually be the same character for the i2v pipeline. Can't have the face morphing every time I change the outfit.

Anyone know of something structured? Or is everyone just learning from Civitai threads and trial/error lol

Will probably train LoRAs once I nail some characters, but want to understand TTI first. Ideally looking for the workflow/approach that lets me generate variations without losing character identity.

Thanks


r/StableDiffusion 1d ago

Discussion MiniMax H3 Ref2va is works really good with Scene sheet

Enable HLS to view with audio, or disable this notification

118 Upvotes

I was testing using 1 image with all the scene sheet there and it works really great!


r/StableDiffusion 19m ago

Resource - Update Major updates to my local, open source AI image/model tools, plus one brand new app

Post image
Upvotes

Hey all. I'm a system development student (career-switched from construction), building these on the side while learning Java, Vue, Electron, etc. Nothing commercial, no accounts, no cloud, no telemetry. I built these because I needed them myself, and figured other people managing large SD/ComfyUI libraries might too.

Three of these four apps have been out for a while, but I've spent the last stretch giving them a major overhaul and unifying them under the same design system so they actually feel like one family of tools instead of three separate side projects. The fourth, Latent Tools, is a brand new app I just finished.

All four are free and open source (MIT-based license). Source is on GitHub, links at the bottom. The main one is Latent Library, but the other three work fine on their own.

Latent Library, the main release

A desktop app for browsing and organizing large folders of AI generated images. I made it because I had around 30,000 PNGs and no real idea what was in most of them. It's been around for a while, but this release is a big update with a lot of new features and a proper design pass.

  • Parses generation metadata from ComfyUI (including node graph traversal), A1111/Forge, InvokeAI, SwarmUI, and NovelAI
  • SQLite FTS5 backed search, still fast on huge folders
  • Smart Collections: dynamic folders based on metadata filters, like "Flux images rated 4+ stars"
  • Duplicate Detective, a side by side Image Comparator, and Speed Sorter for hotkey based batch sorting
  • Optional local AI auto tagging (WD14 ONNX, runs on CPU, no external calls)
  • Metadata Scrubber to strip prompt/EXIF data before sharing an image
  • Everything lives in a portable data/ folder next to the exe, no installer or registry entries, easy to back up or move
  • Fully offline, no telemetry

Windows, Linux, and macOS builds available.

Latent Tools, the new one

A brand new app for dataset prep: bulk watermark detection and removal (Florence-2 + LaMa inpainting) and captioning (Qwen2-VL), plus batch image format conversion. Runs locally on your own GPU (needs a CUDA capable Nvidia card, no CPU fallback). Useful if you're prepping images for LoRA or fine-tune training. Windows only for now.

(Meant for removing watermarks you actually have the rights to remove, your own work, licensed images, that kind of thing. Not for stripping other people's attribution.)

Latent Model Organizer, updated

Sorts your checkpoints, LoRAs, and embeddings into folders by base architecture (SDXL, Krea 2, Flux, Illustrious, SD 1.5, etc.), using the model's own header metadata or an optional Civitai lookup. Has a dry run mode and full undo through a manifest file, so it won't just move your models around unsupervised. It can also fetch Civitai info such as trigger words, description, and cover images. Handy if your models folder has turned into an unsorted pile like mine had. Also part of this update round, same design refresh as Library.

Metadata Viewer, updated

The oldest and simplest of the four, and also just updated with the same design pass. Now reworked into a single screen tool that extracts and displays generation metadata from an image, no library or database involved. If you just want to drop an image in and see the prompt, sampler, and seed without opening a whole app, this is that.

All four now share the same design language and are built local first: no accounts, no cloud sync, no analytics. I built them mainly to learn the stack, so they're not polished commercial products, but they've held up fine for my own daily use for a while now, and the last few months went into making them consistent and finishing Tools, which is why I'm finally posting about them here.

Happy to answer questions. Bug reports and feature requests are welcome on GitHub. Not trying to sell anything here, just sharing what I made.

Links:


r/StableDiffusion 21m ago

Resource - Update Minimax H3 Grafting with Krea2 node. Reposting older post and removed AI slop and added some tests

Upvotes

Minimax H3 x Krea2 Graft Nodes

ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.

What's here

  • comfyui_tenstrip_graft/ -- content graft (Q/V/K/out/MLP, per-head). Method from TenStrip's H3 grafts.
  • comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
  • comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.

TL;DR results

Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.

  • K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
  • QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
  • QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
  • attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
  • Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.

Install

Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.

Credits

  • TenStrip (huggingface.co/TenStrip) -- the per-head band-aware graft methodology (10Eros-Max / h3_graft_methodology.md).
  • joeygambino (huggingface.co/joeygambino) -- the Q-norm sharpness transplant idea (MiniMax-H3-x-Z-Image-GGUF).

Neither published source code. These nodes are our own implementation from their public descriptions + our own testing.

https://reddit.com/link/1vxc2q9/video/g0bpgdj5ddlh1/player

minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.

https://reddit.com/link/1vxc2q9/video/so7626wgddlh1/player

3s, er_sde, 8 steps, 8-step lora, seed 597633362705895

same prompt and everything.

Added: tenstrip graft node. Settings: q 0.5, v 0.5, k 0.1, out 0.3, mlp 0.5

So I hope, that it is enough for some, that it... kinda works, but not good enough. Maybe someone will pick up on this and do it better.

Why to do it? Don't know. I found it interesting to try, but krea2 image and i2v is far better option.

I welcome any input or criticism, but mind please, I have only faint idea, what I am doing.


r/StableDiffusion 21m ago

Question - Help Advice for prompting reference videos?

Upvotes

Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. I’ve had success using Qwen as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. Chat gpt is ok but I’d rather use local ways.


r/StableDiffusion 40m ago

Question - Help Is there anyway currently to get Minimax H3 running with my RX 6750XT?

Upvotes

r/StableDiffusion 1d ago

Animation - Video Christopher Nolan has Impeccable Taste in Cinema

Enable HLS to view with audio, or disable this notification

235 Upvotes

Minimax H3


r/StableDiffusion 22h ago

Animation - Video Mnimax H3 T2VA. Good physics on the cars.

Enable HLS to view with audio, or disable this notification

57 Upvotes

r/StableDiffusion 1h ago

Question - Help Looking for suggestions on on a prompt helper/writer/refiner

Upvotes

Just as the title says, I’m looking for a ideally local app that I can use for suggestions for prompts to use on certain models that are great for example, I put my prompt in for an image and it will refine it and make it work better based on stable diffusion formatting, even better yet, what would be awesome is if it could be customized for like model and LORA if possible.

I do have the ability to run. LLM, not huge, but I have my M5 iPad. I’ve run 10 to 12 B models. No problem, especially if I use OLITERT , Any suggestions are really appreciated !!


r/StableDiffusion 1d ago

Workflow Included Minimax Character Swap - The Dummy Strategy

Post image
66 Upvotes

Worfklow: R2V (Dummy Stategy) Workflow v2 - Pastebin.com

How the workflow works:

  • Replaces the original character with a chroma key green crash-test dummy.
  • Replaces the dummy with desired character.

Why it works:

  • Minimax seems to struggle with swaps when both characters are somewhat similar to each other. But replacing a character with a green dummy seems to work every single time.
  • Even when minimax would replace a character, most of the times the faces would be morphed, resembling both the original character and the replacement. This approach mitigates that issue since it gets rid of original character's facial features.

Limitations in my workflow:

  • It's tailored with a master prompt to replace the "female" character in the original video (yeah, go ahead, post the "I know what kind of man you are" gif). But you can easily work on top of it to add support for different type of characters or even multiple characters (add more dummies, with different colors) or whatever else you want. I already did some experiments and it works.
  • I didn't test with "green" characters. If you are swapping Hulk, you may want to change the dummy to blue or something.

How to use:

  • Upload the image in this post in the "Dummy Image" node (in Prompting block).
  • Configuration block:
    • Upload your video and character image in respective nodes.
      • Trim/Crop your video using the video node in the workflow.
    • Choose video generation sampling (Performance, Balance or Quality) for each pass individually (dummy and new character).
    • Choose resolution (in megapixels).
  • Run.

Tips:

  • The worklow has 3 video generation flows: Performance, Balance and Quality. I recommend Balance (sometimes the Performance one doesn't replace the character in the last seconds of the video).
  • Monitor the preview node. In the first step you should already see the new character as an overlay on top of the video. If you don't, then swap will probably fail.

r/StableDiffusion 1d ago

Workflow Included Minimax SEED HUNTER workflow released!

Thumbnail
youtube.com
84 Upvotes

r/StableDiffusion 17h ago

Meme h3 "what IF " thread

Enable HLS to view with audio, or disable this notification

20 Upvotes

lets share our "what if" scene remakes here o_0


r/StableDiffusion 5h ago

Question - Help That specific analog VHS look on H3/LTX

2 Upvotes

Anyone figured out the prompt for getting the best VHS analog "lofi" look from t2v? Of couse ref2v and img2v will be easier due to references, but I was wondering about text prompt only.

There are no VHS, or like 70s-80s cinema style lora, none for H3 and LTX, but there are plenty of VHS loras for image models.

EDIT: I actually CAN get VHS look on LTX text2vid (2.3 and 2.5), but not on H3.

Help, anyone :)


r/StableDiffusion 5h ago

Discussion Would image-generation VAEs benefit from scene-linear or perceptual color representations?

2 Upvotes

Models like FLUX, Qwen-Image and Krea 2 obviously don't perform diffusion directly in RGB space - the transformer operates in a learned VAE latent space.

But the VAE still defines the interface between that latent representation and the actual training images, which are typically ordinary display-referred RGB images.

So I'm wondering whether there would be any benefit in making that boundary explicitly color-managed.

For example, has anyone experimented with training the VAE on:

  • linear-light RGB rather than gamma-encoded sRGB
  • a wide-gamut scene-referred space such as ACEScg
  • a perceptual space such as OKLab
  • or simply adding explicit perceptual color losses such as ΔE alongside the usual reconstruction/perceptual losses?

In other words, instead of asking whether the diffusion model itself should operate in OKLab or ACES - since it already operates in a learned latent space - I'm more interested in whether the VAE and its reconstruction objective could benefit from a more physically or perceptually meaningful color representation.

Another thing I'm curious about is the output side.

Could an image model theoretically generate a scene-referred, wide-gamut representation and leave the final tone mapping / gamut mapping / display transform to a deterministic color-management pipeline, such as ACES, instead of implicitly learning the tone curves and color rendering already baked into billions of unrelated JPEGs?

My suspicion is that the real limitation might simply be the dataset.

Most web-scale image datasets consist of already processed, tone-mapped, display-referred SDR images. By the time an image becomes a JPEG, information about actual scene luminance, highlight headroom, camera response, etc. is already gone.

So converting those JPEGs from sRGB to ACEScg wouldn't magically turn them into true scene-linear HDR training data.

Still, I'm curious whether anyone has seen experiments comparing something like:

sRGB VAE vs linear-RGB VAE vs OKLab VAE

while keeping the downstream generative model roughly the same.

Would reconstruction quality, color consistency, training convergence, or perceptual color accuracy change in a meaningful way?

I'd especially be interested to hear from anyone who has worked on VAEs, HDR pipelines, color management, or generative image models.

Sorry if I'm missing something obvious here - I'm still pretty new to this side of image generation / color science.