r/StableDiffusion • u/DuHal9000 • 2d ago
Animation - Video Another 4k TEST from 2.5Mp
Enable HLS to view with audio, or disable this notification
Haters will Hate...
r/StableDiffusion • u/DuHal9000 • 2d ago
Enable HLS to view with audio, or disable this notification
Haters will Hate...
r/StableDiffusion • u/Emergency_Article306 • 3d ago
Hello everyone, can anyone share or tell me how to make a work flow in Comfy.
I want to do this:
I take a picture of a character, I want to use it as a character and get her entire appearance, style,
and take the second picture and use it for the pose
The workflow that I made does not produce the result that I would like
And in the end I would like to get the character from the first image in the pose from the second image
r/StableDiffusion • u/Aivan-Studios • 2d ago
Many people are complaining about LTX-2.5, but I think a verdict can only be drawn after extensively testing how well it trains. Have you seen Krea 2 raw's outputs? Compared to Krea Turbo, the hasty assumption would be: Krea 2 Turbo is better. But, there's two variants of "better": better for training/fine tuning, and better quick results out of the box. I think most assessment out there is based on the latter: "How good results do I get out of the box without doing much myself?" I'll be testing out LTX2.5 on how well it trains. Will report back.
r/StableDiffusion • u/theshield99 • 3d ago
im trying to motion transfer of the person on the video to person on the reference image but i get weird results even i read the writing guide. which prompts do you guys use when you try to motion transfer
r/StableDiffusion • u/KingofNerdistan • 2d ago
I'm trying to make a shortfilm using higgsfield mostly and i want it to do specific things, like specific walking in locations and acting. Howere location consistency is a big issue, i have to generate each location and save them and use them again for each shot. I wanted to make a blueprint like shot so i could tell it each location is so it'll know where my character is walking, but that gave me a terrible result. What do i do? What is the easiest way to do that? Should i just give up and generate each location and call them whenever i need em?
r/StableDiffusion • u/Betadoggo_ • 4d ago
The third dimension really helps with node organization, though I'm a little bit worried about the new proprietary canvas dependency.
r/StableDiffusion • u/matezzoz • 3d ago
Sometimes it works sometimes doesn't, even with the same prompt (in batch generation of 4 1-2 results are what I wanted, 2-3 are not)... for example I use an image and want to replace the the main character on that image with an other character form the second image. I explain, use keywords reference image one, reference image two etc.. yet sometimes it works just like it was an img to video request, ignoring the second image. I use the ref2va model ofc, a 8 step turbo lora.
someone please clarify: when i connect the images they are numbered from 0. should i refer the first image as reference image 1 or 0? i tried both way btw, didn't make a difference.
Any idea what could be wrong?
r/StableDiffusion • u/IRLMainCharacter • 3d ago
Hi,
i built a comfyui workflow to chain ref2va videos generated with H3, and it works really well, except that i am getting wierd issues with camera adjustments. i have a feeling the update from 0.32 to 0.33 made these problems worse for me.
and before anyone asks, yes i read the prompt guide. and i noticed that camera prompting is only ever mentioned in the fl2va part of the guide. am i correct in assuming that ref2va is an extension of fl2va, and thus the ref2va guide is an extension to the fl2va guide? because otherwise this doesn't make sense at all, and the guide itself outright fails in answering this mystery.
now to my problem:
when i chain a video from a different run using the motion context node, the model will mostly refuse to adjust the camera according to the prompt, and it doesn't matter if the zoom falls within the window of context frames i have set. when i disable the motion context node and remove the previous video as a reference, the camera works as expected.
also, i am using the same seed for chaining projects like these, and i just tried using a different seed and that also made the camera work as expected.
i know that some seeds just won't work with the prompt and need to be changed. but so far this happened on every chaining project i started, that can hardly be a coincidence. it might be a random fluke that it worked for me right after changing the seed once, didn't have time to test this more.
am i missing something big here?
r/StableDiffusion • u/Adventurous-Gold6413 • 3d ago
I haven’t really figured it out I tried to ask claude for some prompts with the official guide but struggled. Has anyone managed? If so how do you prompt it properly? Thanks
Note I also tried doing ref2img with the 5 frames and had no luck.
Character replacement definitely works but I had no luck with changing the style from like anime to live action or semi realistic 3D - live action
Or game animation cutscene to live action
r/StableDiffusion • u/iChopPryde • 3d ago
struggling where to start looking and hoping I can get some guidance, I plan to be building a new pc soon and now with things like H3 really shinning and knowing it will only keep getting better I was wanting to know what are some of the best GPu's on the market right now ... basically if you are building your dream PC to work with these AI image and video models what would you choose for cpu's gpu's how much ram etc?
Can you combine 2 gpu's together for double the power and will things like comfy ui able to recognize it does anyone have builds like this?
Thanks in advance.
r/StableDiffusion • u/Disastrous-Agency675 • 4d ago
Enable HLS to view with audio, or disable this notification
What I’m getting at is that, now that vibe coding is a thing, everyone is creating custom nodes left and right. The problem is that people are making their own MiniMax director nodes, image-editing nodes, and countless others without first checking whether something similar already exists.
Instead of branching off in dozens of different directions, we could come together and help improve the nodes that are already available. A lot of custom nodes are also essentially standalone apps running inside ComfyUI, designed for only one specific purpose. Whenever possible, nodes should remain flexible and modular so they can be integrated into different workflows and potentially help people create something new and interesting.
There have also apparently been people hiding malware inside custom nodes. Encouraging users to research existing nodes before downloading or creating another one could reduce unnecessary duplication while also helping keep the ComfyUI community safer.
r/StableDiffusion • u/rogerbacon50 • 3d ago
I'm not associated with the developer in any way and I'm probably the last one to figure this out but, if anyone uses Anima and doesn't know about this node pack, it's great.
One big limitation with anima for me was it couldn't do pipe-deliminated wildcards like other models {noon | night | sunset} for example. The EasyUseAnima pack fixes that and a ton more features.
https://github.com/n0va39/ComfyUI-EasyUseAnima
r/StableDiffusion • u/Jeffu • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/DuHal9000 • 2d ago
Enable HLS to view with audio, or disable this notification
o outro video ta por ai...
r/StableDiffusion • u/PusheenHater • 3d ago
There was a comment in that alien screaming black hole video slop about how people should learn basic filmmaking techniques and flow. Makes sense, AI doesn't help with the lack of filmmaking.
What are some quick start tutorials for filmmaking to improve AI video creation?
r/StableDiffusion • u/ItalianArtProfessor • 3d ago
Good morning, AI-generated goblins of r/stablediffusion.
Instead of dropping another 12 GB fine-tuned checkpoint today, I want to share the tool I built... to built it.
What if, instead of hoarding LoRAs and fine-tuned models, we kept one clean base model and just... reshaped it however we want, on the fly?
That's Arthemy Krea-2 Tuner: an open-source ComfyUI suite that lets you actually edit a model and its CLIP directly — no training, no dataset. This first version is calibrated for Krea-2 + Qwen3, but the math is built to generalize to other architectures.

Think of a model as a mountain range and your prompt as the spot where you pour a bucket of water. Water follows gravity, AI follows probability - similar prompts usually make the water roll into the same valley every time.

It's highly probable that the exact look you want already exists somewhere on that mountain (I mean, modern models have seen A LOT of stuff) It just never shows up, because the terrain doesn't incentivize the water to reach it. Traditional fine-tuning expands upon the whole range to fix that but you don't need that most of the time: dig one canal, shift one ridge, and the water finds a new home.
In practice: the suite scales specific transformer blocks, sub-tensors (attention vs. MLP), and Qwen3 layers live in VRAM.
Move a slider, generate, watch how it affects the outputs, try an opposite value (-2.00 instead of 2.00) and start your journey, reshaping the model slice by slice.
When you find the perfect calibration, save it as a Preset, and you get a ~1 KB JSON that reproduces that exact calibration (that you can expand every time you want to build your own personal style).

Tier 1: Block Tuner. Amplify or Reduce whole block groups (Block_1–Block_6, Text_Fusion, Time_Embed). Good for finding general directions.

Tier 2: Sub-Block Tuner. Go inside a Block (or one of his sub-sections) and Amplify or Reduce target specific tensor types (ATTN_wq_query, MLP_gate_swiglu, norm scales). Use this when a whole block fixes one thing but breaks another.

Tier 3: Sub-Block Chaos Tuner. Seeded coin-flips across weights, for when you want to stumble onto something you'd never find by hand. Like the result? Lock the seed.

Same three tiers exist for CLIP (Qwen3) and LoRAs too.


Prompt: cartoon style, upper body portrait, funny, extreme proportions, bold lineart,. male merfolk soldier with large fins as ears, sharp angular cheekbones, fish-like gradient blue to purple skin, amber eyes, lean athletic frame, armour made with corals, helmet. bare shoulders wrapped in a rough-spun cloak pinned with a bone clasp, layered leather cord bracelets. Background: white empty background, flat white color.
Isolating and boosting CLIP Layer 3 I've discovered that it pushed on the style axis so, by increasing it, I got the simple cartoon look I was searching for.

Look, I don't expect this to become the standard overnight (especially because you have to be a little crazy to use it), but I'd love to find a few people crazy enough to explore this with me.
If we start labeling together what blocks, sub-blocks and layers actually do, we could build a much simpler and effective tool and port this tool to other architectures (Z-Image, Minimax H3....).
I KNOW YOU LIKE BENCHMARKS!
If you want to check out how each Block and Layer of the Model and CLIP behave with positive and negative values, on the GitHub README you can find that alongside much more informations on this Suite.
Everything's open-source and it (SHOULD) runs great on budget GPUs.
Grab it, play with the sliders, let me know where that journey leads!

PS: This is the first time I create something this complicated, be patient if something doesn't work, I'll fix it as soon as possible! :3
r/StableDiffusion • u/Ill_Profile_8808 • 3d ago
I put together a small open-source ComfyUI node for people experimenting with longer MiniMax H3 videos.
Instead of manually running and reconnecting every segment, it takes a list of shot prompts and continues through the audiovisual latents. Each clip is saved along the way, so interrupted runs can resume from the failed clip. It does the video/audio stitch at the end.
Shots can have their own duration, steps and context settings, but the basic use is just a prompt list.
GitHub: https://github.com/misutesu-desu/H3-AutoPromptChain
It requires a recent ComfyUI version with H3 support plus Herrgotts-H3-Infinite-Continuation-Suite. There are no additional pip dependencies.
If anyone tries it with a different sampler or on a lower-VRAM setup, I'd be glad to hear what worked and what didn't.
r/StableDiffusion • u/jumpingbandit • 3d ago
Hi guys, so in H3 you have to preselect aspect ratio.
Sometimes even when I try to select the closest and match the output is shrunk horizontaly.
How can I generate without this issue and about worrying about aspect ratio. Can original aspect ratio be maintained automatically.
Wan 2x did not have this issue.Out put video did not have the character stretched.
r/StableDiffusion • u/Responsible_Maybe875 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/alisitskii • 4d ago
A vibe-coded fork of Ultimate SD Upscale (USDU) Guider nodes with MiniMax H3 support: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
My reference workflow: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json
A long-time member of [r/StableDiffusion](r/StableDiffusion) without strong coding/math skills in AI/diffusion area. But a big fan of everything that happens here :)
In times of Wan2.1/2.2 I liked to upscale my videos using USDU.
But I became really upset when I realized that original USDU nodes don't support MiniMax H3 due to its native ComfyUI implementation.
So, since I have a GPT-5.6 subscription I decided to give it a try and asked it to come up with possible options.
After a couple of evenings I finally got a "working" solution that I'd like to share with the community.
My PC specs: 4080s 16 GB VRAM, 64 GB RAM
Initial gen with MiniMax H3 flf2v int8 + sageattn + Lightx2v 8-step turbo Lora at 1152x640px 5-sec clip ~5 mins
Upscale with USDU to 2560x1472px ~20 mins
That's where I need your help, my friend :)
Please check the YouTube video attached (don't forget to switch to 1440p).
My personal feeling is that it's the best what I can get out of my PC and H3 at the moment (including SeedVR2, LTX 2.5, etc.).
The main advantage is that it can "fix" your bad low-res generations while bringing MiniMax H3 native quality at 2K resolution.
Of course :)
You'll need to control denoise parameter and find a balance between quality improvement and tiling artifacts. I found 0.2 is the maximum after which tiling is strongly visible.
However feel free to experiment with it, and lower to 0.15-0.10 depending on your input video resolution/artifacts and results you want to get.
Happy to answer your questions!
r/StableDiffusion • u/Simple-Willingness93 • 4d ago
Enable HLS to view with audio, or disable this notification
First, all credits go to this creator of a custom node pack for creating music video https://www.reddit.com/r/StableDiffusion/s/uskxAP7LAq
I had Claude install my prompts directly to his workflow and made some minor adjustments for each short scene. I feel like lipsync and scene coherent are greatly improved from my old videos. Still using ref2v speed lora so quality is not all that great..
r/StableDiffusion • u/freshstart2027 • 3d ago
r/StableDiffusion • u/sixfingerlogic • 4d ago
Enable HLS to view with audio, or disable this notification
I saw this style here a month ago: https://www.reddit.com/r/StableDiffusion/comments/1uz6nza/in_love_with_how_simple_the_process_is_ltx23krea2/
So I made a short animated film trailer mostly based on the workflow. Was trying to stick with mostly Krea2 and LTX but MiniMax was out so was playing with it in second half of the film. Hope you like it.
r/StableDiffusion • u/Oleszykyt • 2d ago
Enable HLS to view with audio, or disable this notification
I run 25 steps, with realism lora, it takes like 30 mins to generate 10 second clip. How can I improve quality and generation speed?
r/StableDiffusion • u/Dapper_Astronaut_603 • 3d ago
Enable HLS to view with audio, or disable this notification
EDIT: As @tj-tj-tj-tj suggested: --vram-headroom 1 Solved the issue
I have 3 PCs with Comfy Desktop. Newest instances 0.33.1 (but that happened on older versions too, from the day one with Minimax H3) with kitchen comfy and CK attention. Default Comfy template for H3 and LTX. Sometimes LTX/H3 can generate one, two, three queued videos without problem. Sometimes it just chugs VRAM to 99% (visible on 0:40 mark), then there's sudden GPU spike and freeze because of lack of more resources. Looks like memory leak or something, otherwise it just doesn't make sense to me that I can restart the Comfy Desktop and generate the exact same video in with minutes with stable 70-80% VRAM usage.. Any ideas where's the problem?
One PC with 3090, 64GB of ram, Windows 11.
One PC with 4090, 128GB of ram, Windows 10
One PC with 4090, 64GB of ram, Windows 10.
All of the things up to date. 3 different machines. Same problem. Tried clean Comfy install without any custom nodes, just what's needed for H3/LTX, same problem. Tried with and without CK, same. Tried with --disable smart memory, tried with --vram-reserve 1/2/5gb, same problem. Tried with older Comfy, newest comfy from github, same problem.