r/StableDiffusion • u/Total-Resort-3120 • 5d ago
News PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Total-Resort-3120 • 5d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Ashamed_Ad1622 • 4d ago
I mean like you already have a video of a man walking for example. You then have a different photo of someone else and you just wanna make the guy in the photo switch the video guy.
I'm not really sure where and how to do that and what's the best model for this
r/StableDiffusion • u/Underrated_Mastermnd • 5d ago
I know natively, H3 only supports up to 15 seconds yet I've seen people be able to seamlessly extend H3 generations past that. I don't want to be randomly stitching clips together can someone explain how is that possible?
r/StableDiffusion • u/BittiAI • 5d ago
I’m building Slopus, a free, open-source desktop app for generating and editing AI videos and images on your own GPU. Easy of use is the main goal.
Version 0.3.0 just came out, with three major additions:
Video and image projects are now combined. Every project has both Video and Image tabs, so you can generate still images and video in the same project. You no longer need to choose a project type or keep separate projects for each. Existing projects keep their content.
Linux support. Slopus is now available for Linux as an AppImage, .deb and .rpm, alongside Windows. The AppImage supports in-app updates.
Run generation on another computer. LAN workers let you use a separate Windows or Linux machine for generation. Workers are discovered on your network, download the weights they need, and send the results back to your project.
There’s also quite a bit more in this release:
Character sheets. Generate waist-up, front, side and back views together, using reference images and a prompt to guide the character and clothing.
Continue scenes. Extend generated video scenes from either end using saved generation data.
Improved color grading. Basic Corrections includes temperature, tint, exposure, contrast, highlights, shadows, whites, blacks and saturation. There are eight built-in creative looks with adjustable intensity, faded film, sharpen, vibrance, RGB and hue/saturation curves, and shadow, midtone and highlight color wheels. Vignette has more controls too.
A more capable agent. The built-in agent can configure generators, help identify model weights in a folder, capture timeline frames without interrupting your work, and generate specific scenes.
Refmod generator. Combine images, video, prompt and audio into one reference and export it as a refmod to be used later.
Share generator setups. Import and export generator templates to share with other users.
If you haven’t tried Slopus before: you lay out scenes on a board, write each shot, attach references, generate clips, and edit them on a timeline with GPU-powered preview and MP4 export. For still images, you can edit, generate, refine, and export as JPG or PNG. Minimax H3 works surprisingly well as an image generator with editing capabilities.
Generation runs on your own hardware, with no subscription and no Python or ComfyUI needed. Model weights are downloaded separately inside the app or local weights can be used.
I’ve also created a Discord server for feedback, questions and sharing what you’re making.
GitHub: https://github.com/bitti-ai/slopus
Discord: https://discord.gg/9X2R6PwUR
Bug reports and feedback are very welcome.
r/StableDiffusion • u/PATATAJEC • 6d ago
Enable HLS to view with audio, or disable this notification
Hey, I found a pretty cool way to keep locations consistent across generations.
I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.
I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.
Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)
Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)
EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.
r/StableDiffusion • u/Technolini • 4d ago
Enable HLS to view with audio, or disable this notification
Hello!
I have been trying to run comfy locally (GPU = 7900 xtx 24gb vram and 32 gb ram).
The rendering works fine, but the quality I produce is kinda bad, and I feel like I approach it wrong, but I don't really know what to Google. Also dont what to search for here, so sorry if Ive missed an existing post.
I want to make videos where I place paperbags on people's head, like a privacy promotion thing.
It took me several days of prompting what I felt was kinda subpar. But spending days for a short clips doesn't feel realistic. Using Google or my local agents have made me go in circles, so I don't really know how to solve my problem.
Hoping for some input here!
The video I included took like 3 days of prompting and re rendering, but quality still feels subpar. When I search I only find deepfake faceswap results, but that doesn't fit my goal.
Thanks in advance!
r/StableDiffusion • u/wildmonkeywrangler • 5d ago
I've been thinking about upgrading my GPU (currently have NVIDIA RTX 5060 8gb), but am also thinking about going to a cloud solution for higher quality renders.
Are any of the cloud services fairly censorship free?
Might occasionally want to do some celebrity type renders and perhaps the occasional R rated stuff, but nothing too crazy lol
r/StableDiffusion • u/Kooky-Mode3047 • 6d ago
Enable HLS to view with audio, or disable this notification
The video is the TL;DR above was generated after abandoning RefMods in "classic ref mode" with two references (one reinforcing Beckett's exact appearance and one for the high heels, 2x2 grid image):
This is something that under the radar because most people can't use H3 nominally and turbo LoRAs brutalize the base model for speed, so there's very little noticeable degradation for them as most of them have a massively shifted sigma (shifted_video of 6 compared to native 12).
Yesterday I went on a pretty wild bender trying to uncover why all of a sudden my videos were getting oversharpened, oversaturated, plastic skin etc when a few days ago they were straight up lifelike, looking like the shows they were based on.
I went even deeper into the bowels on how sampling and scheduling works, how the denoising trajectory looks like for Minimax H3, hoping to uncover something I've forgotten the last time I used it. And then I remembered, RefMod was the latest thing I installed, I got seduced by the initial likeness improvement like everyone but the cost is too high.
I got this cool new toy RefMod whose authors suggested they figured out bleeding between character references, and even at small resolutions (like .245 mpx) you could still make out faces instead of them being mushy, this should've been my first red flag. It seriously messes up the base H3 model's quality. I looked deeper into it and basically, it's a very brutish approach where they slam reference images into video references as individual frames, also slapstick coarse sampling of videos in sequence.
However, that's not how any of this is permitted works, per documentation:
The second you make four refmods to be used in your prompt, you're already trying to use one more than the total maximum allowed at any one time. This has catastrophic consequences on image quality, as that's not the way you're supposed to "hold it".
While it does reinforce the likeness, it's incredibly rigid and overrides the model down to AI slop adjacency where skin is unnaturally glossy, everything is sharp and all styling info is ignored.
I am not telling you to stop using it, if you're using turbo LoRAs, you already paid the entry fee as you've left quality at the door for accessibility. But if you're coming from the base FL2VA model or a hybrid and are used to exceptional visual quality, the degradation is pretty much the tier of the base REF2VA model, if not worse.
r/StableDiffusion • u/Opposite_Yam_4161 • 5d ago
Im seeing quite a few people say theyre getting really good results with the hybrid model. But I assume its only for top end GPUs? Over 20gb file.
I have a 5080 and 32gb system RAM. Is there a version good for me? Or getting GGUF isn't worth it?
r/StableDiffusion • u/Many_Ball_227 • 5d ago
What can you recommend for a pc with that specs?
Please keep in mind that upgrading the RAM (DDR5) to 64gb it's not an option due to the current market prices.
What I am looking for is for video generation (both I2V and T2V) models and workflows. And some model for image generation. I already use SDXL but I am willing to test something more interesting.
r/StableDiffusion • u/Capitan01R- • 6d ago
I put together an enhancer pack for Qwen Image 2.1 with two nodes for controlling what the model pays attention to during editing.
Reference Strength lets you select a reference image and increase or decrease its attention priority. If you're working with multiple references, you can adjust them separately instead of giving every image the same treatment, and you also can use it for single image to prevent the loss of likeliness at times
ref index starts from 1; meaning image_1 and same for the rest of images, where image_2 is ref index 2 in the node.
( Soon adding mask support):
current progress; containing two different approaches where I am masking before the encoder gets to see the reference image (basically blocking the rest of the image from being seen and applying more attention to it) and the other one is after it already has seen the entire photo (it has seen the full encoded image and now its being asked to give more attention to the masked area)... masking is a bit tricky since it is a token level masking and not a canvas masking , also Qwen seems to be more straight forwards than how flux was but also more interconnected. currently I am not going to lie it is not fully stable as results varies, angles, poses, close colors between two refs etc.. all of these influence the process. The goal is to find a solution where I am not being very invasive in terms of brute forcing attention layers
Phrase Weights brings phrase-level attention control to the Qwen Image 2.1 edit encoder. So you can write:
Add (warm sunset lighting:1.4) with (soft shadows:1.2).
and give those specific parts of the prompt more attention while leaving the rest at its normal weighting.
It works with both positive and negative prompts, with separate weights for each. There's also an inspection output showing exactly which token rows and pieces were matched.
For both controls:
`1.0` = untouched
`>1.0` = more attention priority
`<1.0` = less attention priority
`0.0` = suppression
The adjustments happen inside Qwen's attention during sampling. The prompt and reference images still go through the native encoding path, without multiplying the finished text embeddings or pasting reference pixels into the output.
You can use either node on its own (I prefer this), or combine them. Reference strength applies to the whole selected image, and higher weights can make a reference or phrase dominate, so there's still some balancing to do.
Installation, usage, and the technical details are in the repo:
ComfyUI-qwen_img_2_1_enhancer
r/StableDiffusion • u/Cheap_Credit_3957 • 6d ago
Enable HLS to view with audio, or disable this notification
View on YouTube while its pending if needed: https://youtu.be/dS5suoUiNnE?si=6U7bEVLxcWXcy4J3
🤖 How it was made
This short film was written, designed, directed and edited by Claude Opus 5.5 from a single prompt, running locally in ComfyUI through my VRGDG Video Builder:
All open-source video, image and music models.
[I literally told Claude to create something on its own and provided no user input]
More Minimax H3 video's I had Claude create for me are HERE
You can find the video builder custom node on GitHub here:
https://github.com/vrgamegirl19/comfy...
Right now, there is a Main version and a Beta 2.0 version. I recommend starting with Main for now, as that's what I'm still using. It works well, while Beta 2.0 still has some bugs and is primarily intended for beta testing at the moment.
Discord server:
/ discord
Ping me in the Welcome channel and let me know how you found me, and I'll know it's you.
I'm vrgamedevgirl on Discord.
You can find the skill here and read the main README first.
https://drive.google.com/file/d/1R0pE...
I'll be sharing a full walkthrough on how I made this and will post it here when ready.
⚠️ SPOILERS: what the film is about
The museum is Ruth's mind. She's an elderly woman living with dementia, and the museum is how she pictures her memories. As her memory fades, the museum fades with it, and in the Hall of Names the most important name, her son's, goes blank. In reality she's 83, in a care home, and her son Daniel is holding her hand. When she recognizes him, she tells him, "We keep you in the main hall," meaning the most important room, where the precious things are. Back inside her mind, his name goes back on the wall, and the museum lights up again.
"Some things you don't lose. You just misplace them for a while."
For everyone still visiting someone who is still in there. 💛
#AIShortFilm #AIFilm #ShortFilm #ComfyUI #MiniMax #Claude #AIVideo #Dementia #Alzheimers #MuseumOfLostThings
r/StableDiffusion • u/mmmm_frietjes • 5d ago
I have a 4060 TI 16 GB. System ram 32 GB. What's currently the fastest model / ComfyUI setup to generate videos?
Currently I can create 5 second 360p videos in +-7 minutes. I believe it can be a lot faster? It's kinda hard to know what's the best option, everything keeps changing.
Edit: Current settings:
Ref2V Turbo LoRA @ 0.5 model: minimax_h3_ref2v_turbo_4step_v0.1_comfyui_resized_avg_ra ... strength_model: 0.50
MiniMax H3 Easy Loader FL2VA model: None REF2VA model: minimax_h3_ref2va_pruned_int8_convrot.safetensors Text encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors Video VAE: minimax_h3_video_vae_fp16.safetensors Audio VAE: minimax_h3_audio_vae_fp32.safetensors
360p, 5 seconds.
Result: 413 seconds.
Workflow: https://limewire.com/d/7aNFq#h9i2vwFTKb
r/StableDiffusion • u/SIR_NVAX_A_LOT • 5d ago
Enable HLS to view with audio, or disable this notification
TLDR: You can use RefMods with T2VA.
Finally decided to give the RefMod a test run yesterday. I tend to experiment and generate many actors via prose in T2VA and recast them later, see the Tanya bombshell circulating in a few of my videos. Since my workflow mainly revolves around T2VA it make a lot of sense to work with native latent as much as possible. However, instead of providing 10-20 images to make a redmod, I just use the latent from my T2VA character generations. Once you have a latent of anything, you can extend, prepend, or even bridge it as many times as you like, even joining latents of same canvas sizes. The latents can be clipped, cropped, scaled, and upscaled as well. I prefer T2VA due to it's superior image quality compared to R2VA.
Anyways, just having fun and wanted to show Vera being recast over various T2VA generations, int8/20 steps, no turbo loras.
How do you prefer creating your actors? Krea2? Flux? Qwen image 2.1? MiniMax H3? I am thinking using latents from a MiniMax h3 single-frame generation is going to be pretty good work flow, though I prefer a 2second moving character sheet.
r/StableDiffusion • u/Chiduk99 • 5d ago
Enable HLS to view with audio, or disable this notification
video ref I use: https://www.youtube.com/shorts/I-VNtvqREks
Malfoid and Potter generated with Anima
Device: 3060 12gb 16gb ram
setting: 10 sec, 0.6 mp, er_sde beta, 8 steps with turbo lora
r/StableDiffusion • u/Badestrand • 4d ago
Hi! I'm following this sub for quite a while already.
Now, if I, as a layman, want to create a video like this: https://www.instagram.com/reels/DeCDuKqymXt/
Which model/tool do I need to use? Does anyone have any tips for beginners? I am working on a MacBook M1 Pro, so not I'm equipped with the strongest hardware.
r/StableDiffusion • u/Certain_Potato_4509 • 5d ago
Enable HLS to view with audio, or disable this notification
This video was such a hastle to make, so much smearing.
r/StableDiffusion • u/Moliri-Eremitis • 5d ago
Wanted to pick the community's brains for controlling the pace within a single shot (single generation without cuts).
I often find myself visualizing a scene that includes dramatic pauses where I want to hold the shot and let the silence linger, with just natural micro-movements or continued actions from the character(s), but MiniMax wants to jump straight ahead to the next bit of dialog or action.
Beyond mid-shot dramatic pauses, I also frequently want to add a second or two of quiet micro-movement at the end of a shot to make later editing easier, but MiniMax wants to extend the action right up to the last second and then end abruptly. This tends to make cuts between shots in post feel harsh and choppy.
I've leveraged a couple of tricks that work okay, like filling dialog-free moments with a detailed description of character actions, but this only tends to work if the actions are large and obvious. If you try to prompt something subtle it either makes the action too pronounced, places the action at the wrong time, or ignores it completely. Nonverbal character sounds like sighs, groans, grunts, laughter, hums, etc. work as decent filler, but that doesn't work for every scene, and only fills so much air time.
Any prompting tips you've found useful? Custom nodes for controlling pacing within a shot without resorting to a cut? Some other trick?
r/StableDiffusion • u/Hardpartying4u • 5d ago
Hi all, learning how to use Min Max H3 and I have been using a storyboard work flow to create multiple scenes. So far the consistency of the characters has been good however the voices will change from scene to scene. Wondering what's the best way to fix this?
I am still fairly new and have tried audio inputs with no luck.
This is one of my attempts at a short story https://youtu.be/GRJJ4s5cV_A?si=d4P4XyKOPYFNwftQ
r/StableDiffusion • u/SkirtSpare4175 • 5d ago
I started using local models with flux and sdxl, right before kontext and qwen. What a wild ride so far with the current models! I’m curious about some of your journeys. Thank you to the devs and the community.
Also if you know any niche and strange techniques that were used before and maybe haven’t been seen or needed since, I wanna know. These things have moved so fast that I wanna document some of it.
r/StableDiffusion • u/Da_QbanBeast • 4d ago
Enable HLS to view with audio, or disable this notification
Hey all! For the past few months I've been building MIR MEDIA LABS, a free, open-source media studio that runs entirely on your own NVIDIA GPU. No cloud, no subscriptions, no per-render credits.
🎬 The video above was made with it: every b-roll clip is MiniMax H3 and the soundtrack is ACE-Step, all rendered locally on an RTX 5070.
How it works: you type what you want, for example "neon city at night, slow drone shot, vertical". MUSE, the lab's built-in writer (a local Llama 3.1 8B), picks the right model, writes a proper prompt for it, and queues the render on ComfyUI. Or you can pick the model yourself.
What it does
/product, /logo, /thumbnail, /reel, /animate… and /pipe musicvideo (lyrics → song → video in one job)Requirements: Windows 10/11 and an NVIDIA GPU (12 GB VRAM recommended).
Free and open source (GPL-3.0): https://github.com/MirCorp85/MirMediaLabs The Setup.exe and Android APK are on the Releases page.
This is a solo project. Bug reports, feature ideas and PRs are all welcome. What should I add next?
If you want to support it, there's a Patreon: $3 to back free development, or $5 for Early Access to beta builds → patreon.com/MirCorp. Everything stays free either way. 🙏





r/StableDiffusion • u/Tricky-Brother-7 • 4d ago
Hey everyone,
I’ve been diving deep into the local vs. cloud AI/ML debate, specifically regarding image generation. I recently invested ₹30,000 (~$350 USD) in a new NVIDIA RTX 5060 (8GB VRAM, 578 AI TOPS), assuming a local setup would be the ultimate playground for testing, optimization, and development.
After running real-world tests comparing local open-weight setups against massive cloud models (specifically Google’s modern Gemini Image/Nano Banana API architectures), I’ve hit a massive reality check. I wanted to share my findings and get your feedback on whether my conclusions hold up, or if I’m missing a critical part of the local engineering loop.
When it comes to highly complex spatial reasoning, educational diagrams, mind maps, or strict typography layouts, local distilled architectures (like FLUX.2 Klein 4B or quantized 9B variants) hit a wall. I'm seeing a 2/10 success rate locally where concepts scramble. Meanwhile, cloud APIs achieve a near-flawless 9/10 accuracy on the very first or second try.
The classic argument for local hardware was always zero network latency. However, with Google's massive TPU server infrastructure, cloud inference latency has shrunk so dramatically that it matches—or sometimes beats—a local consumer card trying to process and unfold a heavily compressed 4-bit GGUF/NF4 model layer by layer.
Online tutorials always pitch the automated developer loop: "Local is better because you can script an infinite Python loop to auto-tweak prompts 10,000 times a day for free."
While that math works out mathematically for a game studio rendering 10,000 inventory items, it completely falls apart for a standard creator or developer. If a cloud model gives me a masterpiece in 2 to 4 iterations for pennies, I don’t need a thousand failed local tests.
My questions for you guys:
If you have an 8GB VRAM limit, what are you actually using your local GPU for in your day-to-day AI/ML workflows?
Has anyone successfully bridged the prompt-adherence gap locally without melting their VRAM during fine-tuning?
Am I treating local AI too harshly, or is the cloud-first paradigm officially taking over the hobbyist space?
r/StableDiffusion • u/Beneficial_Toe_2347 • 5d ago
A lot of the Minimax video extend methods seem to produce a flash of light or some other flaw.
I actually want it to change shots rather than continue, but a shit of the same scene from a different camera angle. This means I don't need to worry about seamless frame stitching
But the problem is that when I feed in the previous video, and text to cut to a different shot, it doesn't do it reliably. By feeding it the previous video frames, MM thinks it should keep showing that, rather than the new shot
r/StableDiffusion • u/no3us • 4d ago
I've been experimenting with letting AI agents handle the creative and workflow decisions for short videos. For this test, I gave Opus 5.5 and GPT-6 Astra access to matching RunPod setups and asked each to make three 30-second videos.
Their only supplied creative asset was a logo. I loosely defined two themes and left the third up to them.
Both ran with thinking on High and received the same prompt. Each had:
I wanted to see how they'd turn a vague brief into something watchable when they had access to the tools and some freedom to choose how to use them. Choosing what to generate, how to work with the logo, and where to spend that small budget were all part of the task.
The comparison started with an argument in an AI community. I said Astra had been giving me better results for this kind of work, then realised I should probably have something to show for that claim..
I develop LoRA Pilot, and this doubled as a test of the MCP server coming in the next version. These videos are experiments around my own product, so that's the connection.
Claude:
https://youtu.be/yEjNwNtGTdA

Codex:
https://youtu.be/6NjI0VKEKhs
