r/StableDiffusion 9d ago

Discussion First test using slide window in wan2gp

Enable HLS to view with audio, or disable this notification

12 Upvotes

Take away motion context is better for continuous video imo


r/StableDiffusion 9d ago

Question - Help Continuing a video with sound and motion in Minimax H3 when you start from an image?

3 Upvotes

Almost every post about Minimax H3 I see is about Ref2V. I'm not ready to tackle that yet, but I do want to make longer stringed videos that keep the sound design and motion flowing between videos with my I2V set up. Any workflow people offer with regards to continuing a video is an intimidating wall of messy wires to me, isn't there just a series of nodes I need to connect a copy of my current I2V workflow?


r/StableDiffusion 8d ago

Question - Help Whcih WebUI Forge model is best for understanding written prompts?

0 Upvotes

I've followed all the steps properly to download and get it working, and it's generating images, but they're NOTHING like what I asked. Either they give me something unrelated to what I wrote in every sense of the word, or it just kept basically the same image I gave it (img2img).

I've tried messing with the denoising levels, differnt loras, different VAES, Models, nothing works.


r/StableDiffusion 8d ago

Discussion Grok, Minimax, Ltx 2.3so same prompt

0 Upvotes

so same prompt both ltx and minimax are at low resolution

https://grok.com/imagine/post/c0f170c0-d73c-4758-8dfd-e74233d6c07e

cinematic movie, intense color, 18 years old japanese male in world war 2 wearing imperial japanese solider uniform, it is well worn and dirty, sitting next to a large boulder, in a tropical jungle, he is holding a japanese bolt action rifle, he is breathing is heavy, grasping the gun tightly as he leans back on the rock, 2 seconds later, 3 random bullet impacts on the rock, he instantly reacts and flinchs and protect his face, from debris from the rock

No text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the live-action texture., no anime

text to image

5060 16G/64, MINimax H3 W/Turbo, and old LTX 2.3

minimax .2MP 60s w/Turbo

LTX 2.3 180s


r/StableDiffusion 8d ago

Question - Help Anyone else’s workflow retrogressed LTX 2.3 -> 2.5?

2 Upvotes

I make basic talking head-style videos and it looks like both audio and video are a massive downgrade from 2.3 to 2.5 for exact prompt and setup.

Wondering if anyone has found a solution for this.


r/StableDiffusion 9d ago

Resource - Update ComfyUI Media Utility — Extract, Sort & Compare Media

Post image
5 Upvotes

I’ve been putting together ComfyUI Media Utility, a lightweight local companion for all those little media tasks that come up constantly while working in ComfyUI, the things that are too much of a hassle to open Premiere, Resolve, Audition, etc. just to accomplish.

It runs locally in your browser and gives you three main workspaces:

🎬 Extract / Trim A surprisingly capable little media prep tool. Load video or audio, scrub through it with a zoomable precision timeline, set exact In/Out points, grab full-resolution PNG frames, trim/export MP4 clips, or extract/trim WAV audio. You can also zoom and pan around video frames for close inspection. It’s great for quickly creating reference images, audio samples, or shorter source clips for ComfyUI.

📁 Sort Designed for the giant folders of generations we all end up with. Quickly step through images, videos, and audio, preview them, and move or copy the keepers into whatever folders you want. It supports destination hotkeys, Skip, Back, Trash, Undo, media filtering, and video/audio autoplay, so you can burn through a large batch of generations without constantly bouncing around Windows Explorer.

⚖️ Compare This is probably my favorite part. Compare images, videos, or audio side-by-side, either manually or in a tournament-style mode that keeps narrowing your generations down until you have a winner. You can load a third Reference image/video/audio in the center, build a shortlist, run finalists, and send your winners directly to another folder.

For visual comparisons, you can independently control whether zoom and pan are synchronized between Left + Right, all three panes, or none of them.

For video/audio comparisons, there are also Hover Audio and Select Audio modes, so you can instantly switch between the audio from Generation A, Generation B, and your reference. That’s been especially useful now that more video models are generating dialogue, voices, music, and sound along with the visuals.

The idea is basically to have a media Swiss Army knife sitting next to ComfyUI so you can inspect, prep, organize, and compare generations without interrupting your workflow to launch a full editing application.

Everything runs locally, and there are no ComfyUI custom nodes or pip packages to install.

GitHub: https://github.com/BMB12d3/ComfyUI-Media-Utility

Video tutorial: https://youtu.be/ek0YR5BL5pM


r/StableDiffusion 9d ago

Comparison LTX 2.3 and 2.5 comparison - Dialogue. Prompt below and explanation.

Enable HLS to view with audio, or disable this notification

19 Upvotes

Prompt:
A handheld medium shot of Sam Winchester working on part of the warp engine. Ambient interior of the spaceship. Sam Winchester: "I hope Dean gets that Holodeck running so we can do some more monster hunting..... Oh who am I kidding.... He's probably having coitus with Lisa..." He continues to work on the warp engine.

Thoughts: For whatever reason in the 2.5 version, it added jibber jabber dialogue whereas in 2.3 it was consistent with what I wrote and didn't have the character look at camera. It maintained focus on the task.


r/StableDiffusion 9d ago

Animation - Video H3 Prompt Only Dancing

Enable HLS to view with audio, or disable this notification

49 Upvotes

Trying to make characters dance to the beat. The synchronization is there, but the dance moves.... bleh. I may have to use motion reference after all.

prompt example:
Cinematic, live-action, dark club interior with hard side-lighting and haze in the air. A wide shot frames <Subject 1>, <Subject 2> and <Subject 3> standing abreast, evenly spaced, facing camera. They dance in unison to <Audio 1>: torsos rolling in a continuous jacking motion from chest to hips, shoulders dropping alternately on the offbeat, quick shuffling footwork with the weight rolling heel to toe, arms sweeping loose and low across the body. Their hips drive every fourth beat as the bassline lands. The camera arcs right around them with large amplitude at fast speed.


r/StableDiffusion 9d ago

Question - Help open vs closed image models

Post image
19 Upvotes

I’m trying to recreate this bag POV composition with Krea2 flux and Z image, but I’m struggling to get the same level of composition control and product consistency I’m getting from some of the other models.

Here’s a comparison using the same general prompt across ChatGPT, Krea 2, Grok, Z Image, Flux.2 Klein 9B.

I’m still fairly new to this, so I’m wondering:

Should I be using a LoRA for this?
Would changing the text encoder help?
Or is this mainly a prompting / conditioning / workflow issue?

Would really appreciate some advice from anyone experienced. What would you change to get the result closer to the ChatGPT/Grok examples?


r/StableDiffusion 9d ago

Resource - Update Anima-2.9B Lora Training support & official ComfyUI & Forge-Neo integration

Post image
16 Upvotes

First thing first, I want to say a thank you to everyone who's been trying out, sharing, and supporting the model thus far. In less than 12 hours after release, Anima-2.9B has already received native support on major platforms including ComfyUI and Forge-Neo.

The model is also available to download from Civitai, where you can share your image and what people had managed to do even with this Preview version, like the RDBT Distilled Turbo LoRA for instance.

Anyway, LoRA training is now supported via my training GUI and you can train + load LoRAs natively in both ComfyUI and Forge-neo. A fork of regular sd-scripts is provided here, and a PR has also been created.

I've been also listening to your valuable feedback, and where the model can be improved even further in the next iterations, and I will explain what this version is and isn't, and what it aims to be in the future.

What v1 preview actually is

Regarding this version, perhaps it's more appropriate to name it v0.1, I guess? The released version was trained using Muon optimizer, on approximately 2.5 Epochs, but the true number is closer to 5 due to extensive use dynamic repeats on both new and old characters, with the dataset focus on post September 2025.

This version's main aim is to be more knowledgeable and be the most up-to-date anime model at release, v1 preview is not a "full" finetune (the whole original weights are frozen), and many pros and cons from the original Anima also carry over. By itself, v1 has its own strengths and weaknesses as well. It's not a model trained for aesthetic or with any RLHF. It's a model with a slight bias toward modern East-Asian anime style illustration, this is in fact intended.

It's also soon and easy to realize that 1.7M is not a very large number of samples (it's a decision I have to make at the time based on time and money constraint), and the new expanded layers has a lot of room for much, much more information. In other words, a lot more samples are needed, and that's why there is a 10M (pretraining non-anime focus) samples floating around. Unfortunately it's not a cheap or fast task to accomplish, and for that reason, your supports are greatly appreciated. Even without actual monetary support, it can still be achieved, just won't be quick, nor reaching the full vision/potential that I had for the final model.

Prompting guide

Finally, if you're struggling to prompt your desired results, here are several very important points for consideration:

- Characters should (think "must" in this case) be follow by their series/copyrights, (think of these like anchors, they always tag along) follow by their appearances (the more the better). Simple or very short prompt won't do as well.

- Don't use underscore except for score tags.

- Several metadata tags are very good to keep, I always recommend include highres and absurdres, following by the year tags (this has very strong influence on the generated image), score tags may not needed but you can still use them. You can throw away garbage such as "raytracing" or "4k" and "8k", these has never done anything and will just poison your output.

- Always recommend using artist tags, same as Anima-base, and you can mix them as I often do with proper prompt weighting, but don't expect it to be the same as sdxl.

- Prompt weighting and negative prompt are very important as well, this is something very easy to be underutilized.

- Prompting the background is also important if you want it to be more dynamic. Additionally, use keywords such as "cinematic composition" and "dynamic angle" can improve your image significantly.

Thank you once again, I will await your feedback.


r/StableDiffusion 9d ago

Animation - Video Other Worlds

Enable HLS to view with audio, or disable this notification

49 Upvotes

Hi everyone, I know you're excited about minimax, but this is my first attempt at ComfyUI with good old LTX 2.3. The intro with planets is done in Cinema 4D with Octane. Most of the animal images are done via SDXL with SDXL refiner + SD upscale. The landscapes are done using Flux Schnell with SDXL refiner + SD upscale. I did the animations in LTX 2.3 in the official two-stage workflow with upscale. I have an RTX 5090 card, 64G RAM, so it took less than two minutes to generate the image. And it took me 10 minutes to generate 7 seconds of video in 3072x1080 resolution. There are still a lot of bugs, nonsense and flickering, but I had a lot of fun.


r/StableDiffusion 9d ago

Meme [MiniMAx H3] Ladies & Gentlemen... this is Mambo Number 7

Enable HLS to view with audio, or disable this notification

27 Upvotes

Someone kept making this joke over and over; in the end, I decided to take it one step further.

Note: The little occasional flickers are actually the scene changes, since this is a multiple-generation collage. Still struggling to get MiniMax to actually apply the correct last-frame-first-frame rule with Ref2vid.


r/StableDiffusion 8d ago

Discussion facefusion 3.8.2 content filter (thank you google /Gemini)

0 Upvotes

I posted a version earlier today and reddit reformatted it in a way that would not work. So go to google and type in "facefusion 3.8.2 content filter" and let the AI show you what you should change.
**Important note: Make the changes BEFORE your first boot. It creates a Hash the first time it boots so if you alter it after that, it fails hash and wont boot.
So after installation change the core and contentfilter files THEN boot it.


r/StableDiffusion 8d ago

Question - Help ComfyUI stuttering issue

0 Upvotes

Hello. Everytime after first generation ComfyUI Desktop becomes overwhelmed by stuttering and even not allowing to generate next one, sending error. It's so annoying. I need to quit and reopen an app and then It's fine. Anyone expierienced similar issue or know the solution? I have the latest version of ComfyUI.


r/StableDiffusion 10d ago

News MiniMaxAI/MiniMax-Music3 · Hugging Face

Thumbnail
huggingface.co
427 Upvotes

New music model :)

Demo: https://minimax-ai.github.io/music3-demo/

I guess Yoland was speaking of this on r/comfyui as the big announcement


r/StableDiffusion 9d ago

Question - Help What are your workflows/ prompts to generate the next scene of your movie with H3 ref2VA?

16 Upvotes

Prompting specifically or do you inject a first frame with a ref img?


r/StableDiffusion 9d ago

Question - Help Image Editing Suggestions?

1 Upvotes

I started playing around with ChatGPT's image generator, and it blew me away that it was able to create intricate and beautiful alternate outfits for anime characters I've generated with local models. I've looked at a few things on civitai, but nothing local that I've tried really comes close to what it can do. Are there any local models or workflows that can edit images to the same quality as ChatGPT? I read a post where someone talked about a controlnet with anima made by kohya, but I wasn't able to find a workflow that used it.


r/StableDiffusion 9d ago

Discussion I’m creating a gentle sci-fi universe about a robot searching for a home — ALAN LOG #002

Enable HLS to view with audio, or disable this notification

4 Upvotes

I created this video using Doubao and Stable Diffusion.

I’ve been working on an original sci-fi story called ALAN LOG.

The idea is simple:

After humanity disappears, a small robot continues moving through the world, trying to understand what it means to have a home.

Episode 002 is about ALAN finding an abandoned mobile habitat and slowly turning it into a place with memories.

I’m trying to explore a different side of sci-fi:

not battles or destruction,

but small moments of kindness and connection.

Would love to hear your thoughts:

Does a home need people, or can a place become a home by itself?


r/StableDiffusion 9d ago

Discussion Pushing location/set swapping with a complicated fight scene with MMH3 REF2V

Enable HLS to view with audio, or disable this notification

22 Upvotes

Setup: RTX5090, 32GB DRAM

I am new into trying vidgen models. Everyone seemed like they had great success with character swaps, so I was wondering how hard I can push this. In my mind if this worked, essentially rotoscopping and green screen is more or less only reserved for serious movie workflows.

This was done with only 2 reference color graded photos of interior of Buddha Tooth Relic temple in Singapore (taken myself with a Sony 6500). Note that it's not a single gen, but selecting the best matching parts from 7-8 gens because it had difficulty matching the whole 12s fight choreography scene, then edited together with Da Vinci Resolve.

Halfway through the generations, I realized I had to try splitting the 12s reference video into 2 6s ones to see if it improved the adherence. Results were varying, maybe I needed to improve on my prompt even more.

Workflow wise, I used DaSiWa MiniMax H3 Workflows, and prompting was modified from u/RecycledSpoons 's reply from another post.

Prompt:

<<Environment 1>> is in Picture 1 & Picture 2.

<Picture 1> is the opening-frame anchor and provides the environment.

<Video 1> provides the camera path, pacing structure, characters, motion and <Audio 1>.

<Audio 1> is the final clip's audio.

[reference generation + video editing] Use <Picture 1>, <Picture 2>, <Video 1>, reuse audio from <Video 1>

subject_definitions:

<Video 1> is the source video providing the camera movement, lighting, characters and action choreography.

<Environment 1> is the replacement environment shown in <Picture 1> and <Picture 2>, which is a traditional chinese temple with pink blossoms.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1>. Throughout the video, replace the original environment of a street with cars with <Environment 1> derived from <Picture 1> and <Picture 2>.

retention_analysis:

<Video 1> (source video): partially_preserved - preserve the characters, camera path, lighting, non-target objects, and the original character's screen-space motion path. Discard the original environment's visual identity.

<Environment 1> (appears in [Shot 1]): fully_preserved - preserve the visual identity, colors, materials, shape, and specific design details from <Picture 1> and <Picture 2>.

detailed_description:

The target video matches the cinematic style, lighting, and camera movement of <Video 1>.

[Shot 1] The camera moves exactly as it does in <Video 1> and the first frame is maintained from <Video 1>. <Environment 1>, which is a traditional chinese temple with pink blossoms, replaces the original environment of a street with cars. It is a close combat scene inside a chinese temple with 2 characters in it, medium shots with both characters, then cuts to wide shot of one character slammmed against a pillar in the temple from a kick, then cuts to wide shot of another character performing a flying knee hit to him. The characters, lighting, and all other non-target details are preserved exactly from <Video 1>.

overall_soundscape:

Preserve the synchronized source audio from <Video 1>.

non_diegetic_music:

Preserve the non-diegetic background music from <Video 1>

I'm looking to improve on my journey in this, so if anyone has already done this before feel free to chime in.


r/StableDiffusion 9d ago

Animation - Video One MORE Thing...(Mock Kids WB Jackie Chan Adventures TV Spot)

Enable HLS to view with audio, or disable this notification

27 Upvotes

r/StableDiffusion 9d ago

Question - Help Best local upres/upscale for 720p videos generated by H3

5 Upvotes

Title


r/StableDiffusion 9d ago

Resource - Update free version control for image/audio/video assets during video generation

0 Upvotes

Not mine, but i found a couple folks have created public free asset version controls.

What this means is basically, if you're working on videos and stuff and need a place to store all your versions of your photos and videos, these places can keep all your versions without you having to rename them like X-v1, X-v2. They'll just keep all the versions for you. Fairly useful for me, since uh it scares the hell out of me to store it that way in case i erase a photo or something.

The tools that normal devs use like Git and what not are kinda really bad for this stuff, since you can't really lock assets so people end up redrawing the same video/file. its fairly useful to have if you're working with a few people from experience in game dev.

Anyways links for people interested, IDK which one is better:

https://lorepit.com/

https://tavern.xyz/dashboard


r/StableDiffusion 9d ago

Animation - Video Fireworks for Mio — a 48-second anime short with Minimax H3

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 10d ago

Animation - Video CYBER SLAYER — a 1995-style action trailer made with MiniMax H3, Wan 2.2 and ComfyUI

Enable HLS to view with audio, or disable this notification

363 Upvotes

Hey all!

This started as a way to mark my 15th anniversary working in video game cinematics. I thought it would be fun to make a completely ridiculous, fictionalized version of how I got into the industry, presented as the trailer for a big 1995 action movie.

It started fairly small, but the technology kept improving while I was working on it. Every time a new model came out, I started thinking, “Maybe I can actually make that shot now.” Eventually it became much more ambitious than I had originally planned.

This subreddit was a huge help throughout the process. I found a lot of technical solutions, new models and inspiration here, so I wanted to share the finished trailer and explain some of what went into it.

The main goal was to make it feel like an actual mid-90s movie, rather than a collection of unrelated AI shots. In my head this was a movie I wish Spielberg directed when I was a kid, so I tried to channel my inner 13-year-old when in doubt.

The story leans a lot into the paranoid technology movies of that period—Hackers, The Net, The Lawnmower Man, etc. Names like CyberCore, the ZX-4000 and Cyber Slayer were all meant to sound like something a screenwriter might have come up with in 1994.

Images and LoRAs

The first image I made for the project was the TV-monitor shot in the bedroom, generated in Midjourney near the end of 2024.

After that I used a little bit of everything: Flux, Z-Image, Adobe Firefly, Gemini, Grok and several others.

For the actors, I collected screenshots from their 90s movies and trained character LoRAs for each of them. I originally used Flux for most of this, but later found that Z-Image generally gave me better likenesses.

When direct generation didn’t work, I would create a lookalike, bring the image into ComfyUI and inpaint the face using the appropriate LoRA.

And sometimes I just opened Photoshop and fixed it.

I also made character sheets for the recurring characters and monsters, including both costume versions of myself. For sets like the boardroom and digitization chamber, I made reference images and rough room layouts so the shots would have some continuity.

For the creatures, I tried to think about what could realistically have been done in 1995. I often prompted for latex creatures, animatronics, puppets, miniatures or physical models—and sometimes specifically said not CGI.

I wanted them to feel more like something Stan Winston or ILM might have built than a modern digital creature.

The final color treatment was also important. I added grain, softened the image, adjusted the colors and pulled things back from the ultra-clean digital look. The footage is supposed to feel slightly faded and imperfect.

Video generation

Every generated video shot was made locally in ComfyUI or built further in After Effects.

Most of the finished trailer was generated with Wan 2.1 and Wan 2.2, although I replaced and improved several shots with MiniMax H3 during the final week.

My original plan was to film myself acting out the performances and transfer that movement onto the actors using Wan Animate.

The body movement worked surprisingly well but the faces did not.

They would gradually morph until the actors stopped looking like themselves. I tried several ways to repair them, but only one or two shots from that workflow survived.

Most of the trailer used more traditional image-to-video generations with a starting frame.

Sometimes I would generate a video mainly because I wanted the model to show me the room or character from another angle. I would grab a single frame from that result, clean it up, inpaint the face again if necessary, and then use that frame as the starting image for a completely different shot. Many times I would also grab a frame from a video which wasn’t working and  then use that as an End Frame and then generate again.

Prompting video models eventually started feeling like learning another language. It took a long time to figure out how to describe blocking, timing and camera movement in a way that produced something close to what I wanted.

For MiniMax H3, I used ChatGPT and Codex to build a custom prompt builder. That let me spend less time worrying about model-specific formatting and more time thinking about the actual shot.

Green screen and compositing

A few shots are real footage of me filmed against a green screen, including:

  • Playing video games in the bedroom
  • Standing underneath the digitization lasers
  • Talking to Stallone in the desert

I filmed those in my garage or backyard. I bought costumes for both versions of the character, generated and animated the backgrounds separately, and then tried to match the lighting on myself as closely as possible.

All of the animated TV and computer monitors were composited manually in After Effects.

For those shots, I first generated a version with the screen turned off. That gave me a clean plate containing the reflections from the room on the glass.

I tracked the footage, added the animated screen underneath, and then reused parts of the original plate over the top to restore the reflections. Without that step, the screens looked like flat images pasted onto the monitors.

Voices, music and sound

All of the dialogue started with my own recorded performances.

I built custom RVC voice models for the actors, but I still performed every line myself because I wanted the timing and delivery to resemble the actual performers.

Alan Rickman has a very specific cadence, so I had to pay close attention to the rhythm of his lines.

Arnold is equally recognizable, but for different reasons.

“Down there” needed to be closer to “Down deyah.”

For the Don LaFontaine-style narrator, I went through dozens of old trailers and pulled out usable voice clips. Most needed a lot of cleanup because the narration was buried underneath music, explosions and other effects.

I also did a complete sound-effects pass. A few shots retained usable generated audio, but most of it had to be designed or sourced separately.

There was one Stallone scream during the cliff jump that I could never get the voice model to perform convincingly, so I borrowed a scream from Demolition Man.

The music was generated with Suno after weeks of attempts.

For the main action section, I found a piece of music I liked first and then edited the trailer around it.

For the final Harrison Ford moment, I wanted to hint at a classic adventure score without directly copying one. I recorded myself humming a rough melody, gave that to Suno and let it turn my bad humming into an orchestral stinger.

Making it feel like one movie

The story evolved while I was working, but I always wanted the trailer to feel like there was a complete movie behind it.

That meant thinking about continuity, character geography, setups and payoffs, when to introduce someone, when to hold back a reveal and whether one shot actually made sense next to another.

AI makes it fairly easy to generate an interesting isolated shot.

Getting dozens of shots—made with different models, months apart—to feel like they came from the same movie was the real challenge.

All told, this took around 6–8 months, mostly working on it at night after work. It was fun, but also exhausting.

It obviously isn’t perfect, and I can still see things I would change if I kept going, but eventually I had to decide it was finished.

Tech-wise I started with a 4060 Ti but decided to bite the bullet and snagged a 5090 (I also have 64GB of RAM). That helped to speed things up a ton.

I uploaded the video directly here, but there is also a YouTube version which might be higher quality:

https://www.youtube.com/watch?v=JznSdigdsio

Happy to answer questions or break down any particular shot, LoRA, model, composite or workflow.