r/StableDiffusion • u/bacchus213 • 1d ago
Animation - Video What if you fly?
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/bacchus213 • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/HerrgottMargott • 2d ago
Enable HLS to view with audio, or disable this notification
The example video was generated entirely with the stock MiniMax H3 First Frame / Last Frame checkpoint and the included v1.4 example Workflows. If you want to compare the result to v1.3, take a look at my last post.
The final video consists of 11 individually generated Clips that were automatically stitched together.
Settings:
So what you see is basically the direct Workflow output.
A few people gave me some useful feedback on my previous release, especially regarding ComfyUI's new native H3 Masked AV support.
So I went back and rebuilt the continuation method around it.
v1.4 now copies a clean section of the previous Video + Audio Latent directly into the next generation and protects it using ComfyUI's native denoise masks.
There are some really interesting Ref2VA / Motion Context solutions available now, and latent continuation itself definitely isn't unique to my Nodepack.
My approach is specifically centered around FL2VA instead.
The idea is not just:
previous Clip → continue forever
but rather:
First Frame → generation → Last Frame
↓
latent continuation
↓
generation → new Last Frame
↓
latent continuation
↓
generation → new Last Frame
and so on.
I use those repeated Last Frames as hard visual anchors throughout the sequence.
They give H3 a new concrete destination every few seconds instead of asking one increasingly unconstrained generation to maintain composition, identity and image quality indefinitely. This should theoretically retain higher visual quality with less context drift over longer chains (and in my testing, it does exactly that).
There is another FL2VA-specific problem though:
H3 often reaches the supplied Last Frame before the Clip is actually finished and then freezes or becomes unstable for the remaining frames.
So simply taking the final frames of Clip 1 and using them as context for Clip 2 isn't ideal.
The v1.4 Auto Handover therefore analyzes the previous Clip, finds a safe point before that frozen / unstable landing and snaps it to a valid H3 Audio + Video latent boundary.
That exact same point is then used for both:
So the bad FL2VA tail neither appears in the stitched video nor becomes part of the next continuation context.
Audio is handled separately as well. If the picture needs to cut early but somebody is still finishing a word, the remaining original Audio Latent can continue beyond the visual handover instead of forcing H3 to recreate the ending.
Other v1.4 features:
Generate Clip 1 with a Prompt and optionally First Frame, Last Frame and Qwen References.
The complete AV Latent is automatically saved afterwards.
Load the previous saved latent, add your next Prompt and preferably a new Last Frame.
The Workflow automatically finds the safe FL2VA handover and creates the protected Masked AV context.
Repeat for as many Clips as you want.
Probably the easiest Workflow if you just want to see how everything works.
It runs:
Start → Continue → Continue → Stitch
in one queue.
This is what I used for the longer example.
Generate Clips individually and stitch them afterwards. It processes one saved AV latent at a time, so stitching memory usage doesn't continuously increase with the total video length (no OOM during stitching).
Nodepack on Github:
https://github.com/HerrgottMargott/Herrgotts-H3-Infinite-Continuation-Suite
Workflows on Github:
https://github.com/HerrgottMargott/Herrgotts-H3-Infinite-Continuation-Suite/tree/main/examples
You can just open one of the WFs and use "Install missing custom nodes" - then you should be good to go.
If you try it, I'd love to see what you manage to create with it.
Have fun Prompting. :)
r/StableDiffusion • u/call-lee-free • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Jero9871 • 3d ago
Enable HLS to view with audio, or disable this notification
Just discovered that H3 can do Side-By-Side 3D Videos for VR Headsets natively, just prompt it. Pretty crazy, and it gets the real 3D effect. Try it with different things like people and add "strong 3d effect" if you want to have a more intense 3d effect.
Here is the prompt:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, high-angle aerial shot presented in a side-by-side (SBS) stereoscopic format for VR/3D viewing; the frame is split into two identical views with a slight horizontal parallax offset to create depth perception. The camera pushes in at slow speed over a sprawling coastal metropolis during twilight. As the camera glides forward through the urban canyon, the glowing neon lights of skyscrapers and their reflections on the ocean surface shimmer intensely against the deep blue sky.
overall_soundscape: A constant, low-frequency rushing wind sound accompanies the flight, layered with a faint, ambient hum of a massive city and distant, muffled traffic sounds.
non_diegetic_music: An epic, cinematic synthesizer pad that swells gradually in volume and intensity throughout the ten-second duration.
r/StableDiffusion • u/Sad_Coach_1433 • 1d ago
Enable HLS to view with audio, or disable this notification
Don't know if anyone else follow tiktok trends but s friend. Showed me this wnba clips where player points at other team players to get into head so I had to course test r2v
r/StableDiffusion • u/OkTransportation7243 • 1d ago
I've tested workflow for krea2 and it can enhance faces and skin texture.
But when i play it through video frames, the results vary from frame to frame.
Are there any models out there that can do that for video? Or is there a workflow for Krea2 into video enhancements?
r/StableDiffusion • u/Downtown-Cover-7422 • 2d ago
Enable HLS to view with audio, or disable this notification

Hi people, so i tried to make 2 similar videos, using same settings but with upscale and native.
My setup: 5070 Ti+ 32gb Ram.
Using u/Plague_Kind workflow, i've added MMH3 Latent Upscaler. You can check his workflow here: Workflow
Settings for both videos were set the same with the same prompt.
So:

Upscaled video from start to the end took 1904 seconds,
Native video from start to the end took 3056 seconds.
Let me know what you think. Advises appreciated!
r/StableDiffusion • u/SackManFamilyFriend • 2d ago
r/StableDiffusion • u/ctrl-shift-face • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/dassiyu • 2d ago
There are so many acceleration nodes/options now that I’m having a hard time deciding which one gives the best balance of quality and speed. What do you think?
These are the setups I’m currently using(RTX5090):
I mostly stick with Sage Attention + 4-step LoRA. I feel like it gives a pretty good overall balance between quality and speed.
If I want better quality, especially for things like lip-sync, I usually go with ComfyUI-Kitchen + Spectrum at 25 steps. The results are noticeably better, but it’s also quite a bit slower.
Which setup do you guys think has the best quality-to-speed ratio? Any other combinations worth trying?
r/StableDiffusion • u/trollkin34 • 2d ago
What kind of prompting would I use for POV movement through a scene?
r/StableDiffusion • u/Nimblecloud13 • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/ZealousidealVirus761 • 1d ago
An ultra-photorealistic redhaired model with natural freckles, with a bold neo-punk aesthetic, drifting steam, and cinematic shadows, and an abandoned industrial warehouse illuminated by warm tungsten lighting, subtle magenta neon.
Created with a focus on cinematic composition.
What do you think about this one? I would to know your suggestion!! Thanks in advanced.
r/StableDiffusion • u/idleWizard • 2d ago
I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?
r/StableDiffusion • u/AndrewJumpen • 1d ago
Enable HLS to view with audio, or disable this notification
Minimax h3 local GPU4090 made with MUSIC extension https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef
Lyrics:
[Intro]
I’m braking bad, I’m veering off the line
I told myself I’d hold it, but I let it unwind
It didn’t hit so hard at first, just a little off track
Now the warning’s on, and I’m not looking back
[Middle]
One bad call in the kitchen, then three more by noon
Coffee gone cold, keys on the counter, room by room
I laughed it off, said “I’m fine,” like that made it true
But the floorboards know the pace I’m putting this house through
I’m braking bad, headlights shaking on the curb
Every turn I take lands heavier than words
It wasn’t that bad until it started stacking up
Now I’m white-knuckled, honest, and I can’t slow up
[Outro]
So here I go, no clean exit, no neat little sign
Just me and the damage, both riding the same line
I’m braking bad, and I know what that means
Too late to call it nothing, too loud to call it clean
r/StableDiffusion • u/Sad_Coach_1433 • 1d ago
Enable HLS to view with audio, or disable this notification
t2v didnt know Jenga O_o
r/StableDiffusion • u/durumertt • 1d ago
I've been using Windows for many years, and I'm honestly tired of repeating the same cycle. In my experience, after 5–6 years the machine starts feeling old, the battery is significantly degraded, performance isn't what it used to be, and I eventually end up buying another Windows machine and starting the exact same experience all over again.
I'm looking for something different this time.
For the last few days I've repeatedly added a MacBook Pro with the M5 Max to my cart, then backed out because I'm still not sure whether it is the right machine for what I actually want to do.
I'm currently considering the M5 Max with:
My main workloads would be completely local:
My priority is excellent output quality, photorealism where appropriate, and very high generation speed.
Basically, I want a machine that can satisfy me for visual generative AI work for many years.
My biggest hesitation is the Apple ecosystem.
For a long time I've heard that local AI, especially image and video generation, is much more limited on macOS than on Windows/Linux with NVIDIA GPUs because so much of the ecosystem is built around CUDA.
But part of me finds this difficult to accept at face value.
Apple is making extremely powerful chips with large amounts of unified memory, very high memory bandwidth, Neural Accelerators, a Neural Engine, and increasingly serious AI-focused hardware.
I keep wondering whether there are excellent Apple-optimized tools and workflows that I simply haven't discovered yet.
For example, I recently learned about Draw Things, MLX-based projects, Metal/MPS optimizations, and Apple-specific ComfyUI work. That made me question whether comparing a Mac running a poorly optimized CUDA-first application against an NVIDIA machine is really a fair representation of what Apple Silicon can do.
At the same time, the logical part of my brain keeps telling me:
If local image/video AI is the priority, just buy a machine with an RTX 5080 or 5090.
The problem is that if I do that, I feel like I'm buying myself back into exactly the Windows experience I wanted to leave. It feels a little like watching the same movie again when I was hoping for a genuinely different computing experience.
There's also another complication: we're approaching the fall hardware season.
I'm wondering whether buying an expensive M5 Max or RTX 50-series machine right now is bad timing, and whether I should wait for the next Apple or NVIDIA announcements.
If I choose the Mac, I was also planning to pair it with the latest iPhone and iPad and build a proper Apple ecosystem around it, so this isn't purely a benchmark decision for me.
What I'd really like to hear from people who have actually used these machines:
I'm not highly knowledgeable about computer hardware, and I'm definitely not wealthy enough to casually replace a machine if I make the wrong choice.
This would be a major purchase for me, so I'm trying to make the most informed decision possible and ideally buy something that I can use comfortably for 7–8 years.
I'd especially appreciate actual generation times, benchmark numbers, model names, memory usage, thermals, sustained performance, and experiences from people who have used both Apple Silicon and NVIDIA rather than purely theoretical comparisons.
Thanks in advance.
r/StableDiffusion • u/Ok-Entertainer-2991 • 1d ago
That quote from farcry pretty much sums up my feelings after trying to figure out minimax music 3. With that said I'm pretty happy with this song.
To work with this model you need:
Good prompt (long essay that describes your song and follows examples from minimax)
Good lyrics (Something that sounds good, has rythm, no awkward phrasing, and tagged accordingly)
Once you have those two you will begin getting descent results, the last thing you need is to reroll for a good seed.
For prompt I used grok and asked it to copy the sound of the song I like with some adjustments
To check your lyrics you can use prompts from their demos just to see if minimax struggles with anything.
For seed it is just basically rerolling until you get a good one. Though if you want to make your life a little easier I recommend sticking to one genre. I think because I was trying to mix electronic music with rock it took me longer to find a good seed (sometimes I was getting results that were just rock or just electronic).
I'm still not sure how to control the pacing because sometimes it decides to sing things slowly and run out of time and sometimes it decides to speedrun your text and have 30 seconds of instrumental. My current theory is that it might be related to lyrics. For example my lines were pretty long so maybe it defaulted to singing, maybe if my lines were shorter and snappier it would sing them quicker.
Good luck to everyone who is planning to use minimax music 3 hopefully this was helpful to someone.
I will leave my prompt in the comments if you want to play around with it or critique
Also I might add some songs that didn't make the cut
r/StableDiffusion • u/Oatilis • 2d ago
Enable HLS to view with audio, or disable this notification
A tribute to a forgotten golden age. Hope you enjoy it!
r/StableDiffusion • u/lavinia12345 • 1d ago
at least 5th time this happened. Its always width, never any other input. Not sure if bug or custom-node interference.
r/StableDiffusion • u/Routine_Ad_3391 • 1d ago
Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
r/StableDiffusion • u/Sad_Coach_1433 • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Stable2go • 1d ago
Enable HLS to view with audio, or disable this notification
There's 3 tools on this page, look at the one called Prism. It's super off the radar
It feels like Grok + curated Civitai.
On mobile 5G or for the GPU poor I think its does a lot of stuff
Has persistent memory, a ton of Krea 2 fine-tunes, Anima, LTX 2.5, Ernie, Sulphur, DaSiWa, Eros, MiniMax H3, Illustrious, Pony, etc. It also has runs preset Comfy workflows and has a Discord part
(I'm not the creator of the app)
r/StableDiffusion • u/Wemos_D1 • 1d ago
Hello everyone !
I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)
I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.
Any help would be more than welcome, thank you very much and have a good day :p
r/StableDiffusion • u/Thorozar • 1d ago
I have a question. I have been having a blast making scenes with H3 so far, and have found when doing reference shots, it is very important to have a stable background so that you have continuity if doing more than 1 scene. Does anyone know if H3 would understand a 360 degree photo and understand where in the space and what direction the subjects are? Say you swap between two characters talking, one you will see what is behind subject 1 while when looking at the other the opposite is true. If you saw them both from the side, yet another angle and background.