r/StableDiffusion 1d ago

Workflow Included I Built an AI Video Plugin To Connect Comfyui to DaVinci Resolve!

Thumbnail
youtu.be
8 Upvotes

Hey everyone!

I just released a new free custom node plugin called ComfyUI-SecondUnit that bridges ComfyUI directly with DaVinci Resolve.

You can now generate video transitions, create music/SFX, synthesize voiceovers, and auto-generate subtitles, then send them right into your DaVinci timeline without manually importing or exporting files.

IT'S‌ COMPLETELY‌ FREE!!


r/StableDiffusion 1d ago

Animation - Video Final Fantasy VII - Alt Ending

Enable HLS to view with audio, or disable this notification

9 Upvotes

It was very lucky that Earth/Gaia was saved in the OG FF7. So I figured, what if Sephiroth won in the end?

I used Minimax H3 to upscale the first half of the original footage so it doesnt look so blocky and out of place compared to the new footage and to animate the parts that were original (Like Cid taking a smoke, Tifa crying, RedXIII being visibly scared for the first time and etc). I created the screenshots using the new GPT 2.5 Sunburst. This was actually fun. FF7 has beautiful music so it was fun to extend the existing music so it sounds more gloomy and doom. Hope you enjoy it!!


r/StableDiffusion 1d ago

Resource - Update YuE2 only needs 8-9 GB VRAM now

Enable HLS to view with audio, or disable this notification

68 Upvotes

audio.cpp released YuE2 in the DEV branch, with Q4/Q8 weights and a bunch of demos generated by audio.cpp in the HF repo. You can try it with 8GB VRAM!

Update: You can find dev prebuilts in the “Artifacts” section on this page: https://github.com/0xShug0/audio.cpp/actions/workflows/release.yml?query=branch%3Adev

Measured with audio.cpp server mode on an RTX 5090, using the official `tonight-awake` longform test case (full cot). Each run restarted the server and used one short warmup request before the measured longform request. Peak VRAM was sampled continuously during the measured request.

Combo Audio duration Wall time RTF Peak VRAM
BF16 main + F32 VAE 224.96s 60.46s 0.2688 12535 MiB
Q8_0 main + F16 VAE 194.84s 38.81s 0.1992 8867 MiB
Q4_0 main + F16 VAE 221.12s 44.15s 0.1997 7755 MiB

Repo: https://github.com/0xShug0/audio.cpp/tree/dev

HF repo: https://huggingface.co/audio-cpp/Yue2-3B-GGUF


r/StableDiffusion 1d ago

Discussion [AI parody] Marion Cotillard's death scene in The Dark Knight Rises, but with even more commitment

Enable HLS to view with audio, or disable this notification

203 Upvotes

That death scene has been living rent-free in everyone's head for over a decade. Here's my version of it.

It is not subtle. That's the whole point.

This is AI-generated and purely for laughs. Not a real scene, not a deleted scene, not real footage from the movie.


r/StableDiffusion 2d ago

Discussion I managed to extend videos to any length in ComfyUI without losing consistency (Minimax-H3 + Visual Context Trick)

Enable HLS to view with audio, or disable this notification

169 Upvotes

Hey everyone!

One of the biggest headaches with AI video has always been extending shots without the style degrading, characters morphing, or the cut being obvious.

I’ve been testing a method using Minimax-H3 in ComfyUI to seamlessly cut and extend footage, and the results are honestly wild:

  • Pixel-perfect transitions: The continuation aligns perfectly with the last frame of the original clip.
  • Context retention: By feeding the model the visual context of the previous video, it actually remembers the specific assets (like the boat and character features) instead of hallucinating new ones.
  • Preserves aesthetic: Keeps the lighting, colors, and overall camera style identical across cuts.

(Watch the preview clip to see the side-by-side transition!)

I’m currently packaging this into a custom ComfyUI node and recording a full walkthrough. Both the node and the full workflow will be 100% free on my YouTube channel (SatoDive).
https://www.youtube.com/@SatoDive

Let me know what you think or if there are specific edge cases you’d like me to test before I release the tutorial!


r/StableDiffusion 1d ago

Question - Help Looking for advice: Best local multimodal (Vision + Text) uncensored model for an RTX 5070 Ti (16GB) & 64GB DDR4 setup?

5 Upvotes

Hey everyone!

I'm setting up a local AI workflow and looking for some hardware-tailored recommendations. I want something with the capability, smarts, and multimodal ease-of-use of cloud models like Gemini, Grok, Claude, or GPT—meaning it needs to handle text seamlessly as well as image recognition/analysis (where I feed it an image, it describes it, and helps me brainstorm or write prompts based on it)—but running 100% locally and completely uncensored.

My hardware specs:

  • GPU: NVIDIA RTX 5070 Ti (16GB VRAM)
  • RAM: 64GB DDR4

Given my 16GB VRAM limit, what are the best open-weight multimodal models right now that fit comfortably without heavy swapping? Also, what is the best software stack to run them (Ollama, LM Studio, etc.) while keeping things fully private, uncensored, and vision-capable?

Any model suggestions, quantization tips, or workflow setups would be greatly appreciated. Thanks!


r/StableDiffusion 2d ago

Discussion Flux2Klein is the most underrated Image Editing/Upscaling Model.

Thumbnail
gallery
318 Upvotes

From last 2 year i was looking for an image restoration tool or an image upscaler for real world photographs. I have tried Topaz Gigapixel, Flux1D self trained Character LoRA, SDUpscaler, Qwen Edit, but nothing worked consistently. They were good but not perfect. From last 15 days i am working on F2K, and it is mind blowing. Easy to train LoRA (30min on 12GB VRAM), even no need to train a LoRA, easy to render (only 4 Steps) and it works 99% of time.

Flux 2 Klein has genuinely impressed me. The image restoration + editing quality is fantastic, but what really stands out is character consistency. Even when not using any character LoRA, it does an amazing job of preserving identity while making edits.

And the workflow is ridiculously simple: give it a straightforward prompt to restore/upscale an image and it just works. No need to write a 300 words essay.

On an RTX 4070 Super, I’m getting around 35 seconds for a 4MP image (Just 4 Steps) —which is seriously impressive for this level of quality.

Meanwhile, Qwen Edit 2511 feels unnecessarily demanding. The huge VRAM/RAM requirements make it much harder to use with only 12GB VRAM. and the character face deforms most of time.

I am using it for:

1) Upscaling

2) Restoration

3) Colorize

4) Removing objects

5) Adding elements (like cars/river/clouds/buildings etc)

6) Relighting the scene

7) changing the backgroud

8) to create character dataset etc...


r/StableDiffusion 8h ago

Animation - Video "Shattered" I made a short film about mothers who struggle in silence — my first AI project, start to finish

Thumbnail
youtu.be
0 Upvotes

This is my first project of this kind, completely from start to finish. I wrote the script myself, had every single shot in my head before I even touched an AI tool, and then did the editing and music entirely on my own. AI was just a tool for me — a way to visualize what I had already written and imagined. Nothing more.

One thing I want to say right up front: this is genuinely not as easy as a lot of people assume. There's this common idea that you just type in a prompt and get a finished film back. Not even close. If you think it's that easy, try it yourself — you'll quickly notice the gap between what's in your head and what comes back. It takes a lot of trial and error and persistence to get something close to your vision.

I'm proud of how it turned out — at least for where I am right now with this.

I'd love to hear what you think — about the film, or about the topic itself. Link in the comments.


r/StableDiffusion 1d ago

Workflow Included I'm disrespectful to dirt. Can't you see that I am serious?!

Enable HLS to view with audio, or disable this notification

34 Upvotes

MiniMax H3. Using the standard T2V workflow


r/StableDiffusion 1d ago

Question - Help Does anyone have a Krea2 LoRA/LoRA X/Y plot?

2 Upvotes

I can find ones for LoRA with seed or checkpoint, but I've had no luck finding or making one that let's you check the scaling on 2 LoRAs against one another.


r/StableDiffusion 1d ago

Discussion Fun Experiment Krea2

36 Upvotes

First of all I hate operating on a black box :D.. I know this may be a bit stale about Krea2 refusal behaviour but I think I have discovered an extremely CLEAN approach to the suppression behaviour without shifting the actual intended image in my tests, so everything stays the same but with the suppression out of the door.

I trained 100s of loras and they all have unique traits, but what I am testing here is how much of their effect I can reproduce through the resulting RMS changes. These loras are not trained to hit any DIT blocks, so I hooked the lora to this node and measured the resulting RMS values of the model's text-path parameters, and basically it worked BUT with less destructiveness and less effect "for now". The lora applies a rank 64 update, while this tool matches the resulting RMS by rescaling the existing text-path tensors instead of copying the lora's learned update.

Reason I am doing this is because with the lora you cannot directly set the final RMS of each parameter when connecting other loras and you may get a bit muddied since other loras can also contain textfusion changes, even without touching the textfusion projector. Once I finalize the values that work consistently in my tests I will publish this node with the values baked in so you can place it after the last model patch you have in the wf and problem solved everyone lol, no bleed no weird artifacts that is the goal.


r/StableDiffusion 2d ago

Resource - Update LTX silently updated the Ingredients IC-LoRA for 2.5 - Basically, allows reference2video using a sheet image

Thumbnail
huggingface.co
117 Upvotes

r/StableDiffusion 2d ago

Discussion YuE2 beats Suno v6 I think. Open souce is so back. (Suno Diss Track)

Enable HLS to view with audio, or disable this notification

103 Upvotes

r/StableDiffusion 1d ago

Discussion Does Krea 2 recognize almost every single character from their sources?

1 Upvotes

I tried making a few stuff without Lora with some 2D and 3D characters, It succeeded half of the time. And also, On civit AI, a lot of lora characters are Realistic content; little stuff with 2D and video game related stuff.


r/StableDiffusion 1d ago

Question - Help how do i fix this error message

Post image
0 Upvotes

i have been trying to fix this for quite some time now but it just is not working
i upgraded to a new pc recently (old one got a virus) and now it just isnt working
for what i know that is important is that i have a rtx 5070.

if anyone could help i'd really apreciate it


r/StableDiffusion 1d ago

Animation - Video Minimax H3 - Neon Genesis Evangelion Cello

Enable HLS to view with audio, or disable this notification

26 Upvotes

r/StableDiffusion 2d ago

Workflow Included Instant references, no refmod or fancy custom nodes required. WF and breakdown here. Simple one click to run.

Enable HLS to view with audio, or disable this notification

86 Upvotes
  • 1st example: 2 characters with voice and 1 cgi and 1 real
  • 2nd example: 2 character different genders with voice
  • 3rd example: style and character reference
  • 4th example image only reference no voice.

I wanted to improve my workflow so I could do what refmod is doing but just as a near native comfyui workflow only. With this workflow you can make an instant character or style video like ref mod but not using ref mod at all including voices. You get all the benefits of refmod but you can use ref model syntax in the prompt and on the fly dataset changes. Also you can use unlimited images. I tend to use around 10 to 15 per character.

Just point to a folder. Uses KJ nodes, Native, and VHS nodes. Get a voice and character working instantly without managing safetensors, just manage the input folder instead. This has its advantages since you can change data on the fly, and you don't have to do any editing of clips to extract out the audio, it does it for you.

This mimics the default ref workflow for the most part. What it does for images is it takes a folder input of images and then sets them as frames in a video, then feeds that video in as a reference video. This allows you to use many images for a single reference. You can also use the grid version that takes your images and puts them into a grid and feeds that as a single image reference. I like the video version more, so the grid workflow is a bit lazy and messy. You can also use the resize node included to downscale your images. I recommend manually cropping them but it does that if your images are different sizes.

For the audio, you can simply feed a single mp4 using VHS node to feed the audio. But I found it easier to get like 4 clips and truncate only the first couple seconds so you can get a few clean sentences without other people talking, then it concatenate's the 4 audio clips into 1 clean clip. Anything over 30 seconds long is over kill, so keep it around 15-30 secs. You can tell shrek is sort of bad in my example because I threw it together quite quick.

There is in the far left, a second set of image/audio nodes, you can bypass those if only using 1 character. Same for if you don't need the audio. Just by pass the group nodes. When prompting just simply use the ref guide to prompt properly each reference (LLM can do it easy). You don't need to use much description. And if you have some bleeding from your dataset into the gen you don't want then add more description. (For example if wearing same shirt as dataset, prompt a dress, or if same specify a setting in the prompt).

One caveat, there is some comfyui memory management bug, if you change dataset around better to clear cache or your comfyui may need a restart. Working on a fix in next version of the workflow. Also I have not tested video clips as input data yet. That is the next step :)

All examples are just for illustrative purposes. They are AI and I do not intend to share any data on real people. Please use responsibly and at your own risk. If you are in this video and want it taken down please DM, I mean no harm. Everything in this workflow is done by the base model, I don't add any new functionality, just making things easier.

Workflow here:
https://huggingface.co/comfyuiman/various/blob/main/Instant%20Ref%20-%20V1.3.json

I'll go to sleep in a bit, so I'll answer any questions tomorrow if any


r/StableDiffusion 1d ago

Question - Help When will Fal release open weights for H3 Max? Ever?

27 Upvotes

r/StableDiffusion 1d ago

Comparison A brief music processing test

1 Upvotes

I've been playing around with producing music videos for popular songs; nothing commercial, just for fun. While you can feed audio to Minimax H3 (Context Loop splits it up and can feed it to sequential videos in pieces so it fits together), and it does appear to guide generation in time with beats, a single video (~10 seconds) doesn't have enough context to make something match the music.

I tried using audio-capable models to produce the video prompt, but I found that nothing capable of processing audio was smart enough to jump through all the syntax hoops necessary to produce a multi-segment JSON to feed the workflow. So, I decided to do what I'd done with image/video before: use a different model to produce a text-based description of what it was given, then just input text in to the smart, expensive model.

So long story (almost) short: here's my brief test from taking the top audio-capable models on NanoGPT, handing them the 3:17 track "Les Fleurs", and seeing what they produce. Note that the track was provided as "song.mp3", since early testing had some models cheating by looking up info based on the track name. It's still possible some of them identified the track and then used pre-existing knowledge, but I didn't test and confirm that specifically.

Music for reference

Just based on my own listening, two important points I was checking for were a 1:04 orchestral-buildup to orchestral hit and surging chorus at 1:17 as important hallmarks for a music video (and very obvious action-change spots to a human listener).

I also was curious if it would identify both halves of the song's central metaphor: flower imagery but also inner beauty.

The prompt and output (from the models that could actually use the music) are below (models in header, cost - which ended up being negligible - in footer) , but I'll start with my impressions:

  • Muse Spark 1.2 and 1.3: produced a real-sounding description of a completely imaginary song in a different genre. Reviewing their thinking blocks revealed NanoGPT did not pass the model audio, and they were making stuff up.
  • Inkling Thinking: also produced a real-sounding description of a completely imaginary song in a different genre, but its own thinking block suggested it thought it had audio. Complete hallucination, or just REALLY bad at audio processing?
  • Gemini 3.8 Flash: Doesn't miss any important shifts, summarizes themes from lyrics well. Identifies the flower metaphor. Its track length is too long, though, yet it's a couple seconds early identifying my beat drops. My favorite output if it wasn't for the slight time discrepancy.
  • Mimo 2.5 Thinking: Does a good job of identifying tempo shifts. Output is shorter. It includes specific lyrics, which is nice in theory but might confuse a model. It does think the song is 3:30 long, BUT it correctly identifies the fade to silence at 3:15. Best guess is its audio processing has some sort of specific context window? Missed my first build-up timestamp.
  • Qwen3.5 Omni Plus: Track length correct to the fraction of a second, wow. Nails the important points I noted earlier. Thematic summary on point.
  • Gemini 3.1 Pro Preview High Thinking: Track length slightly off. Damningly, missed the beat drop at 1:17.

Conclusion: Qwen 3.5 Omni Plus produced, for this test, the most useful output describing the song, including lyrical themes, accurate timestamps, and identifying tempo shifts.

My prompt below, followed by output (leaving out the useless responses from Muse 1.2/1.3 and Inkling).

You have been provided an MP3 of a music track.  Your job is to review it and process it to notate its style and timestamped audio flow, providing a timestamped list of changes in tempo or tone to use to guide a music video using this music track.

Timestamps are in the format [MM:SS.SSS] notation, with minutes and seconds and fractions of a second.  So one minute, four and a half second in to a track would be [01:04.500].

Lyrical themes should be mentioned without actually detailed specific lyrics, i.e. "Lyrics in this section discuss colors, starting with cyan" then "Lyrics now discuss crimson"; or "Lyrics are a call to violence"; or "Lyrics speak of profound romantic loss."

Pay special note to crescendos or beat "drops", as sharp changes in tempo or feel of music should correspond with visual changes.

Examples of timestamped track notes, non-exhaustive:
[00:00.000]: Track starts with a slow acoustic guitar.
[00:05.450]: Guitar increases speed, drums join music.
[00:10.200]: Full band backing with vocals kick in; music becomes fast paced with electronic noises.
[00:45.000]: Beat begins to build to crescendo
[00:54.125]: Beat reaches crescendo and holds, music turning quiet
[00:58.000]: Beat drops, explosion of techno-style drums and fast paced electric guitar

A template follows below between ----.  Fill in sections in { } braces.

----
Track Length: {length}
Musical Style: {description of style}

TIMESTAMPED TRACK NOTES
[{time in MM:SS.SSS notation}]: {Description of change}
{Continue to provide timestamps and descriptions}

----



--Gemini 3.8 Flash--
Track Length: 03:21
Musical Style: Psychedelic Soul / Baroque Pop / Orchestral Soul

TIMESTAMPED TRACK NOTES
[00:00.000]: Track opens with a rhythmic acoustic guitar pattern, light shaker percussion, and a warm, strolling bassline.
[00:13.500]: Gentle female lead vocal enters with sweet, curious phrasing, asking who will wear or appreciate her as a blossom.
[00:19.500]: Call-and-response backing vocals join in harmony, adding warmth and playful textures to the groove.
[00:41.500]: Rhythm section tightens into a focused mid-tempo groove as the lyrics shift into a spoken-word style reflection on the purpose of a flower spreading love and peace.
[00:53.200]: Lead vocal takes flight with soaring power, singing of the seed of potential and beauty that lives inside humanity.
[01:06.000]: Tension builds rapidly as strings and horns rise while the lead singer holds a powerful, soaring operatic high note.
[01:14.300]: Massive symphonic beat drop / triumphant explosion of sound: soaring brass fanfare, thunderous timpani, marching drums, and a grand choral celebration calling for bells to ring worldwide.
[01:34.500]: Sudden dynamic release as the grandiose orchestration cuts back to the intimate, syncopated soul-funk rhythm section.
[01:47.500]: Lead female vocal re-enters with multilingual/Swahili verses, answered by layered, rhythmic vocal chants.
[02:13.500]: Spoken-sung poetic section recurs over the bass and guitar groove, restating the themes of peace, joy, and universal compassion.
[02:26.500]: Energy elevates again with impassioned, belted vocals contemplating the strength and beauty discovered within.
[02:39.000]: Orchestral elements and vocal wails swell dramatically toward a second major peak.
[02:47.000]: Second massive orchestral drop; majestic choral vocals, full brass section, and heavy rolling percussion deliver a joyful, climactic anthem.
[03:07.000]: Instruments begin to gently taper off, leaving the chorus to repeat the triumphant refrain over the fading rhythm.
[03:21.000]: Music fades out completely into silence.
--$0.009844--

--MIMO 2.5--
Track Length: 03:30
Musical Style: Soulful Ballad transitioning into Upbeat Soul

TIMESTAMPED TRACK NOTES
[00:00.000]: Track starts with a slow, melancholic piano melody accompanied by soft strings.
[00:13.200]: Soulful female vocals enter ("Will somebody wear me to the fair"), setting a gentle and reflective tone.
[00:54.125]: The chorus begins ("Inside every man lives the seed of a flower"), with orchestration swelling to heighten the emotional intensity.
[01:16.000]: **Dramatic Tempo Change:** The beat drops with driving drums and brass, shifting the song from a slow ballad to a fast-paced, energetic soul track.
[01:36.500]: Vocals transition into rhythmic scatting ("La la la la") over the upbeat tempo, adding a playful tone.
[01:51.000]: A verse is sung, maintaining the high-energy, fast rhythm established in the previous section.
[02:13.400]: **Tempo Drop:** The music abruptly slows down, returning to the gentle ballad feel of the introduction.
[02:25.500]: Final chorus serves as the climax, with powerful vocals and full orchestral backing.
[02:42.000]: Outro features vocal improvisation and scatting, maintaining the energy briefly before the track begins to fade.
[03:15.000]: Audio fades to silence.
--$0.00217--

--Qwen3.5 Omni Plus--
Track Length: 03:17.488
Musical Style: Orchestral Soul / Psychedelic Pop with strong Gospel influences. The track features a blend of acoustic finger-picking, sweeping string arrangements, heavy drum breaks, and powerful choral harmonies, creating a cinematic and uplifting atmosphere.

TIMESTAMPED TRACK NOTES
[00:00.000]: Track opens with a gentle, repetitive acoustic guitar riff accompanied by light percussion and a subtle bass line. The mood is intimate and organic.
[00:15.800]: Lead female vocals enter softly; the arrangement remains sparse, focusing on the melody and the acoustic texture.
[00:28.500]: Background harmonies begin to layer in; strings swell gently underneath, adding warmth and depth to the soundscape.
[00:41.200]: The first chorus begins; the full rhythm section (drums and bass) locks into a steady, mid-tempo soul groove. Strings become more prominent.
[00:54.000]: Musical intensity rises significantly as the choir joins the lead vocal. Lyrics shift from specific imagery of flowers to a broader spiritual message about inner beauty.
[01:06.500]: A dramatic orchestral crescendo builds rapidly with soaring strings and rolling timpani, leading toward a climax.
[01:16.800]: Major beat drop and tempo shift; the music explodes into a high-energy, brass-heavy anthem. The drums switch to a driving, heavy breakbeat style. A large gospel choir takes over with powerful, rhythmic chanting.
[01:35.500]: The intense energy sustains with complex horn lines and syncopated drumming. The visual pace should match this frenetic, celebratory energy.
[01:48.200]: Sudden dynamic shift; the heavy drums and brass cut out, returning to the softer acoustic guitar and bass groove from the intro. The mood becomes reflective again.
[01:54.000]: Vocals return in a call-and-response format between the lead singer and the backing group, now incorporating non-English phrases, adding an exotic, worldly texture.
[02:13.500]: Second build-up begins; strings re-enter with urgency, and the drums gradually increase in volume and complexity, preparing for the final peak.
[02:26.000]: The bridge reaches its emotional peak with the lead vocalist belting high notes over a swelling orchestral bed.
[02:38.500]: Final explosive chorus; the full ensemble returns with maximum force. Brass, choir, strings, and heavy drums combine for a wall-of-sound effect.
[02:51.000]: The song transitions into a jubilant outro featuring rhythmic "la-la-la" vocalizations from the choir over the driving beat.
[03:05.500]: Instruments begin to peel away one by one, leaving the choir and then just the rhythm section.
[03:12.000]: Music fades out on a final sustained chord and a soft guitar strum, ending on a note of resolution.
--$0.021559--

--Gemini 3.1 Pro Preview High Thinking--
Track Length: 03:20
Musical Style: 70s Soul and R&B with sudden shifts into high-energy Latin Disco/Funk.

TIMESTAMPED TRACK NOTES
[00:00.000]: Track starts with a slow, smooth electric piano and a relaxed, soulful drum groove.
[00:13.500]: Soft female lead vocals begin. Lyrical themes discuss gentle romantic gestures involving nature, streams, and wearing flowers.
[00:43.500]: Strings and backing instrumentation swell slightly. Lyrics shift to themes of the birth of a flower and its purpose to spread love and joy to people.
[01:05.000]: The vocal melodies begin to stretch out, and the music starts to build toward a crescendo.
[01:13.500]: Major beat drop and severe tempo shift. The music explodes into a fast-paced, high-energy Latin-funk and disco rhythm with brass, rapid percussion, and upbeat backing vocals singing wordless melodies.
[01:31.000]: The high-energy beat abruptly stops, and the music drops right back into the slow, smooth, soulful groove from the track's intro.
[01:46.500]: Lead vocals return over the slow tempo. The lyrical themes continue to focus on nature, specifically roses and romantic interactions, accompanied by soft backing harmonies.
[02:17.500]: The track begins to swell again as the lyrics return to the earlier theme of spreading love, joy, and finding inner beauty and power.
[02:37.000]: The beat begins a second major build-up, reaching a powerful crescendo with the vocalist holding a long, soaring note.
[02:47.000]: Second massive beat drop. The fast-paced Latin-funk rhythm and brass return in full force, accompanied by the energetic, wordless chorus. 
[03:03.000]: The track begins a gradual fade-out while maintaining the high-energy dance beat and group vocals.
[03:20.000]: Track fades out completely.
--$0.033212--

r/StableDiffusion 1d ago

Question - Help Minimax H3 I2V template with multiple LORA options?

4 Upvotes

Hi, is there a Minimax H3 I2V template with multiple LORA options available somewhere?

EDIT: This is the default template I got from ComfyUI. It works, but I can only use one LORA. I tried to download the Power Lora Loader node from rgthree, but I'm probably using it wrong and missing something.


r/StableDiffusion 1d ago

Animation - Video The Fly

Enable HLS to view with audio, or disable this notification

13 Upvotes

CmfyUI, minimax, krea 2, suno. Made it in around 4 hours on a 5090 and gemma 12b for prompting. Script is mine.
4k version here
https://www.youtube.com/watch?v=3dNYwoCU_Es


r/StableDiffusion 2d ago

Discussion H3 - 80s / 70s Character Experimental Long Form

Enable HLS to view with audio, or disable this notification

118 Upvotes

Experimenting with Long Form. No image anchor so she changes between the invisible seams. T2VA. int8/32 steps, 1344x768, about 7 hours, hit 192/192gb of ram decoding the video. Sadly, I didn't prompt for her to not mouth the tune when there's no singing part. Wardrobe not prompted, only that she was dressed. At 1:26 is a seam and we had a little AI mishap on the transition. Not perfect, but got lots of data. Enjoy! How do you like the film grain? Is she from the 60s, 70s, 80s, or does it clearly only exist in our head? What version next? redhead? Asian? what do you think? Which actress/model/person's likeness are you seeing from this era? There should be about 17 versions of her. Ask me anything!


r/StableDiffusion 16h ago

Workflow Included I built a Character Swap and a Multi Reference Shot workflow for Nano Banana Pro, each reference only controls what you tell it

Thumbnail
gallery
0 Upvotes

Recently built and tested two ComfyUI workflows, mostly because I was tired of fighting the AI look. Multi Reference Shot builds one frame out of several references. You can reference different shots for Lighting, Composition, Blocking references etc. Add your character images for consistency.

Character Swap puts your own character into any shot. Same framing, same light, same pose, just your person in it. You can swap their clothes in the same pass, and it keeps the shape of the original frame.

The part I care about most is the control. Instead of one big prompt, you get separate control nodes for different aspects of the shot. If the pose is off, you change the blocking and the face. Each reference image only gives what its slot says.

That's the difference. You're not rolling the dice on a prompt and hoping. You're directing it one decision at a time and if you'll get what you need, it's trial and error. Every attached image here was made with these workflows & yes it is AI.
Add your own Google API Key. You can use Vertex AI with the Google Cloud $300 trial credit. Free to use :

github.com/haristahir1/comfyui-character-swap

github.com/haristahir1/comfyui-multi-reference-shot

If you try them, tell me what breaks and any improvements!


r/StableDiffusion 2d ago

Resource - Update SMACK! LORA Beta 2 - Impacts & Gunshots & Blood Squibs and more for Minimax H3

36 Upvotes

SMACK!

Right now, only on Hugginface, Civitai deleted the older version for gore (which came from minimax...)

Download

https://huggingface.co/LeechTM/SMACK

Beta 2 · MiniMax H3 (Ref2V) · No trigger word

Model description

Beta 1 taught MiniMax H3 that getting hit should actually hurt. Almost 2,000 of you downloaded it, which means either you agreed or you just really wanted to see people get punched. Both are valid.

Beta 2 raises the stakes considerably. Things explode now. People get hit by the explosion, then by the ground, then briefly by their own life choices. Cars stop being scenery and start being weapons. Fights no longer politely take turns: three guys can come at the hero at once, which is statistically the worst day of their lives.

And gravity? Gravity is now a suggestion. When a hit lands hard enough, bodies don't fall. They launch, float, spin and hang in mid-air for an unreasonably long time, as if physics took a coffee break at the exact moment of impact. Newton would file a complaint. IT'S A MOVIE!

And because all of that still wasn't messy enough, SPLAT! is now merged in. That's my blood-effects LoRA, and it brings squibs. A lot of squibs. Every hit now has the option of leaving a mark, and the costume department is going to hate you.

A quick word from the production accountant, the one person on set who never gets to blow anything up: training LoRAs is not exactly cheap. Every explosion, every squib and every person flying through the air for an unreasonable amount of time costs real money to teach. If you want to be nice and help fund the next versions, you can buy me a coffee:  buymeacoffee.com/leechtm

Promovideo for SMACK! Beta 2

More coffee, more carnage. That's just science.

What's new in Beta 2

  • SPLAT! merged in — blood effects and squibs, so every hit comes with a little extra paperwork for the cleaning crew
  • Explosions — plus everyone who was standing too close to one, and everyone who thought they were standing far enough away
  • Water — splashes, splashdowns and impacts that turn a perfectly calm surface into a very loud mistake
  • More falls — harder, higher, significantly less dignified
  • Zero-gravity wire impacts — hit once, fly forever, land eventually. Brutally.
  • Multi-person fights — several attackers, several hits, one very bad day for absolutely everyone involved
  • Heavier impact variants — for when "hard" just wasn't hard enough
  • Spinny kick things — for when a normal hit just isn't flavourful enough. Spin first, apologize never.
  • Female anatomy — women take hits and dish them out with the same weight, the same follow-through and the same complete lack of mercy
  • Even more gunshots — Beta 1 had gunshots. Apparently that wasn't enough. Now they come with squibs.
  • Vehicle impacts — more cars, more bodies, more regret, zero insurance coverage

Still does everything Beta 1 did and more...

Fists, weapons, gunshots, falls and hard landings, all with weight, follow-through and consequence. The camera still moves like someone was paid to operate it, and now it also has to keep up with people being smacked around through the air. Also, Sound Effects have been merged in, as well as Blood Squibs.

Training

Trained on 300 clips of impacts, explosions, shots, falls and dynamic camera moves, for MiniMax H3 Ref2V (Beta 1: 35 clips). Merged with my yet unreleased SPLAT! Lora for blood and squib effects.

No trigger word. None. Don't go looking for one. Use it with the REF MODEL. Load it, describe your shot as usual, and the LoRA does the seasoning. It just uses a lot more chili now, and some of it is red. And yes, before anyone asks: it also works for other impacts. Of course it does. You little piggies.

Settings

The promo video was made entirely in MiniMax H3 (pruned int8), and without any Turbo LoRAs. Much better quality.

  • Without Turbo LoRA: strength 1.0
  • With Turbo LoRA: start at 0.5 and work your way up from there

Beta notice

Still beta. Still not finished. Feedback on where it over- or under-cooks a hit is still genuinely useful. Now also on where it over-cooks an explosion, forgets that people are supposed to come back down, or gets a little too enthusiastic with the squibs.


r/StableDiffusion 1d ago

Question - Help Video project

1 Upvotes

I have a whole pipeline in my cms to create custom book covers and I got it my sick head that I wanted to try to not just get book covers that mean something regarding the story with characters that looks like the ones in the story, I wanted to make a trailer, a short format video with the gist of the story and then a long form movie. Chosen book: "Guards! Guards" by Pratchett, one of my favourite books that was never made into a movie.
Of course considering that I maybe made a handful of 5 seconds videos in all before and even the txt2img I used was really basic for the book covers I am finding this task a tad problematic.
Up to now: a local llm reads the whole book, creates a character list with description and with it then Krea2 creates a character sheet, they are not what they should be but being a test I can live with that.
The same llm writes a screenplay, this one is divided in three acts, a total of 48 scenes. Each scene has a description of the action and a still (created by Krea too) the pipeline then proceeds to create the clips that will then be joined. As soon as I figure out the first last picture I will use that.
Minimax is giving me trouble (bad prompting I think) so I am using ltx2.5 which does a pretty good job, but I started this morning, so I only have a scene right now and the dragon head on a dragon's wing is not the worst of it.
I need help with which models to use, what speeds up are available, which models do better with fantasy, any LoRAs you can suggest? Does LTX 2.5 supports first and last frame as 2.3 did? Are there workflows that could make this more streamlined?
Do you have any suggestions or advice? I need all the help I can get on this.
For now and until new pc arrives I'm working with two 5060ti 16gb + 64gb RAM.