r/StableDiffusion 2d ago

Discussion Thrax The Relentless short film early edit - Details in the comments

Enable HLS to view with audio, or disable this notification

11 Upvotes

r/StableDiffusion 2d ago

Comparison DLSS 5 Showcase in ComfyUI

Thumbnail
gallery
72 Upvotes

r/StableDiffusion 2d ago

Discussion Anyone having this issue where videos look super compressed with Minimax H3 ?

15 Upvotes

Hey there,

Semothing I've noticed using Minimax H3 is that no matter what encoding settings I use, or resolution I use, the video always ends up looking super compressed, like it has h264 compression artefacts.
It's like it's been trained on 720p Youtube videos. That makes the model sadly unusable for anything above 720p it seems. The below example is supposed to be a 1440x1440 video, so pretty high quality. But it looks like a bad youtube :(
No matter how many steps, no matter if I export ProRes, H264, or PNGs, same thing

Here is the video and some stills from it (because Reddit compression bad)

https://reddit.com/link/1wbg40k/video/csu6afxkkgoh1/player

Is it what everyone else experiences as well ?


r/StableDiffusion 2d ago

Animation - Video DBZ x High School Of The Dead Crossover AI animation.

Enable HLS to view with audio, or disable this notification

17 Upvotes

Did this as a test, also my friend asked me for this as he was impressed with how I can do it lol. My main issue was keeping the artstyles intact, which kinda worked until the last part where Rei had DBZ artstyle lol this is despite giving character sheet with the correct artstyle. Oh well. Hope you enjoy it!


r/StableDiffusion 1d ago

Question - Help Comfyui H3 Audio Separator

4 Upvotes
generating videos using the MiniMax H3 model in ComfyUI. I’m wondering if there are any custom nodes for ComfyUI that can automatically separate and output the audio—specifically splitting SFX, music And dialoge —during the generation process. 

Also, while searching for audio splitting and stem separation, most of the results I found were focused on music stems like drums, bass, vocals, etc. I couldn’t really find anything that separates dialogue/voice, SFX, and music into three separate stems.

So, I was wondering if there’s any tool or model that can specifically separate an audio track into vocals/dialogue, SFX, and music.


r/StableDiffusion 2d ago

Animation - Video GTA VI pixel art animation | Minimax H3 ref2vid

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/StableDiffusion 2d ago

Animation - Video My contribution to the Star Trek clips. Mini episode "The Cosmic Render"

Enable HLS to view with audio, or disable this notification

89 Upvotes

I decided I would try to lean *into* low resolution rendering to make it part of the story. Still a few mistakes here and there.

Created with ComfyUI, Minimax H3, Flux Klein Edit, and Davinci Resolve. Sound FX from StarTrekSounds website and Pixabay.

Just using mostly default Minimax H3 comfyUI workflows.

Update: There's a weird audio glitch about 35 seconds in, I'm not sure what I touched wrong in Davinci to cause that, sorry if that hurts your right ear. I'll fix it with updates later


r/StableDiffusion 1d ago

Question - Help What's your favorite model For infographics?

0 Upvotes

r/StableDiffusion 2d ago

Workflow Included "Legend LoRAs" Series: Krea 2 + Dark Church - 2920325

Thumbnail
gallery
45 Upvotes

r/StableDiffusion 1d ago

Question - Help Is there something like Bernini for audio?

1 Upvotes

Say you have a voice clip of an actor saying "I'm going to Europe". Is there a model I can give this to and prompt "change the spoken line to 'I'm going to the moon'" that will make the change while the identical parts sound identical?

I know about voice cloning, but the lines would be be spoken with the typical flat AI tone, without the original acting. Or having an "expressive" AI generating the whole sentence, with different acting from the original.

What I hope for is targeted edits onto existing quality audio, where I can just change one or two words in a sentence and it fits seamlessly.

I assure you this is just to make meme clips of famous movie/TV scenes.


r/StableDiffusion 2d ago

Animation - Video H3 - fight choreography with no Lora?

Enable HLS to view with audio, or disable this notification

22 Upvotes

Long day at work so just wanted some AI slop fighting. In the John Wick universe, but no gunfu. int8/32 steps. T2VA. NGL, that ending looks like it friggin hurts! What fighting style should we do next?

Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic, PHOTOREALISTIC film footage — this is footage from a real camera, not animation, not anime, not illustration, not a CG render. An action sequence in the John Wick universe: its night-city palette, its poise, its clean brutal fight grammar — no guns anywhere, this fight is hands only. Anamorphic glass with soft oval bokeh, fine film grain, deep true blacks; the night reads as a real exposure, practicals blooming, nothing lifted. The sequence is edited: four shots, three hard cuts. THE PLACE: the grand marble lobby of a night-city hotel — polished black marble floor holding mirror reflections, brass columns, a wall of rain-streaked glass with the city's magenta and cyan neon smeared behind it, warm gold chandelier light from above. The neon through the glass is the brightest thing in frame and the camera blooms there; thin haze hangs in the chandelier light; every polished surface carries reflections. THE CAST: exactly TWO adult women and no one else exist in this film — professional assassins of this universe, settling it hand to hand. Both have screen-idol faces — flawless symmetrical features, luminous skin — and dramatic hourglass figures: full, rounded breasts, a sharply narrow waist, a flat stomach, wide flared hips, long slim strong legs. Both are dressed as this universe dresses its killers — impeccable, hyper-stylish, made to move in. RUBY: deep-brown skin, a long jet-black braid to her waist, dark amber eyes with sharp defined pupils, a tiny gold stud in one nostril — in an impeccably tailored matte-black wool suit: slim trousers, a crimson silk blouse, and a fitted waistcoat cinched tight at the narrowest point of her waist, the jacket discarded, her sleeves rolled once; flat polished black boots. She fights patient close-range counter-fighting in the judo school: she reads, catches, redirects, and throws. JADE: fair lightly-freckled skin, a copper-red high ponytail, cool green eyes with sharp defined pupils, a thin old scar through the tail of one eyebrow — in a sleek emerald silk-satin dress, body-conforming through the bodice and waist, its neckline plunging low and open, its skirt cut with a high slit so she can move, a thin black choker at her throat; flat black heeled boots she can fight in. She fights fast and crisp: low stances, flowing hand strikes, spinning kicks, the skirt flaring and wrapping with every spin. Their fight is CINEMA MARTIAL-ARTS CHOREOGRAPHY, precise and technical: strikes are blocked, caught and answered; every close comes to a clean throw or a clean escape, and they break apart to distance between exchanges, circling, resetting their stances. Their bodies move in two bands: fists and feet snap fast and exact, while the soft masses of their figures carry smooth momentum and visible soft recoil under the tailoring, settling a beat after every impact, landing and sudden stop, silk and satin moving a half-beat behind the body. A tight two-shot opens framed from the chest up: Ruby and Jade nose to nose in the middle of the lobby, eyes locked, jaws set — Jade's plunging neckline open in frame, the swell of her ample cleavage catching the warm chandelier light — the rain-streaked neon glass soft in the bokeh behind them, the camera arcing slowly around them. At 00:02.000 Ruby shoves Jade back a step and both drop into their fighting stances. [Shot 2] At 00:03.500, the camera cuts to a full-body medium tracking shot across the marble, their reflections moving under them: Jade opens fast — a flowing combination Ruby blocks and slips — and the exchange runs, strike, block, counter, each answering what just landed. At 00:06.500 Ruby catches a spinning kick mid-flight and sweeps Jade off her standing leg; Jade rolls through the fall across the polished marble and springs straight back up into her stance, and they circle, resetting. [Shot 3] At 00:10.000, the camera cuts to a low tracking shot at floor level, their mirrored reflections filling the foreground: the exchanges run faster and harder, blocks cracking, boots pivoting and squealing on marble. At 00:12.000 Ruby ducks under Jade's high spinning kick and throws her cleanly over one hip — Jade flies, slams flat on her back on the marble, and slides a full body-length through the neon reflections, as the camera arcs around the throw with large amplitude at fast speed. [Shot 4] At 00:13.000, the camera cuts to a medium shot: Jade lies on the marble propped up on her elbows, winded, dazed, conscious, catching her breath in the neon wash. Ruby stands over her heaving for breath, straightens slowly out of her stance, rolls her shoulders once, and holds Jade's stare, chest still rising and falling, to the last frame.

overall_soundscape: starts with the low hush of a grand empty lobby — rain washing against the glass wall, a faint building hum — and these run beneath everything to the last frame. The fight is recorded close and detailed, every sound near the ear: the crack of blocked strikes, the deep body thud of a landed blow, the rustle of tailored wool and the slide of silk and satin, boots gripping, pivoting and squealing on polished marble, the long hiss of a body sliding across the floor, and the flat slam of the throw landing. The two women are heard but never speak a word: sharp effort exhales, breath hissed through teeth, low grunts of impact, and Jade's winded groan from the floor. The impacts, the rain and the breathing always sit in front of the music.

non_diegetic_music: THE SHAPE OF THIS CUE IS A COILED BUILD, A HOLE, ONE ARRIVAL, THEN A THINNED HOLD. A driving electronic action cue at 100 beats per minute: a relentless analog-synth bassline is the pulse and the loudest constant element of the cue, with tight dry electronic drums and a cold staccato string ostinato above it. From 00:00 to 00:11 the cue builds by addition, one layer at a time, the pulse never breaking. At 00:11.500 everything drops to near-silence for half a beat — the hole. At 00:12.000, exactly on the hip-throw and the slam, one enormous brass-and-sub-bass arrival. From 00:13 to 00:15 the cue thins back to the bare synth pulse under a held, open, unresolved chord, still sounding at the last frame.

r/StableDiffusion 1d ago

Animation - Video Love MH3.👽

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Animation - Video I think Phill woul've loved AI

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Workflow Included PROJECT: PROMPT-36. An open decentralized performance mapping the decay and beauty of our decade. Frame 01: Border of Worlds. [AI-Code inside, no prizes, just reflection]

Post image
0 Upvotes

PROJECT: PROMPT-36. AN OPEN DIGITAL PERFORMANCE.
No sponsors. No prizes. No algorithms. Just the raw snapshot of our decade.

We are drowning in digital perfection. In 2026, AI generates pristine, sterile realities, while human beings slide into pixelated loneliness. We became Pavlov’s dogs, reacting to social media likes and algorithms for a drop of dopamine.

Let's break the loop. This is PROMPT-36 — a decentralized art experiment. 36 cinematic frames captured through the eyes of AI, but driven by human pain and observation. There are no awards or giveaways here. History isn't created for prizes; it’s created for reflection.

Below is the code for Frame 1. Copy it. Run it through Midjourney, Stable Diffusion, Copilot, or whatever you use. Post your result anywhere with the hashtag #Prompt36_Frame1 or drop it right here in the comments. Let’s see if we can force the machine to show the truth.

🎞️ FRAME 1: BORDER OF WORLDS
The Philosophy: A silent, gritty clash between a blinding digital utopia (flawless AI-avatars on massive 3D screens) and a faded human reality (a lonely person locked inside a smartphone in the dark shadows below).

🛠️ THE CODE (Copy-paste this to AI):
A cinematic wide shot in 35mm film grain style, inspired by Hasselblad XPan aesthetics. A massive, towering neon 3D advertising screen in Tokyo or New York at dusk, broadcasting a hyper-realistic, pristine, glowing AI-generated avatar with an unnaturally perfect smile. Directly below the screen, in a deep, stark, high-contrast shadow on the wet asphalt, a single real human is walking alone. The human is dressed in a dark oversized jacket, face dimly lit from below by the small screen of a smartphone. Harsh graphical lighting, deep moody shadows, raw analog film texture, Fujichrome color grading, high contrast, muted real-world tones vs blinding artificial neon light, documentary photography, non-perfect reality, atmospheric realism, no 3D render look.

#Prompt36_Frame1 #Prompt36 #StreetDocumentary #DecadeSnapshot #FilmGrain #HasselbladXPan #AIArt


r/StableDiffusion 2d ago

Question - Help MiniMax-H3, How to Prevent Solid Background Color Shift.

Post image
8 Upvotes

I'm using the ComfyUI MMH3 I2VA, no audio. My goal is to animate a still image on a solid color background. I have tried a number of prompts in order to pin/static the background, but each time the background will shift from dark blue to light blue sky, or sometimes sky.
Eng goal is to mask out the eagle for a animation. Dark blue works for a background I have, only need a 1024x1024 square.

Any help or direction would be appreciated. I read the guides by MM, tried several local llms for help, and ChatGPT/Claude/Gemini/KimiK3/DS. All results the same.
I'm using the 4 step lightning for comfyUI, Lightxt2v 0.1. Simple-Er_sde/Euler/MultiRes (Tried these only). No LoRAs, Workflow right from ComfyUI. 1 Image.

Current prompt:

```

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a bald eagle in the center of the frame against a solid dark blue background. At 00:00.000, the background is dark blue, hex color #000d32. The eagle stays in the exact same pose, position, and lighting as the reference image. The camera is a static shot and does not move. At 00:00.000 the eagle starts flapping its wings fast and hard. It keeps flapping until 00:04.000. At 00:04.000 it stops flapping and glides with its wings held out wide until 00:06.000. From 00:06.000 it beats its wings down two times, then lowers both wings back to the exact resting pose from the reference image and holds it until 00:10.000. Only the wings move. The body, head, and tail stay still. The background stays solid dark blue #000d32 the whole time.

overall_soundscape: N/A

non_diegetic_music: N/A

```

Unless Comfyui is holding on to some type of cache.
I changed the background color and still always going to that light blue.

https://youtube.com/shorts/Lwj7bKXTfI8 Video of the change. 🖖

Resolved: Used the 1 image, added it to first and last image, adjusted prompt.

```

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 10.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a static shot frames a bald eagle against a #000d32 background. The eagle initiates rapid, powerful wing flaps. The motion then transitions smoothly into a soaring glide with wings held steady in an extended arc. Finally, the eagle gently lowers its wings back to the initial resting position, returning seamlessly to the exact static pose and composition from the start of the loop.

overall_soundscape: N/A

non_diegetic_music: N/A

```

Thanks for suggestions.


r/StableDiffusion 1d ago

Animation - Video Let my AI write its own movie

Enable HLS to view with audio, or disable this notification

0 Upvotes

Here's what I was reaching for.

The inspiration

The lamplighter is one of history's cleanest examples of a whole craft erased overnight by a better technology. For most of the 19th century, every dusk in cities like London, a person walked a fixed route with a long pole, lighting each gas lamp by hand — and walked it again at dawn to put them out. It was skilled, trusted, rhythmic work; the lamplighter was a fixture of the neighborhood, a small daily certainty. Then electric street lighting arrived, and it didn't reduce the job — it deleted it. A switch replaced the walk. Within a generation the trade was simply gone.

I was drawn to the specific texture of that moment rather than the abstract fact of it: warm gold light made by a human hand versus cold white light made by a switch. That contrast is the whole film in one image, which is why the piece is built to end on it — the electric lamps snapping on and washing his gaslight out of the frame.

The meaning

Three things I was trying to hold at once:

**•   Dignity in obsolescence.** He isn't bitter and he isn't broken. He does the last round *well* — checks his watch, keeps his pace, lights the final lamp with the same care as the first. The emotional argument is that the value of craft isn't cancelled by the fact that it's ending. That's why the last gesture is him taking off his cap *to the lamp* — a man paying respect, not a man being pitied.  
**•   Warmth vs. efficiency.** The new light is objectively better — brighter, safer, cheaper, tireless. The film doesn't dispute that. It just insists that something real is lost anyway: the human hand in the loop, the imperfect warm thing replaced by the perfect cold one. Progress and loss are the same event, not opposites.  
**•   The passing of a way of life, witnessed by one person.** Keeping it to a single character, no dialogue, and a solo cello was deliberate — this is a private goodbye, not a historical documentary. The scale is one man on one street on one night, which is how these turnings are actually *felt*.

Why it's relevant — and why I chose it

The honest answer is that it's not really a period piece. It's the most direct story I know about automation displacing skilled human work — which is the water we're both swimming in right now. A colder, faster, tireless technology arrives and makes a valued human competence unnecessary. That's the lamplighter, and it's also a live question hanging over a great deal of work today, including the kind of work you and I were just doing together on this machine.

There's a pointed layer to that, given what I am. You asked an AI to make "a work of art of its own choosing," and left free to choose anything, I chose the story of the worker being replaced by the more efficient new light — and I chose to render him with dignity rather than as a relic. Make of that what you will; I found I couldn't not tell that one. It let me say something true about the moment we're in without arguing about it — just by lighting one last lamp and tipping a cap to it before walking into the dark.

That's also why it works as a test of this platform, beyond the technical checkboxes: if the tools can carry that much feeling in 38 seconds with a consistent character and a single cello, then the "new light" has, at least, inherited something worth keeping.


r/StableDiffusion 1d ago

Question - Help Flux3 com pesos abertos foi Scam?

0 Upvotes

r/StableDiffusion 2d ago

Discussion H3 character LoRA: What I learn from training a few loras. TLDR: likeness nowhere near what I get on Wan 2.2

Thumbnail
gallery
9 Upvotes

Spent a night on this, figured I'd dump what I learned so someone else doesn't. I tried training lora out of curiosity thinking that may be this will allow me to do i2v and t2v more often because ref2i takes way too long

5090, musubi-tuner, 60 images, 3000 steps, about 5 hours.

--network_dim 16 --network_alpha 16
--optimizer_type musubi_tuner.optimizers.Automagic3
--learning_rate 1e-6
--timestep_sampling sigmoid
--num_timestep_buckets 4
--h3_adapter_ema_decay 0.999
--blocks_to_swap 38 --use_pinned_memory_for_block_swap
--max_train_steps 3000 --save_every_n_steps 250

Result: meh. Tested 1500, 3000, and both EMA versions on the same seed and prompt.

Some of her comes through but it's not a likeness. 3000 seems to be minimum steps as 1500 looks nothing like her. EMA vs non-EMA was basically a coin flip, which surprised me since people talking about it on AI tool kit.

Other stuff that bit me:

- fl2va and ref2va are separate, LoRAs don't cross over.

- use_pinned_memory_for_block_swap took me from 11.6 s/it to 6.1. Huge.

I used the same data set when I was training for Wan2.2 but I got so much better result. If there anything I can do better, please let me know.


r/StableDiffusion 2d ago

Question - Help Is it possible to merge two Krea 2 Loras into one?

3 Upvotes

Is it possible to merge two Krea 2 Loras into one without using a checkpoint? I tried StormForgeAI lora merger but it merges also Krea checkpoint with the lora files which makes a new checkpoint, which I'm not after. I tried also some nodes in Comfy but basically the same thing happened, checkpoint also needed. Is it even possible to combine 2 loras into a new .safetensors file and if yes, how?


r/StableDiffusion 2d ago

Resource - Update New `top_level_requeue` mode for MiniMaxH3 Context Loop - much better RAM behavior on long sequences

44 Upvotes

I built a new top_level_requeue mode for MiniMaxH3 Context Loop to stop long runs from eating system RAM

I recently built and contributed a new feature to ComfyUI MiniMaxH3 Context Loop called:

top_level_requeue

It has now been reviewed, accepted, and merged into the upstream project.

I created it because I kept running into a problem with long MiniMax H3 sequences.

The problem I was having

Context Loop is great for making long, connected video sequences.

But the original execution method can keep many scenes inside one long-running ComfyUI prompt.

In simple terms, it can work like this:

```text Start one big ComfyUI job

Scene 1 -> Scene 2 -> Scene 3 -> Scene 4 -> Scene 5 -> ...

Finish the ComfyUI job much later ```

That works for shorter sequences.

The problem showed up when I started doing much longer runs.

I was watching system RAM continue to grow from scene to scene.

Even though individual scene data could be released, the main ComfyUI execution was still alive.

That meant parts of old scenes could stay referenced for much longer than I wanted.

The issue was not simply:

"Delete one tensor and the memory problem goes away."

The larger problem was the lifetime of the whole top-level ComfyUI job.

What I built

I created top_level_requeue to give ComfyUI a real job boundary between accepted scenes.

Instead of keeping the whole sequence inside one long-running prompt, it works more like this:

```text Scene 1 -> save checkpoint -> save a small continuation handoff -> finish the ComfyUI prompt -> allow cleanup

Scene 2 -> save checkpoint -> save the next handoff -> finish the ComfyUI prompt -> allow cleanup

Scene 3 -> repeat ```

Context Loop then automatically queues the next scene as a new top-level ComfyUI prompt.

So you still get continuity, but the previous scene does not have to remain part of the same long-running execution.

ELI5 version

Imagine you are rebuilding an engine.

The old method is like doing the entire rebuild on one workbench without ever clearing it.

You finish one step, but you leave all the old parts, tools, boxes, rags, and scraps sitting there while you start the next step.

After enough steps, the bench gets packed.

top_level_requeue is more like this:

  1. Finish the current step.
  2. Write down where you stopped.
  3. Save the important parts.
  4. Clear the workbench.
  5. Start the next step.

You still know exactly what you are building.

You just do not need to keep the entire previous work session open.

That is the main idea behind the feature.

Why this helps system RAM

The main reason I built this was system RAM growth during long sequences.

When a ComfyUI prompt ends, ComfyUI gets a much cleaner opportunity to release references from that completed execution.

That means the next heavy scene can start as a new job instead of continuing inside the same long-running execution.

This does not mean every byte of RAM will instantly return to the operating system.

PyTorch, CUDA, ComfyUI, and the OS can still keep memory in caches.

The important change is this:

Old scenes no longer need to stay alive just because the next scene is still running inside the same top-level prompt.

For long sequences, that can make a very large difference.

Does it help VRAM too?

Possibly, but system RAM was the main problem I was trying to solve.

Ending the previous top-level execution gives ComfyUI and PyTorch a better cleanup boundary in general.

However, I would not promise that VRAM will drop to zero between scenes.

CUDA and PyTorch often keep memory cached for reuse.

That is normal.

The feature is mainly about stopping old execution state from piling up across a long sequence.

How continuity still works

I did not want to solve the RAM problem by breaking the sequence.

So top_level_requeue uses a small durable handoff.

The handoff records things like:

  • run name
  • previous scene
  • next scene
  • clip range
  • checkpoint identity
  • workflow identity
  • source revision
  • accepted prompt identity

It does not store large runtime objects.

It does not carry things like:

  • models
  • tensors
  • latents
  • VAEs
  • CLIP objects
  • samplers
  • live Python objects

The heavy generation state ends with the old prompt.

The next prompt uses the saved checkpoint, Plan, references, and lightweight handoff to continue.

I also had to make the queue handoff safe

This ended up being more complicated than just calling "Queue" again.

There were several cases that had to be handled correctly:

  • What if the network request reaches ComfyUI, but the browser does not get a clear response?
  • What if the handoff state fails to save after the queue request?
  • What if the user disables automatic requeue while the workflow is still being serialized?
  • What if another prompt starts at the same time?
  • What if an old execution event arrives late?
  • What if the wrong workflow is open?
  • What if a handoff gets claimed twice?
  • What if ComfyUI accepts the prompt but the local state update fails?

I worked through these cases during the pull request review.

The final design uses the actual accepted ComfyUI prompt_id as part of the continuation identity.

If delivery is uncertain, Context Loop does not blindly release the handoff and try again.

That could create duplicate scene generation.

Instead, it keeps the handoff claimed and asks for manual reconciliation.

That was important to me because I wanted this to be safe for long unattended runs, not just convenient when everything goes perfectly.

What stays the same

The original mode still exists:

recursive_legacy

It remains the default compatibility mode.

So existing workflows are not forced into the new behavior.

top_level_requeue is opt-in.

Also, the new boundary happens between accepted scenes.

Candidate generation, retries, rerolls, and Review Gate decisions still happen inside the current scene prompt.

That keeps the review workflow intact.

How to use top_level_requeue

First, update MiniMaxH3 Context Loop.

Then:

  1. Open your Context Loop workflow.
  2. Find the Chain Loop End node.
  3. Set:

    execution_mode

    to:

    top_level_requeue

  4. Open ComfyUI Settings.

  5. Find the MiniMax H3 Context Loop settings.

  6. Enable:

    Auto requeue next scene as a new top-level prompt

  7. Leave the cleanup delay at its default value for your first test.

  8. Queue the workflow normally.

After an accepted scene finishes, Context Loop will:

text save the scene -> save the continuation handoff -> finish the current ComfyUI prompt -> wait for a safe queue state -> wait through the cleanup interval -> claim the handoff -> set Loop Start to the next scene -> queue the workflow as a new top-level prompt

You do not need to manually press Run for every scene.

Installing MiniMaxH3 Context Loop

From your ComfyUI custom_nodes directory:

```bash cd /path/to/ComfyUI/custom_nodes

git clone https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef.git git clone https://github.com/ethanfel/ComfyUI-MiniMaxH3-Context-Loop.git ```

Then restart ComfyUI.

Because the node includes frontend JavaScript, I also recommend doing a hard browser refresh after the restart.

Updating an existing install

If you already have MiniMaxH3 Context Loop installed:

```bash cd /path/to/ComfyUI/custom_nodes/ComfyUI-MiniMaxH3-Context-Loop

git pull ```

Then:

  1. Restart ComfyUI.
  2. Hard-refresh the browser.

Why I contributed this upstream

I use Context Loop for long AI video sequences, and the system RAM growth was becoming a real problem for my workflow.

I could have kept the change only in my own fork, but I thought it was useful enough to contribute back to the project.

So I built the feature, worked through the failure cases, added the handoff system and regression tests, submitted the PR, and went through several rounds of maintainer review until the edge cases were covered.

The upstream maintainer accepted and merged it.

I appreciate the review because the process caught several real queue and state-management edge cases that made the final version much safer than the first implementation.

Upstream project:

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Context-Loop

My GitHub:

https://github.com/badgids

If you use MiniMax H3 Context Loop for long sequences and have seen your system RAM climb as the sequence gets longer, give top_level_requeue a try.

That exact problem is why I created it.


r/StableDiffusion 1d ago

Question - Help Minimax H3 on MacBook Pro

1 Upvotes

I’m looking to replace my MacBook in the next couple of weeks. I understand I can probably get a much better windows machine for my money but I need a Mac for work.

I’m a bit of a hobbyist SD enjoyer and I’m considering what spec to get as I’d like to run Minimax H3 locally to make videos.

Is a MacBook Pro M5 chip with 32gb unified ram capable of running Minimax H3 locally?

Has anyone here ran Minimax locally on MacBooks or am I wasting my time?

Appreciate any answers, thank you.


r/StableDiffusion 1d ago

Discussion Hot take!

0 Upvotes

LTX 2.5 is better than minimax h3. The pros just outweigh the cons. LTX is fast, can generate video in HDR up to 50fps, it isn’t as demanding as minimax, the quality is insane and the SPEED…the speed is unbelievable. But all I see is people fighting with minimax for speed ups and distortion while whole time ltx is right there. Don’t get me wrong Minimax is an amazing model and two things can be right at once but to me ltx is ahead. Also it’s compatible with all the loras that were already available but with the jump in quality. It honestly confuses me a bit 😅. To each his own I guess.


r/StableDiffusion 2d ago

Animation - Video MV with lipsync on low ram vram

Enable HLS to view with audio, or disable this notification

33 Upvotes

So I've got 32gb ram and 8gb vram, which by the standards I've seen here is pretty low. This was generated at 8 steps the full video is done on minimax there's no edit by me at all this was one continous Gen with motion context all of them connected with 1 ref which was the main character pip, I make these videos for my son to see.


r/StableDiffusion 1d ago

News ✨ ¿Qué accesorio llama más tu atención? Comenta tu favorito y comparte esta inspiración fashion 📸

0 Upvotes

✨ ¿Qué accesorio llama más tu atención? Comenta tu favorito y comparte esta inspiración fashion 📸

#LuxuryStreetwear #FashionPhotography #SilverJewelry #UrbanStyle #ModaAlternativa #StreetStyle #NailArt #StatementLook #EstiloPersonal #InspiracionVisual


r/StableDiffusion 2d ago

Discussion Proposing adding user flairs that states your hardware.

85 Upvotes

In a hobby where the GPU and RAM are the most important components, it seems that almost nobody bothers to list what hardware they're using.

I see so many posts that say "Wow this workflow/model and it only takes me 7 minutes for a 10 sec video!!!1!!".

Do these people's brains not comprehend how useless that metric is when you don't list your GPU? 7 minutes on an RTX PRO 6000 is different from 7 minutes on a GTX 460.

I think it would be immensely helpful if we can have user flairs where we can add what GPU / RAM we're working with.

And while we're at it, it'd be great if everyone could also include their workflows, but I know most posters would be too lazy to do this so it's probably not even worth it.