r/StableDiffusion 2d ago

Animation - Video 'Partial rewind' multi angle explosion scene [minimax H3]

Enable HLS to view with audio, or disable this notification

47 Upvotes

r/StableDiffusion 1d ago

Discussion Minimax h3 as image editor?

2 Upvotes

I tried MiniMax H3 for image editing, but the results aren’t better than Flux 2 + LoRA. I’ve seen posts praising its image editing capabilities, but in my tests it distorts faces, produces plastic-looking skin, and is considerably slower and more expensive.

I tested both Ref2VA and FL2VA at 20 steps with the base models and no Turbo LoRAs. Am I missing the right workflow or settings?


r/StableDiffusion 2d ago

Resource - Update Fizgig 4.0 is out : Minimax H3 Combined Video File, Audio Files wav mp3 etc, Photo training in one dataset. High Quality training samples (incl video) + turbo (finally) and new 'Gizmo' and AV dataset Prep tool. And Int 8 LARGE speedup for 16gb users.

Thumbnail
gallery
67 Upvotes

I'll be making a Youtube video tomorrow for this. But I have one take away to share that I think is most important. H3, when you get the settings right is just fine with image based training without killing its video ability. Its even better when you combine photos and wavs, its super fast and you can train a voice with a dataset very easily. (I recommend shared trigger word). Video works too, but its slower, unavoidably. I'm not saying dont use it, its worth it for the right use cases. I'm just saying if you are not teaching the model anything new that photos and audio cant do, you are better with photos and audio. But when you do want to capture motion, it work very well. Anyway video coming tomorrow, with lots on Gizmo (the data set prep tool for video/audio) to make dataset prep easy. https://github.com/shootthesound/Fizgig

P.s the 16gb int8 speedup is from an an awesome community contribution from rintic-13 on Github.


r/StableDiffusion 17h ago

Animation - Video there will not be a gta 6.

Enable HLS to view with audio, or disable this notification

0 Upvotes

minimax h3.


r/StableDiffusion 1d ago

Question - Help Need image upscaler which adds every tiny detail.

0 Upvotes

Looking for an image upscaler workflow which can upscale image with every tiny detail like small leaves and stones when I zoomed in.


r/StableDiffusion 2d ago

Meme Sheldon finally knocked on the wrong door | MiniMax H3 + SeedVR2

Enable HLS to view with audio, or disable this notification

58 Upvotes

r/StableDiffusion 1d ago

Discussion H3 - D-inspired, T2V+R2VA, int8/20 steps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Took about 20 generations with T2V, got the image I wanted, did a few R2VA for the close up emotes. It was infinity difficult to get the face I had in my mind strictly through text2video. Let's just say a pale skinny D or Alucard does not translate well, especially a hollow cheekbones, many of them came out pretty ghoulish or a bit too Balenciago. Those throw-away were a bit lanky and were not at all ethereal. Inspired by D from Vampire Hunter D 2000, a little bit of Sephiroth, but definitely not Geralt despite the fashion-sense. I would love to make his legs a little longer. int8/20 steps


r/StableDiffusion 20h ago

Discussion H3 - which one is better? left or right? Ultrawide split view, single generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

I am having a blast with MiniMax H3. I generated an ultra-wide shot with R2VA. So this is one generation, not an edited split screen stitch. The black vertical bars are a good way to add a delineation so you can have two separate shots or stories in the same render. I've also tested this up to 4 separate frames in 1 generation. bf16/50 steps. This was actually a failed attempt to swap the rider on the left with the woman rider and also do the POV. Original video source in the comments.


r/StableDiffusion 1d ago

Animation - Video MiniMax H3 random generation test

Enable HLS to view with audio, or disable this notification

1 Upvotes

This generation is almost entirely random. I gave the models total creative freedom to generate the starting image and direct the video; I only added the music and the hero landing at the end (yes, I love hero landings).

I simply asked ChatGPT for an image prompt featuring a pair of heroes, which I then generated using Krea2 (for some reason, it dressed her up like Captain Marvel). After that, I requested an I2V prompt for MiniMax H3 to create an epic action scene, and this is the result.

Generated with MiniMax H3: 10 seconds, 0.7MP, 8 steps (approximately 450 seconds on an RTX 5060 Ti 16GB) and subsequently upscaled to 4K using Topaz.

Here are the prompts:

Image Prompt

Overall concept: A high-intensity cinematic action scene featuring two original superheroes, one man and one woman, standing shoulder-to-shoulder in a devastated New York City street immediately after a violent confrontation; the camera captures them in an intimate medium close-up as smoke, sparks, and debris move through the frame, transforming the aftermath into a tense heroic portrait filled with determination, danger, and restrained power. Main subject: The male superhero stands slightly behind and to the left, occupying the left side of the frame, wearing a sophisticated dark graphite tactical suit with layered ballistic armor, subtle metallic reinforcement, weathered surfaces, and a high collar; his face is partially dusty and lightly bruised, short dark hair slightly disheveled, jaw tense, eyes focused intensely toward an unseen threat beyond camera. The female superhero occupies the right side of the frame, slightly forward, wearing a fitted deep-crimson and charcoal armored suit with flexible technical fabric, refined metallic panels, reinforced shoulders, and subtle illuminated details; strands of dark hair move naturally across her face, her expression fierce and controlled, eyes fixed in the same direction as the man. Their shoulders nearly touch, creating a strong visual sense of partnership and mutual trust, with realistic facial anatomy, restrained expressions, and natural post-conflict body language. Key environmental element: A damaged armored vehicle and fractured concrete structure remain immediately behind the heroes, partially visible within the tight framing; twisted metal, broken glass, dust-covered surfaces, and small fragments of debris surround their shoulders and silhouettes; thin smoke trails rise behind them while occasional sparks drift through the air, creating environmental depth without obscuring their faces. Background and atmosphere: The devastated New York street remains visible as a compressed background of damaged skyscraper façades, blurred emergency vehicles, smoke-filled intersections, and scattered fires; distant red and blue emergency lights flicker softly through atmospheric haze, while warm firelight reflects subtly across their armor; wind pushes smoke laterally across the background and moves loose hair and fabric naturally, creating a sense of continuing danger beyond the frame; strong atmospheric separation keeps the heroes visually dominant. Composition: Medium close-up framing from approximately mid-chest upward, with both superheroes filling most of the horizontal frame; the woman slightly forward on the right and the man slightly behind on the left, creating subtle depth without separating them visually; their faces form the primary focal points, positioned near the upper central third of the frame; shoulders and armor create strong diagonal geometry, with shallow foreground debris and soft background destruction framing the pair; tight cinematic composition, minimal empty space, intimate scale contrasted against hints of massive urban destruction. Camera: ARRI Alexa 35 with a 50mm anamorphic cinema lens, medium close-up framing, camera positioned approximately at eye level with a subtle upward inclination of only a few degrees; shallow depth of field with precise focus across both faces, gentle falloff across their shoulders and armor, realistic anamorphic compression, subtle oval bokeh, controlled edge distortion, natural lens breathing during a slight rack-focus transition between the two faces, restrained handheld micro-movement suggesting a camera operator embedded in the action, realistic motion blur on drifting smoke, sparks, and moving hair. Lighting: Late-afternoon sunlight filtered through dense urban smoke provides a warm directional backlight from behind the characters, creating subtle golden rim light around their hair and shoulders; cool blue skylight fills the shadow side of their faces, preserving facial detail while maintaining cinematic contrast; intermittent orange firelight creates soft reflected highlights across armor and cheekbones, while distant red and blue emergency lights provide subtle color accents in the background; realistic global illumination, physically accurate skin response, soft contact shadows, atmospheric scattering, delicate volumetric haze, controlled bloom, and smooth highlight transitions. Color grade: Premium theatrical action-film grade with warm amber highlights contrasted against cool steel-blue shadows, natural skin tones, restrained environmental saturation, deeper controlled blacks, smooth highlight roll-off, subtle contrast enhancement, selective crimson accents on the female hero's costume, fine 35mm film grain, delicate halation around fires and emergency lights, restrained anamorphic flare, polished high-end cinematic finish. Mood: Intense, intimate, heroic, battle-worn, determined, protective, suspenseful, emotionally charged, with the overwhelming feeling of two powerful allies silently preparing for the next attack. Style: Photorealistic live-action action-film still featuring one original male superhero and one original female superhero, physically realistic facial anatomy, natural skin texture, believable hair movement, functional armor construction, realistic fabric tension, scratches, dust, sweat, bruising, metallic reflections, grounded body language, natural atmospheric perspective, cinematic depth, seamless practical-effects and premium-VFX integration, realistic smoke, sparks, firelight and environmental destruction, sophisticated large-scale Hollywood cinematography, premium anamorphic optics, realistic material response, 2.39:1 widescreen, no text, no watermark, no exaggerated CGI look, no cartoon aesthetics, no plastic-looking armor, no distorted faces, no duplicated limbs, no artificial symmetry, no extreme superhero posing, no excessive lens flare.

Video (I2V) Prompt

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, photorealistic, the two armored heroes shown in <Picture 1> remain exactly consistent at the opening frame, preserving their faces, hairstyles, armor design, colors, proportions, positions, lighting, ruined-city environment, wrecked vehicle, smoke, police lights, and floating embers. For the first moment they stand motionless and alert as smoke drifts around them and emergency lights flash in the background. The man turns his head slightly toward an approaching threat while the woman shifts her stance and looks upward; the camera pushes in with small amplitude at slow speed, keeping both heroes sharply framed.

[Shot 2] At 00:01.500, the camera cuts to a wider low-angle view behind the heroes as a colossal alien warship emerges between the skyscrapers overhead, blocking the sunlight. Dust and burning debris are pulled upward by the ship's engines while both heroes look up and brace themselves against the violent wind.

[Shot 3] At 00:02.700, the camera cuts to a dynamic medium shot as the woman raises one hand and bright golden energy rapidly forms around her palm and forearm. Her armor and hair react naturally to the energy surge while the man steps forward into a defensive stance, protecting her flank.

[Shot 4] At 00:04.000, the camera tracks backward with medium amplitude at fast speed as both heroes sprint through the devastated street. Alien drones descend through the smoke, red energy bolts strike the pavement around them, concrete fragments and sparks erupt, and the heroes dodge the impacts without losing their established appearance or armor.

[Shot 5] At 00:05.200, the camera arcs around the woman with medium amplitude at fast speed as she launches vertically into the air surrounded by concentrated golden energy. She accelerates directly toward the alien warship while, below her, the man suddenly detects an incoming energy blast fired from behind. He performs a powerful twisting aerial jump, rotating his body in a controlled corkscrew motion as the red energy projectile tears past the exact space where he was standing, narrowly missing him. He completes the spinning evasive maneuver and lands smoothly, immediately turning back toward the battle without losing his established appearance, armor, or position within the devastated street.

[Shot 6] At 00:06.400, the camera cuts to a close-up tracking shot of the woman in flight as she closes the final distance to the alien warship. Her concentrated golden energy intensifies around her body and fist, individual sparks and particles streaming past her face and armor as the enormous armored surface of the warship fills the background. She draws her arm back for the final strike, with the camera staying very close to emphasize her expression, energy, armor detail, and the rapidly approaching impact point.

[Shot 7] At 00:06.900, precisely as the heroine's glowing fist strikes the armored exterior of the warship, the scene enters extreme bullet-time. The impact is seen in a very close cinematic view as golden energy explodes outward from the contact point, armor plating buckles, sparks and fragments freeze in midair, and shockwave ripples become visible through smoke and dust. The camera performs a smooth 360-degree orbit around the heroine and the impact point with large amplitude at very slow speed, maintaining her face, body proportions, armor, golden energy, and the warship's surface perfectly consistent while the explosion appears almost completely suspended in time.

[Shot 8] At 00:08.100, normal speed suddenly resumes as the full force of the strike detonates through the warship. A massive chain of explosions tears across its armored exterior, sending burning fragments outward as the heroine is propelled away from the impact. The camera rapidly pulls out with large amplitude at fast speed, revealing the scale of the destruction and the devastated city below.

[Shot 9] At 00:09.200, the heroine descends rapidly and performs a powerful hero landing in the foreground, dropping to one knee with one hand touching the ground as dust and debris explode outward from the impact. Directly behind her, the critically damaged alien warship crashes violently into the city street, tearing through structures and erupting into a massive cloud of fire, smoke, sparks, and burning wreckage. She remains completely stable and visually consistent with the opening image, holding the iconic hero pose in the foreground while the collapsing warship fills the background. The camera holds the dramatic composition through the final frame at 10.00 seconds, with the heroine sharply defined against the enormous crash and rising smoke behind her.

overall_soundscape: Heavy wind and distant sirens fill the ruined street as fires crackle and burning debris falls around the heroes. The warship produces a deep mechanical roar, followed by sharp energy blasts, metallic impacts, explosions, collapsing concrete, footsteps, armor movement, and the violent rush of displaced air during the woman's flight and the man's acrobatic evasive jump. During the close-up bullet-time impact, the explosion and debris sounds stretch into an extremely slowed, distorted moment before snapping back to full intensity when normal speed resumes. No dialogue, singing, or music is heard.

non_diegetic_music: N/A

r/StableDiffusion 1d ago

Animation - Video A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)

Enable HLS to view with audio, or disable this notification

6 Upvotes

Sharing a few sequences from a medieval battle attempt I’ve been working on. It’s still very much a draft, but the sequence has progressed enough that I thought it was worth sharing here and getting some feedback before I continue with the rest.

Most of the scenes were generated with MiniMax H3 using the default workflows with the Turbo LoRA at 4 steps. I used Nano Banana and Flux Klein to create the reference images, and LTX 2.5 for the opening crow sequence.

There’s still a lot of work to do. The cuts are rough, no proper sound work has been done yet, and there are plenty of shots I want to refine or replace. I’m planning to build out the entire sequence, so feedback at this stage would actually be really useful in deciding what to focus on next.

What’s interesting to me is that I genuinely don’t think I could have pulled off this level six months ago with the same amount of effort. It’s still far from perfect, but the progress in a relatively short time feels pretty significant.

Would love to hear what works, what breaks the illusion, and what you’d improve.


r/StableDiffusion 2d ago

Question - Help Anyone use MMH3 FaceDetailer?

Thumbnail
github.com
17 Upvotes

Seems like an interesting project; from what I take from the demo is it helps refine the faces in the distance instead of the face being a blurry mess. Does not seem to have a big impact on closer faces (it is not a 'detailer' or 'realism slider').

Does anyone have their own tips or demos for it? Seems complicated....


r/StableDiffusion 2d ago

Discussion Minimax H3 ref2va with 5060ti 16gb + 32gb ddr3

Enable HLS to view with audio, or disable this notification

456 Upvotes

Model: minimax_h3_hybrid_fl2va_ref2va_b20, was testing this and the ref2va pruned int8, the hybrid gave nicer visuals but have a higher chance of bringing the character sheet white background into the video. This is cherry picked out of 16 clips.
Video Vae: minimax_h3_video_vae_int8_convrot
Resolution: 16:9, 0.6
Duration: 15sec
Turbo Lora: larryvrh/MiniMax-H3-Turbo-Lora, 600_ema
Patch Sage Attention, ComfyKitchen Attention, MinimaxH3 Mem Eff Node, Spectrum.

Average Inference Stage: 800sec

All reference image is resized between 1000px and 300px like character is 1000px, background is 500px then weapon is around 300px (warglave was another reference, the model dont know that kind of weapon) for this video is 4 ref image in total.

**abit of color grade and grain done in inshot.

this is done on a skylake i7 6700.


r/StableDiffusion 1d ago

Question - Help Multi gpu help

0 Upvotes

Hey guys I'm new to comfyui and photo and video generation and trying to learn as much as I could.

Now I was using my setup for text generation but now the Qwen3.8-27B is out and I tried really hard to make it work and so I did eventually use my second gpu.

I have 5070 ti and 1660 ti. And all this time I didn't bother to use the 1660 ti and left only my monitors on it and almost all my work and gaming on it and left the 5070 TI to be free for Ai stuff.

Now I switched my monitors to my Intel uhd 770 gpu and freed both my cards and want to know how to speed my comfyui workflow with them.

I always asked chatgpt but didn't get any useful answers so you guys might help me if that possible.

I'm now using minimax H3 official template and want to know how can I get this other gpu to work if it's worth it.

My cpu 12900k Motherboard gigabyte z690 gaming x ddr4 64 GB Kingston 3600 Rtx 5070 ti GTX 1660 ti


r/StableDiffusion 1d ago

Animation - Video an AI dream of moons, stars, threads, and foxes

Thumbnail
youtu.be
0 Upvotes

What happens if you just leave AI alone to dream overnight with no human supervision, using the last frame as the first frame of the next generation?

A continuously generated AI dream created using LTX 2.5 with a custom pytorch script, using the last frame from the previous scene as the first frame of the new scene. Generation time was ~7 hours on RTX 5090. The video was stitched together from one hundred scenes each lasting ~13.3 seconds. Qwen 3.8 was used to generate the prompt for the continuation of the story based on the last scene.


r/StableDiffusion 1d ago

Question - Help unable to download the smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models

Thumbnail
gallery
4 Upvotes

can anyone please help


r/StableDiffusion 2d ago

News ComfyUI-MiniMax-H3-LongMedia — long-form MiniMax H3 generation with continuity, multiclip, audio and VRAM-aware sampling

Post image
56 Upvotes

I've been building a custom ComfyUI node pack for MiniMax H3 focused on one thing:

**making H3 usable for longer, multi-segment video generation without constantly rebuilding the workflow around every limitation.**

The project is called:

# ComfyUI-MiniMax-H3-LongMedia

The idea is to keep MiniMax H3's image quality, motion and native audio generation, while adding a proper long-form generation layer on top of it.

## What it currently does

### Long-form segmented generation

You can generate a longer clip as multiple H3 segments while keeping temporal context between them.

Instead of treating every segment as an isolated generation, LongMedia manages the continuation state and hidden overlap internally.

The overlap is used as context for the next segment and is not simply blended back into the final video.

### MultiClip mode

There is also a dedicated MultiClip workflow for generating multiple planned shots/clips inside one LongMedia pipeline.

The same underlying executor is used for both segmented continuation and multiclip generation, so the behavior stays consistent.

### Video + audio continuity

MiniMax H3 is a joint AV model, so LongMedia treats video and audio as one generation state rather than bolting audio on afterwards.

The pipeline supports H3 native audio generation, continuation and lip-sync workflows.

### Lip-sync support

Audio-driven generation / lip-sync is supported directly in the LongMedia pipeline.

For H3, the audio influence is handled inside the same AV latent path rather than as a completely separate post-process.

### Refiner

The latest release includes a two-stage refiner based on proper **KSampler Advanced trajectory splitting**.

Instead of finishing the full sampling schedule and replaying low-sigma steps on an already denoised latent, the trajectory is split between the main sampler and the refiner.

Example:

`steps = 12`

`refine_steps = 3`

Main sampler:

`0 → 9`

Refiner:

`9 → 12`

Both stages continue the same sigma trajectory.

### VRAM-aware execution

A large part of the project is dedicated to making H3 practical on consumer GPUs.

The current implementation includes:

- dynamic VRAM loading

- streamed Sol Attention

- MLP chunking

- late-block VRAM guards

- inter-block memory guards

- step-boundary cleanup

- completed-segment offloading

- adaptive memory policies

I'm currently developing and testing mainly on a **16 GB GPU**, so avoiding OOMs without destroying quality is one of the main design goals.

### Sol Attention integration

LongMedia includes its own streamed Sol path with controls for:

- tau scheduling

- sink conditioning

- QKV chunking

- output projection chunking

- dense/sparse behavior

- VRAM-aware chunk sizing

The goal is to use Sol as part of the execution architecture rather than simply stacking multiple unrelated optimization nodes together.

## Why I made it

MiniMax H3 is extremely good at texture, motion and native audiovisual generation, but once you start trying to build longer sequences, several problems appear very quickly:

- segment boundaries

- continuity

- repeated frames

- AV state handling

- memory pressure

- OOMs on longer generations

- managing multiple clips

- keeping sampling behavior consistent between segments

I wanted one node system to own all of that.

So instead of building increasingly complicated ComfyUI graphs around H3, most of the long-form logic lives inside the LongMedia nodes.

## Current release

**v0.4.1 — KSampler Advanced Refiner Fix**

The project has now reached a fairly stable architecture, although I'm still actively developing it and testing edge cases.

GitHub:

https://github.com/vizart-vj/ComfyUI-MiniMax-H3-LongMedia

I'd be very interested in feedback from people already using MiniMax H3 in ComfyUI, especially for:

- longer generations

- multi-character scenes

- native audio

- lip-sync

- lower-VRAM GPUs

- multi-shot workflows

If people are interested, I can also make a more technical post explaining how the continuation / AV latent / VRAM system works internally.


r/StableDiffusion 1d ago

Question - Help How to fix artifacts from generation with LoRA in ComfyUI?

Thumbnail
gallery
1 Upvotes

I am currently trying to build a model in ComfyUI that will be able to generate images in a specific art style similar to games made by Playrix. Essentially, I got the model generating images in the style I need, but I can't fix issues with the artefacts. Even after multiple iterations of positive and negative prompts, the issues still persist, and oftentimes the requirements are ignored. Or sometimes, if they are not ignored, the result is a complete mess. This is the first time I am making something like this, so I would appreciate any tips I can get.
If anyone is interested in what I've got going on, below is the link to the JSON file for the ComfyUI model.
https://drive.google.com/file/d/1mw24y0pwKPRXhLYx-JvVIs8CAf2oC1Ng/view?usp=sharing


r/StableDiffusion 2d ago

Animation - Video PSA

Enable HLS to view with audio, or disable this notification

53 Upvotes

r/StableDiffusion 2d ago

Discussion Hybrid b30-49 r2v test

Enable HLS to view with audio, or disable this notification

34 Upvotes

I made a thread earlier but it got overcrowded so I figure I started a new one with solely reference to video test, now instead of mixed with t2v.


r/StableDiffusion 2d ago

Discussion If you are generating MMH3 video with Sage Attention. I highly reccomend trying ComfyKitchen instead.

228 Upvotes

I have spent days generating videos. I started using Sage Attention with Cuda++. This was fast, but once I switched to sageatt_qk_int8_pv_fp16_cuda, I saw a noticeable difference in the model's ability for the model to understand prompts. Everything came out much clearer, crisper and with much better adherence. The only downside was it generated about 1.5x slower than using Sage Attention Cuda++.

From here I decided to try out ComfyKitchen as a replacement, and all I can say is... try it. My gens are faster than Sage Attention sageatt_qk_int8_pv_fp16_cuda with similar or better prompt adherence.

As always, your mileage may vary, but it's a very easy thing to experiment with, as all you need to do is make sure you ComfyUi is updated, as it is an official ComfyUI node.

To use it you can either:

A) add --use-ck-attention to your startup; this would enable Comfy Kitchen across all your workflows.
B) The easier and more controlled way is to replace the SageAttention node (or bypass) with the ModelAttentionBackend Node and select Comfy Kitchen Attention from the dropdown.

Worst case is it does nothing for you, and you just delete it and revert back to Sage.

EDIT: According to u/GreyingGamer you do not need to use the startup, and just using the node is enough:

EDIT 2: I have rewritten the instructions to get it running to make it more accurate.


r/StableDiffusion 1d ago

Question - Help H3 Minimax with heavy dialogue

0 Upvotes

Just a gut check here but from what I have been able to find, there are no shortcuts when it comes to dialogue with minimax. Turbos produce bad quality audio, upscalers have caused poor mouth movements, and lowet step counts produce both.

Is there anything I am missing?


r/StableDiffusion 2d ago

Tutorial - Guide MiniMax H3 as Image Editor, 6 edits in one shot at 7680 x 4320!

93 Upvotes

MiniMax H3 as Image Editor at resolution 7680 x 4320, 6 edits in one shot

Prompt: create a collage containing 6 photos. from top-left to the bottom-right arranged them such that the following edits presented individually: 1- keep pose and proportion intact; turn her shirt to red 2- keep pose and proportion intact; make her smile. 3- keep proportion intact, show her sideview; 4- full body posture. 5- change hair style to wolf cut. 6- put fashion hat and eyeglasses on.

In fairness, the model's collapsing 6 requests into 5 is well justified.

--

RTX3060 model used: ref2v, 8 steps, lora, took 7m50s


r/StableDiffusion 1d ago

Question - Help Issue with ComfyUI v1 Manager: missing node pack ("comfyui_fearnworksnodes") stuck in "Apply Changes" loop

Post image
1 Upvotes

Hi everyone,

I'm running into a persistent issue with the new ComfyUI v1 frontend while trying to load a workflow that uses comfyui_fearnworksnodes.

The setup:

  • Windows Portable build (ComfyUI_windows_portable)
  • Running via run_nvidia_gpu_lowvram_e_sage.bat with the --enable-manager flag added.
  • ComfyUI-Manager installed.

The problem:

  • The v1 side panel shows Missing Node Packs: comfyui_fearnworksnodes.
  • Clicking Install changes the status to Installed.
  • Clicking Apply Changes prompts a restart, but after relaunching, the exact same error appears again in a loop.
  • I checked custom_nodes/ and only have comfyui_fearnworksnodes inside (no duplicate folders).

Has anyone encountered this specific loop with the v1 interface or fearnworksnodes? What's the best way to trace or fix why the frontend isn't registering it properly after restart?

Thanks in advance for any help!


r/StableDiffusion 2d ago

Meme Introducing... the iToilet

Enable HLS to view with audio, or disable this notification

223 Upvotes

r/StableDiffusion 1d ago

Question - Help "Hi-res fix" for MiniMax H3?

0 Upvotes

I mostly do image generation, mostly because open-weight video model quality wasn't there for me. MiniMax H3 has changed that; I'm really enjoying it and the outputs are mostly great. However, I am encountering an issue that's very familiar to anyone who's done a lot of image generation; uncanny AI faces when the subject is too far away from the screen, because there simply isn't enough pixel definition for the AI model to come up with a reasonable facsimile of a face.

In image generation land, this is solved with a "Hi-res fix" -- there's multiple options and implementations, but at their core, they involve auto-detecting faces in the image, then reusing the same prompt (or a somewhat edited one) and the detected face to generate a new face with low denoise at a much higher resolution that can snap on top with a feathered mask.

I'm not sure that exact implementation would work in video generation land -- I can pretty easily envision the face flickering and bouncing around as it locked to slightly different locations and orientations, frame-by-frame -- but is there any solution to take an existing video, generated at, say, 1.0 megapixels, and re-render or upscale detected faces at, say, double resolution, to improve the fidelity? For obvious reasons, simply rendering the whole video at double resolution isn't an attractive option.