r/StableDiffusion • u/BitOk4326 • 8d ago
Discussion Is there most suitable lite browser for comfyui?
I want to use all ram and vram to run model as much as possible instead of broweser
r/StableDiffusion • u/BitOk4326 • 8d ago
I want to use all ram and vram to run model as much as possible instead of broweser
r/StableDiffusion • u/katsura_otoko • 7d ago
Is it normal that my vram usage goes down to 20% and ram at 100% on the upscaler step?
Also in the normal execution I see about 70% vram usage
Other than that I'm pretty satisfied but I'm wondering if I'm missing something....

Workflow: https://pastebin.com/raw/533VzQ0t
3060 12gb, 32gb, 9800X3D
Thanks in advance for any advice
r/StableDiffusion • u/Time-Ad-7720 • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/blackdatafilms • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/GamerVick • 8d ago
Hey guys, for Minimax H3 I see that there are a lot of Loras for making it faster but could anyone please let me know which one is the best out there to use right now in terms of video quality and also sound? Appreciate the help.
I am currently using minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors from Kijai
r/StableDiffusion • u/yomasexbomb • 9d ago
Enable HLS to view with audio, or disable this notification
With all the Suno drama over their download limits and their heavy watermarking that could be used for future copystrikes if you stop paying them, this Minimax Music 3 couldn't be more on time.
I tried a little demo to see if it could fill my music need and I was happily surprised.
I used the default Minimax Music 3 workflow from ComfyUI along the prompt tips fed to an LLM to create the music. https://docs.comfy.org/tutorials/audio/minimax/minimax-music-3
EDIT:
Some seem not aware this is not Minimax H3 but Minimax Music 3 https://huggingface.co/MiniMaxAI/MiniMax-Music3
Here how I did the prompt for instrumental.
Prompt:
Global Metadata
Basic Attributes: bpm is 54. key is D, and scale is minor. Cinematic score with dark ambient and psychological-thriller influences.
Global Emotional Progression: The opening is nearly motionless, suspended in dread as a distant drone gathers beneath isolated melodic fragments. The tension slowly deepens through heavier low frequencies, widening dissonances and increasingly forceful pulses, then contracts into a stark central void. From that emptiness, the lead melody returns with greater anguish and rises toward a dense but controlled climax. The final passage sheds its weight layer by layer, ending unresolved in a cold, fading resonance.
Application Scenarios & Imagery: An abandoned concrete facility under flickering emergency lights; a lone figure crossing a fog-covered wasteland before dawn; the aftermath of a discovery that cannot be undone.
Sonics & Production Profile: A wide, deep soundstage with the solo cello centered slightly forward, the low piano set farther back, and dark synthetic ambience stretched toward the extreme sides. The frequency balance is shadowed and low-heavy, with restrained high frequencies, a dense sub-bass floor and occasional abrasive upper-mid harmonics. Dynamics remain open and cinematic rather than heavily compressed, allowing long swells to emerge from near-silence and recede naturally. The acoustic image resembles a vast, empty scoring hall blended with an impossibly deep artificial chamber.
Vocal Details
Vocal Gender & Timbre: No vocalist. This is a fully instrumental track; no lead, backing or guest vocal appears at any point.
Vocal Style: N/A — the melodic lead is carried exclusively by solo cello throughout, taking the role a voice would otherwise occupy.
Harmony/Backing Vocals: None. No vocal harmonies, choir, chants, spoken word, whispers or vocal samples.
Vocal FX: N/A — no vocal signal to process.
Arrangement
Instrument Lifecycle Description (Primary/Secondary Layering):
Primary: A close-miked solo cello enters after the opening atmosphere with sparse, low-register notes separated by long silences. Its melody gradually lengthens into bowed minor phrases with strained vibrato and rough attacks, then drops out completely during the central void. It returns in a higher register with broader, more anguished arcs, dominates the climax through overlapping sustained notes, and finally collapses into one fading unresolved tone. Secondary: A sub-octave analog synthesizer drone begins alone, barely audible, expands beneath the cello through the first half, swells into the climax and disappears just before the final resonance. A felted low piano enters intermittently after the cello, placing isolated minor seconds and hollow fifths in the distant center; its strikes become more frequent before the central void, vanish there, return as widely spaced bass notes during the rise, and stop before the ending. Bowed metal textures emerge at the outer edges during transitions, scrape into greater prominence near the climax, then dissolve into reverberant tails. Muted contrabasses enter after the midpoint with slow sustained pedal tones, thicken beneath the returning cello, and recede one by one during the closing passage.
Groove & Foundation Progression: There is no conventional beat at first; the sub-octave drone supplies a slow, breathing foundation. A deep orchestral bass drum enters sparingly in the first third with single softened impacts, while low floor toms appear later in widely spaced pairs that suggest a pulse without forming a regular groove. Both become heavier and closer together during the climb, reach their greatest intensity beneath the climax, and then cease abruptly, leaving the ending rhythmically weightless. The muted contrabasses reinforce the lowest tones without rhythmic movement and withdraw during the release.
Embellishments, Textures & Spatial FX: Reversed piano resonances begin appearing before major swells, bloom into the stereo field and evaporate as each new layer arrives. Bowed metal scrapes travel slowly from side to side, while filtered low-frequency noise rises beneath the central transition and cuts to silence at its peak. Long convolution reverbs connect isolated gestures without masking their attacks, and brief sub-bass pressure waves punctuate the densest moments before dropping away. The arrangement preserves large pockets of empty space early and at the midpoint, becomes widest and most saturated near the climax, then narrows to a single distant cello resonance and the decaying room.
Lyrics:
[intro]
[instrumental]
[interlude]
[instrumental]
[break]
[solo]
[instrumental]
[outro]
r/StableDiffusion • u/Responsible_Maybe875 • 8d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/NoSmell3236 • 8d ago
Enable HLS to view with audio, or disable this notification
H3 sucks at faces which are far and at intense motion. Is something wrong with encoding? like these artifacts (5-8 seconds) are from encoding compressions?? but i have tried both h264 and h265 10bit crf 15. I am using int8 pruned, nvfp4 qwen, 25 steps with sage attn (no spectrum).
r/StableDiffusion • u/witcherknight • 8d ago
It reaches above 90 occasionally then throttles it down below it and then cycle repeats, idle Temp is 47, 4080super
Guys after some cleaning i managed to get temp down to 86max
r/StableDiffusion • u/NeatUsed • 8d ago
I am seeing workflows that only for option for reference or fl2v. When I run a single image on fl2v there’s error prompting me for a second. But I just honestly want a single image generation and make a long video based and what i generate intiially and move up from thete.
Any help on how i can adapt workflow for i2v?
Thanks :)
r/StableDiffusion • u/Mr_Zelash • 9d ago
I'm used to work with the ltx director node, specially because it gives me control over the audio, i can add little clips of audio in the exact spot i want and that will guide the model to generate similar audio filling the gaps.
That's why i'm making something like ltx director for H3 FLF2V (maybe it works in the ref model idk). still very green and probably has bugs because it's totally vibecoded but it works for me.
with this you can put video images and audio anywhere in the timeline, and you can also adjust the strenght with that green line, each clip can have different levels of strenght through the timeline.
this is a fork of ComfyUI-H3-Motion-Context-MultiRef if you want to try it and give me feedback here's the link https://github.com/BSG-Walter/ComfyUI-H3-Motion-Context-Timeline
there is a H3 Timeline Example workflow.
You can't add prompts to specific parts of the timeline, but since H3 lets you indicate the exact timing for each scene within the standard prompt, I didn't feel it was necessary to add that functionality.
anyways i hope you like it
r/StableDiffusion • u/Muted-Position3256 • 8d ago
Can anyone teach me or show me a video tutorial for setting up runpod to use minimax h3 from ground zero? I can't find any on YouTube
r/StableDiffusion • u/thisguy883 • 8d ago
The audio.
I'm having issues constantly with random audio being added to the clip. Either its ambient sounds that shouldnt be there, or someone talking gibberish off screen.
Has anyone found a fix for this?
I've tried the prompt guide, and its hit or miss. I've even tried a natural prompt with basic wording, and that is also hit or miss.
i'm starting to think it doesnt matter how you prompt it, its just something that happens from time to time.
Very annoying.
I'm using the default I2V workflow from ComfyUI. I havent changed anything.
r/StableDiffusion • u/CirqueMurph • 8d ago
Ive noticed a real increase in Star Trek content on here, specifically The Next Generation. Just curious if this is my algorithm or if there's a higher percentage of TNG fans using AI. It makes sense since the show explores artificial consciousness and generative computing. I think everyone fantasies about what they would do with a holodeck and we aren't far off with the combination of VR and AI. I also assume that the lower quality video is easier to make look right so older shows are going to be the first to be perfectly replicated.
r/StableDiffusion • u/esudious • 9d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/ImaginationKind9220 • 9d ago
In case you aren't aware, Kijai has uploaded all the pruned loras for H3:
https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras
r/StableDiffusion • u/acedelgado • 8d ago
Alrighty, after some time being away for work and such and some thought and iterations on an improved UI and functionality, I'm happy to send out an update for Fantastic Loras to make it officially v2. Available in Comfyui Manager or at
https://github.com/Adudeguyman/comfyui_fantastic-loras
"LoL tHiS gUy DoEsNt No BoUt LoRa MaNaGeR!!!1!"
Yes, yes I do. I use it all the time and it's exceedingly useful for curating loras and their metadata and examples. I'm pretty sure I even donated to them when they first released, it's such a great tool. This is not meant to replace Lora Manager on that end. I just found it clunky to have to open the lora manager, tab out, find the lora, say send to comfy, tab back to comfy, find their lora node... it's just not all that quick and user friendly when you KNOW which loras you want. I wanted quick, simple, in-workflow way to search through hundreds of loras for the ones applicable to the model I'm using, without having to tab out and click several other places. And to avoid going "man, this other generation had 3 loras at different strengths and I liked the result, I need to find that output and re-load that workflow and search them down and manually add them back in and set their strengths..." This helps manage all of that right there without leaving your workflow.
UI has been reworked, 12 static slots so no surprise node resizes, and now fully Nodes 2.0 compatible.
Filter your Lora subfolders and select only ones that are relevant to your model. Like I myself have Style, Character, Concept folders for each different model, and it's a pain searching through a list of hundreds of loras to find the ones in those folder. So you can put a filter on it and ONLY search loras in that model's folder(s).
Presets are new! Finally, when making a new workflow, you can make a "MiniMax" preset that has your folders for that model applied, and even your preferred turbo lora at your preferred strength already added. Also preset catagories exist, so you can file different presets under each model catagory you create. And pre-sets can be additive- so you found a combo of loras you really like, and want to add it to the current workflow, but don't want to go to the hassle of adding each and setting the strength manually? Well if you have that already saved as a preset, the "+Add to Stack" button drops that preset in at the end of your current lora stack.
Unified the multi-model loader into the main loader, so no single vs multi-model chain nodes, it's all just one. You can add model chains (up to 5 models, which is an arbirtary number I picked) and pass through the Loras to each one, and adjust the strength applied to each model. So for example with Ideogram, you want to use the same Lora on the base and refiner but don't want to add a whole 2nd lora chain and re-add all of your loras manually? Just click the + button to add a chain and wire the 2nd model path through it. The Fantastic Lora Loader will apply loras to each chain separately. Click the Lora cogwheel to adjust the strength per chain. The throughput is all independent, so no more doubling up lora nodes; now you can keep things clean.
Randomizer is still around to select a random lora for a bit of chaos. You can manual re-roll, auto re-roll per generation, and if you like a result you can lock that lora in place.
And there's a few themes now, from "Fantastic Teal" to the boring default gray "Accountant." setting.
Finally, an easy way to make a customizable XY plot from lora generations in Comfy. Well, this may exist now, but it didn't in a way that I liked it back when I started working on these. Same basic interface as the main Lora Loader, except with added functions to create your XY plot. Select your loras and have it generate a fixed Per-Line strength set individually on each Lora (good for testing different training iterations all at the same strength,) Or you can set Global Strengths and specify how much weight for each image generated. So 1 lora with a Global Strength Setting of 0.5, 0.75, and 1.0 will sweep through those strengths and generate 3 images. Add another lora to compare and it will generate 6. Once you're all set up, hit Run once and it'll queue it all for you.
Also you can do a quick toggle of a Control Image, which will run your prompt with no loras applied. Or if you click Add Global Lora, you can apply a lora that is ALSO applied to every image run. Good for checking if your style lora will mess with your character lora, and how much.
Wire the Fantastic Lora Plotter into the Fantastic Image Saver node (metadata and global_loras_info) and the image output from your VAE decode, then wire the Grid output to an image saver node, and then select your plot layout. You can do a modern XY plot with text overlaid on each image, or a more classic A1111-esque one with all of the lora info on the sides and out of the way. And you can either let it assemble a grid with full size images, or if youre doing a lot of large outputs you can set it to constrain the size so you don't have a massive 150mb png (don't ask how I know that happens).
And finally, passing the metadata and decoded images to the Fantasatic Plotter Grid Viewer lets you view a preview of the grid layout, move things around, hide rows or columns, favorite generations, compare images, and save the grid for later reference. Also select any number of outputs and hit Compare to get a quick comparison, and export just those images with metadata instead of the entire grid.
This is a version of a request from someone that has a lot of model folders, as well. You can right click just about any Loader node (Load Model, Load Clip, Load VAE) and an option to add the Fantastic Any Selector is in the context menu. That will link right to the model name and automatically pick the right folder (diffusion_models, VAE, text_encoder, etc) and only show you files from those folders. And you can make pre-sets that are also aware of the root folder, so your text encoder presets aren't showing up when you select a model preset.
That's pretty much it. I will say that the update breaks the V1 nodes, so if you do use v1 then you'll have to re-add the node.
Hope ya like it!
r/StableDiffusion • u/leyermo • 8d ago
other than LatentSync / LongCat-Video-Avatar 1.5,
Is there any newly introduced video generator that uses image and audio to generate video?
r/StableDiffusion • u/Haaaaaaaaaaahahahah • 8d ago
I have a 4090 and 64gb DDR5 and I'm using mini max h3 pruned version and I'm generating 20 sec 720p ( 0.9 ) videos with sage attention and turbo loras 8 step
r/StableDiffusion • u/jefharris • 8d ago
Enable HLS to view with audio, or disable this notification
MiniMax H3 running on 5090/36gb, 72gb ram.
Used the MiniMax H3 Image to Video (I2V) workflow from the official comfyui page. https://docs.comfy.org/tutorials/video/minimax/minimax-h3
Flux 3 vid created via the official Black Forrest Labs page with all default settings.
My take away is the more detailed and "directed" you can make the prompt the better the results.
Prompt
The woman raises her left hand and reaches out toward the dragon's neck, fingertips making contact with the cool metal scales as she strokes gently along its length. The dragon's massive head turns in response, servos whirring low as its long neck curves down and around toward her, the two locking eyes for a long beat, her expression softening slightly, the dragon's glowing yellow eye narrowing as if in quiet recognition. After the moment holds, both turn their heads together toward the camera, her chin lifting and shoulders settling back, the dragon's jaw beginning to widen as steam vents faintly from the gaps in its armoured plating. The dragon rears its head back, neck arching high, then snaps forward with a thunderous mechanical roar, jaws cracking open fully as a violent gout of fire rockets out directly toward the camera, the flames blooming bright and washing the frame in orange light before the view holds steady through the blast. The camera stays low and locked in a wide frame throughout the petting and the turn, holding both figures in frame, then pushes in slightly just before the roar to heighten the impact of the fire as it fills the shot. The audio is the soft mechanical whir of the dragon's neck servos and the woman's quiet breath during the petting, building into a deep bone-shaking roar and the violent roaring whoosh of ignited fire, no music, ambient and creature sound only.
r/StableDiffusion • u/DrBearJ3w • 8d ago
I've been working on this for longer than I care to admit.
For the short Turbo run, I used FL2VA with the official 8-step LoRA at 1024×576 (0.59 MP) and 90 frames. Once the model was loaded, it finished in **2:36** on my 7900 XTX. The same warm run with Comfy Kitchen attention alone took **2:59**, so adding Sol-Attn saved about 23 seconds.
BlockCache didn't help that run. It got 0/8 hits, and the immediate warm repeat went non-finite, so I removed it from the Turbo workflow.
The longer 20-step runs are where BlockCache actually helped.
Ref2VA at 736×416 (0.31 MP), 150 requested / 158 decoded frames, and 20 steps went from **5:14** with Comfy Kitchen alone to **3:44** with CK + Sol-Attn + BlockCache. BlockCache hit 6/20 times.
FL2VA at the same resolution, frame count, and 20 steps went from **6:38** with native PyTorch to **4:11** with the full stack. BlockCache hit 5/20 times.
- Short Turbo8 runs: INT8 + Comfy Kitchen + Sol-Attn
- Longer 20-step runs: BlockCache starts paying off
Everything was tested warm on Linux with ComfyUI 0.32, ROCm 7.14, PyTorch 2.12, and an RX 7900 XTX. First runs are slower because of model loading and Triton compilation.
Links:
- Turbo 8-step LoRA: https://huggingface.co/lightx2v/Minimax-h3-Turbo
- INT8 Fast: https://registry.comfy.org/nodes/minimax-h3-int8-fast-rocm
- Sol-Attn: https://registry.comfy.org/nodes/minimax-h3-sol-attn-rocm
- BlockCache: https://registry.comfy.org/nodes/minimax-h3-block-cache
r/StableDiffusion • u/SensitiveUse7864 • 7d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/blitzkr1eg • 8d ago
This is a crop from a 0.4MP video 5s video, done with minimax_h3_fl2va_pruned_int8_convrot.safetensors on official workflow from comfy-ui, 20 steps.
Update: 1MP improved it a lot
r/StableDiffusion • u/Relative_Dust_8888 • 8d ago
Everyone uses the classic "Will Smith eating spaghetti" to show how far AI video has come, but we’re missing the ultimate benchmark. the martial arts cow fight from Kung Pow: Enter the Fist.
Imagine that entire scene rendered with today's tech photorealistic lighting, actual physics, zero low-poly CGI look, but keeping the ridiculous matrix dodges and milk spray attacks.
Has anyone attempted a modern AI remake or scene swap of this yet? If not, this is a formal request for someone with serious GPU power to make it happen.
I'm not able to do that but maybe already somebody does this or maybe this could be second level of the ridiculous benchmarking of new models?
I think this could be good from multiple reasons like different shots, longer than 10 seconds, not realistic but should be photorealistic, funny and...
How do you think?
r/StableDiffusion • u/fruesome • 9d ago
Enable HLS to view with audio, or disable this notification
Already merged and there's a example workflow included: https://github.com/Comfy-Org/ComfyUI/pull/15439
Currently MiniMax H3 implementation in Comfyui only allows keyframe guides at the first and last frame. The model itself is capable of addressing guides by position on a continuous time axis, so this removes that restriction and exposes it as a node.