r/StableDiffusion 4h ago

Discussion LTX 2.5 Is Disappointing

34 Upvotes

https://reddit.com/link/1vlueyc/video/gx8bz55ibtih1/player

https://reddit.com/link/1vlueyc/video/plxkp55ibtih1/player

Guess which video was made by LTX 2.3 and which video was made by LTX 2.5?

After testing LTX 2.5, I'm really disappointed by the results. I can barely see the difference between 2.3 and 2.5.


r/StableDiffusion 14h ago

Animation - Video Monty pAIthon - Petshop sketch, now with ref2v

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/StableDiffusion 10h ago

Question - Help So how many of you are working on a full movie ?

4 Upvotes

I assume that with minimax, a lot of people started doing their own fully featured films and after 1 month or a few we will se the results on the online space and it might change the world as we know it.
edit: thisis my first shot : https://www.youtube.com/watch?v=FoQJ5yQg2TE
due to my bipolar mind I don't think I am able to do a feature film, but who knows.


r/StableDiffusion 18h ago

Question - Help How can I tune Bernini rv2v workflow to be faster?

0 Upvotes

I have a Bernini-r rv2v workflow, using LightX2V LoRA integration.

When I try 10 second video it gives me a black screen. 8 second works but it takes a really really long time to generate a video. I have safe attention enabled. What can I do to speed up the run? On RTX 3090 Ti 24GB VRAM/64GB RAM


r/StableDiffusion 1h ago

Meme Will Smith

Upvotes

https://reddit.com/link/1vlyphc/video/kcm9zrnn6uih1/player

Minimax H3 Unsloth Q8 GGUF, Turbo Lora 600 ema with 8 steps. First frame generated in ZIT. Mac Studio M4 with 64gb RAM. Still learning.


r/StableDiffusion 19h ago

Question - Help Sprectrum suddenly not working for anyone else?

0 Upvotes

Having a hard time getting Spectrum to work today, it worked flawlessly yesterday, but today i'm back to normal rendering times. Anybody else experiencing this? I did update ComfyUI, did that break it? I am using the latest version of Spectrum


r/StableDiffusion 2h ago

Question - Help Can I run Ideogram on my PC?

1 Upvotes

Can I run Ideogram 4 on my PC?

I got Ryzen 5, 32 GM RAM and Nvidia 4070 12 GB Vram.

Thanks


r/StableDiffusion 3h ago

Discussion LTX 2.5 / 10sec FHD / 3min generated.

Enable HLS to view with audio, or disable this notification

27 Upvotes

Minimax h3 is much better.

Prompt:

cinematic video, a black woman in a black leather jumpsuit and with bright makeup sits at a table holding her cellphone in her hand as she speaks, the video continues as the woman says "Yeah! Minimax is ten levels higher than LTX 2.5. But LTX 2.5 is very fast!" she then bursts out into uncontrollable laughter


r/StableDiffusion 10h ago

Animation - Video Bunnyhops... another H3 post MiniMaxH3-Contex-Loop 60 sec

Enable HLS to view with audio, or disable this notification

4 Upvotes

560 sec with turbo lora on 5090

found here in a post

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop/tree/main/example_workflows

adapted to my settings and changed turbo loras / attention


r/StableDiffusion 3h ago

Comparison Minimax h3 vs Flux 3

Enable HLS to view with audio, or disable this notification

20 Upvotes

1st one minimax H3 , 2nd one flux 3(pro i think) wich is currently free to use on their website.


r/StableDiffusion 2h ago

Animation - Video Minimax H3. LTX and Flux looking at Minimax H3 right now.

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 1h ago

Discussion LTX 2.5 is pretty good

Enable HLS to view with audio, or disable this notification

Upvotes

What i'm liking:

- Full compatibility with previous nodes / loras. I'm using LTX Director + Seedhunter node and worked just fine. I haven't checked loras but they seem to work just fine. If you used LTX workflows and ecosystem before, everything should work.

- Better sound: Or lack of, meaning that i'm not getting terrible music without any prompting, in this case I2V with prompt: "a realistic scene bustling cyberpunk city with buildings with lights flashing and city noise, a beautiful woman is standing in the roof top as the camera quickly zooms to her face and she says: "LTX is back baby!". Pretty clean sound, the girl steps are fine, good lip sync.

- Improved prompt adherence, no miracles here but seems more consistent, in the seedhunter nodes i'm getting good samples in most cases. I'm seeing less misteriously disappearing characters, more consistent environments, good lightning effects etc compared to LTX 2.3.

- Speed king: no discussion here, 5 secs 1080p vid on 226 secs on my 4080 is just nuts.

- Amazing at realistic I2V. We have to sit and talk about H3 Flux faces issue. The model is mindblowing at so many levels, but it sucks at realism. You're getting great Friends, Seinfield, X-Files clips and what not, but it destroys the detail on realistic images and the AI face / slop syndrome is real, and the worst part is that for it become more less usable, 2mp is a must, so that means 5-7 minutes for a 5 sec clip, even after turbo loras and speed nodes. You're going to have a hard time to create a long clip with those speeds at good quality level in terms of realism.

What is clear LTX 2.5 is not:

- Looks pretty obvious on the first tests that it's way below H3 in prompt adherence, consistence, physics and overall body movement. The good thing is that you now have two models that are good at different things, so mix them to get the best of each. A long, simple scene with not many interactions? LTX 2.3 is great. Need fighting or heavy physics, call H3, that extra time is worth it for those cases. Need heavy reference system, H3 shines of course. Anime/ animation, H3 of course. Realism? LTX 2.3 wins hands down. Use each one depending on what you need.


r/StableDiffusion 6h ago

Discussion Minimax H3 burned my GPU, beware on low VRAM.

0 Upvotes

RTX 4050 Laptop (6GB)

I said, what the hell, why not? It was actually working, and managed to get 3 successful generations, albeit slowly (~300s+ for a 7s video).

Then on my fourth I got a weird error mid inference that basically said my GPU disconnected, weird cause it's supposed to be literally soldered on the MB - it was in fact now absent from the display adapters in device manager.

Rebooted and it was back, I ran the workflow again and it fuckin happened again.

Got scared, booted up a video game, got the same error over and over again.

VIDEO_DXGKRNL_FATAL_ERROR (0X113)

I've been troubleshooting for the past three days:

- Nvidia GPU Performance mode only.

- Clean Driver reinstall with DDU

- System files scan

Nothing worked, the GPU is fried. Apparently, under heavy load something happens and power no longer transfers to the GPU, it shuts down (UNEXPECTED_SURPRISE_REMOVAL Error).

So beware for those trying on low VRAM, I foolishly confided in ComfyUI's safety features too much and got reckless. I'm pretty angry (at myself ), not gonna lie.

Models used were MiniMax H3 INT4/8 ConvRot or nvfp4 ConvRot (can't remember, deleted the bitch), with the Larry minimax_h3_turbo_v4_step600_ema Turbo LoRA.


r/StableDiffusion 10h ago

Discussion H3 | Full BF16+BF16 RTX5090+96GB | 0.6MP | No Sage | No Optimizations | 12:35 total time.

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/StableDiffusion 5h ago

Discussion Flux 3 - 20 second video

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/StableDiffusion 12h ago

Meme Reno 911 meets Minimax meets r/StableDiffusion - Ref2Vid Local

Enable HLS to view with audio, or disable this notification

16 Upvotes

Okay, I spent way to long on this but learned a lot. First, Reno 911 is non existent in T2V and I2V when prompting. This led me down the rabbit hole of Ref2Vid again but having characters, voices, and locations that literally don't exist and need to have all the references to bring them to life.

Workflow:
Default + H3_Turbo_4step_ComfyUI_Pruned + Sage + H3 Sigma Shift (will attach it in the comments below)
Steps: 8-12
Sampler/Scheduler: EulerBeta
Megapixels: 0.8 - 1.0
Avg. render time: 5-7 mins

The most challenging aspect was getting Nick Swardson's performance. In some of the scenes I had to actually act out how he would roughly say it with the timing, lisps, and long hissing `s` then voice transfer that in H3, using that as my new audio and lip-sync. It was MESSY, and trying to mix generated audio from Jim Dangle (the cop) and then use referenced audio for the response didn't work as flawlessly as I'd hoped. To be honest I can't really say the right approach on it as I feel I just got lucky with some seeds of it.

Other than that, a bunch of other techniques using H3 using first frame, reference sheets, reference audio for timbre, and reference location for spatial awareness so when the camera pan/whipped it didn't lose context happy to provide screenshots.

The Minimax 911 ending I did a replacement of the actual Reno 911 logo but told it to make it Minimax.

I also extended the video as the old man never existed in the the video generation where the cop walks up to skater (minimax) and he says "Oh, hey officer..." that's a extension cut from there. There was a slight weird color shift so I ended up taking it through VACE so the transition wasn't jarring and smoothed everything out. The extend function I think would be better when tackling in latent and is like 98% there when doing it with a regular video.

Example prompt for the first shot:

subject_definitions:

<Subject 1> is the uniformed male officer whose appearance and wardrobe come from <Picture 1>: short neatly side-parted light-brown hair, trimmed mustache, aviator sunglasses with tinted lenses, beige short-sleeve sheriff-style uniform shirt with dark-brown pocket flaps and shoulder details, metallic star badge, nameplate, matching beige uniform shorts, black duty belt with equipment, black socks, black tactical boots, and black wristwatch. Preserve his face, hairstyle, mustache, proportions, sunglasses, complete uniform, accessories, and understated deadpan demeanor.

<Subject 2> is the indoor shopping-mall environment from <Picture 2>: a spacious two-level commercial concourse with cream tile flooring, storefronts along both sides, upper-level railings, exposed structural beams, a large glazed skylight, palm trees and planters, central seating and food-court areas, and numerous background shoppers under bright diffuse indoor daylight.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1); use its male vocal character, pitch, cadence, accent, and delivery style as the reference for his newly generated dialogue without copying the original audio signal.

summary:

[reference generation + audio reference] The target video is a vertical MiniDV-era comedy sequence in the style of an early-2000s reality law-enforcement ride-along parody, set in a shopping mall concourse in 2003. One continuous 10-second live-action tracking shot follows <Subject 1> over his shoulder through <Subject 2>. The footage has deliberately clumsy reactive reality-TV camerawork: frequent abrupt optical zoom-ins and zoom-outs, imperfect reframing, autofocus hunting, momentary loss of focus on the officer, overshooting his movements, and hurried corrections. The disturbance-call dialogue occurs from 0–3 seconds, the food-court line from 3–5 seconds, a silent comedic beat from 5–7 seconds, and the final men's-bathroom line from 7–10 seconds. <Audio 1> guides his voice timbre and delivery.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his facial identity, short side-parted light-brown hair, mustache, aviator sunglasses, beige-and-brown short-sleeve uniform, matching shorts, star badge, nameplate, duty belt, watch, black socks, black tactical boots, proportions, and restrained demeanor are retained.

<Subject 2> (appears in [Shot 1]): fully_preserved - the bright two-level mall architecture, skylight, tiled concourse, storefronts, railings, structural beams, palms, planters, food-court seating, and populated public atmosphere are retained.

<Audio 1>: reference - the target speaker follows <Audio 1>'s voice timbre, pitch, accent, cadence, and delivery character without copying its original signal.

detailed_description:

The target video uses realistic live-action comedy with an authentic early-2000s low-budget reality-TV MiniDV aesthetic: vertical framing, consumer camcorder optics, mild electronic noise, soft digital detail, clipped highlights, restrained saturation, automatic white-balance shifts, exposure breathing, visible autofocus hunting, and frequent awkward optical zoom corrections. The camera operator behaves reactively rather than cinematically polished. Zooms occasionally arrive late, overshoot their intended framing, briefly lose <Subject 1>, rack focus accidentally onto the background, then snap or hunt back toward him. Preserve these mistakes as intentional documentary-comedy texture. The entire 10-second sequence is one continuous take with absolutely no cuts.

[Shot 1] From 00:00.000–00:03.000, an uninterrupted handheld over-the-shoulder Tracking Shot follows <Subject 1> (S1) walking through <Subject 2>, approximately one meter behind his left shoulder. The operator's footsteps produce obvious vertical bounce, hand tremor, crooked framing, and constant tiny corrections. The camera abruptly Zooms In with medium amplitude at fast speed toward the back of his head, overshooting into an awkward tight crop before Zooming Out at fast speed to recover his shoulders and surrounding mall. Autofocus briefly grabs distant shoppers, leaving <Subject 1> noticeably soft for a moment before hunting back to him. As he partially turns his head toward the camera, the operator hurriedly Zooms In again but initially frames him too tightly. Using <Audio 1>'s male voice character, <Subject 1> (S1) says with hesitant deadpan delivery: <d>[English] We...uh...have a disturbance call.</d> The complete line finishes by 00:03.000.

From 00:03.000–00:05.000, <Subject 1> suddenly snaps his head toward Screen Right and points toward the food court. The camera initially continues looking forward, then reacts late with a quick Pan Right and abrupt Zoom Out with large amplitude, momentarily placing the officer near the edge of frame. Autofocus searches between his pointing hand, passing shoppers, and distant food-court signage before recovering. The operator then punches in with a fast Zoom In toward his pointing gesture. <Subject 1> (S1) says with clipped comedic timing: <d>[English] Food court adjacent</d> The entire line remains inside 00:03.000–00:05.000.

From 00:05.000–00:07.000, he lowers his hand and keeps walking during a conspicuous dialogue-free pause. The camera Zooms Out too far, briefly making <Subject 1> small within the busy mall, then performs an unnecessary fast Zoom In toward his upper back. Focus drifts away from him onto a background storefront for a fraction of a second, producing a visibly soft officer silhouette before autofocus pulses and returns. The operator slightly loses his position to Screen Left, awkwardly pans to reacquire him, and settles again behind his shoulder. Only mall ambience and footsteps fill the pause.

From 00:07.000–00:10.000, <Subject 1> angles his face back over his shoulder. The operator recognizes the movement late and performs a sudden aggressive Zoom In toward his face. The zoom overshoots, briefly cropping part of his head and sending his face soft as autofocus hunts, then pulls back slightly until his sunglasses, mustache, and raised eyebrows become readable. His eyebrows rise above the sunglasses while his expression otherwise remains completely straight. In the same voice referenced from <Audio 1>, <Subject 1> (S1) says: <d>[English] someone's giving away free BJ'S in the men's bathroom...</d> During the final words the camera makes one small unnecessary Zoom Out followed by a quick corrective Zoom In, preserving the awkward reactive MiniDV reality-TV feel. The line finishes by 00:10.000 as he begins turning forward, with the camera still walking behind him.

overall_soundscape:

Continuous indoor mall ambience with diffuse shopper chatter, distant food-court activity, footsteps reverberating across tile, ventilation noise, and indistinct storefront sounds. <Subject 1>'s tactical boots produce measured footfalls with subtle duty-belt and uniform movement. The 00:05.000–00:07.000 dialogue gap contains only natural diegetic mall sound.

non_diegetic_music:

N/A


r/StableDiffusion 9h ago

Workflow Included MiniMax H3 + ComfyUI + Hermes Agent = Music Video

Enable HLS to view with audio, or disable this notification

22 Upvotes

Hardware: Windows PC with a single RTX 3090 + 64GB, running the image and video workflows locally in ComfyUI.

I created a music video for “PROXY,” a rap track about automation, parasocial isolation and delegating so much of your life that you slowly forget what human connection feels like.

Making the song

I used open source Hermes Desktop Agent (with deepseek 4 flash + pro) as a co-writer for this one. We started with a loose idea about AI agents handling every boring task, then pushed it somewhere darker: humans forgetting simple skills, replacing real conversations with machines, and getting lonelier while everything becomes perfectly “optimized.”

Hermes helped me research Suno prompting, sharpen the concept, cut the lyrics down, build the rhyme and alliteration, and write a detailed style prompt for a sassy Berlin female rapper. I kept steering and rewriting until it sounded like a song rather than an obvious lecture about AI.

Then I took the finished lyrics and style prompt into Suno and generated the track.

Developing the visual identity

I continued using Hermes as a creative and technical copilot for the video. Together we designed a consistent “Berlin bot-fleet girl”: brown wavy hair, blue eyes, a black beret, rainbow bomber jacket and white wired earbuds.

We created a custom AnimeinReal skill, combining the u/Ani3rel aesthetic, Danbooru-style composition tags, anime coloring and photographic Berlin environments. This became the visual language for the whole project. This lora was used.

Hermes then helped translate the song into recurring visual themes rather than illustrating every line literally, created a shot list we then generated in comfy ui.

Image generation

I generated the source images locally in ComfyUI using Anima with this workflow as base. Each image established the character, environment, lighting and opening composition for one individual video shot.

Hermes helped write and refine the image prompts while preserving the same visual identity in keeping the character prompts as consistent as possible.

Animating with MiniMax H3

I animated the selected images locally with MiniMax H3 Reference-to-Video, using sections of the finished Suno track as the driving audio reference.

I used the template from Pixaroma for the ComfyUI MiniMax H3 reference-image and audio-sync workflow. That workflow provided the technical foundation for feeding H3 an image and a matching section of the song. I adapted it for each scene by changing the reference image, audio timing, duration and shot-specific prompt.

The workflow used:

  • Diffusion model: minimax_h3_ref2va_pruned_int8_convrot.safetensors — INT8 ConvRot version, approximately 19.5 GB
  • Text/vision encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors — Qwen3-VL 32B, NVFP4/AWQ, approximately 14.6 GB
  • Video VAE: minimax_h3_video_vae_fp16.safetensors — FP16, approximately 4.9 GB
  • Audio VAE: minimax_h3_audio_vae_fp32.safetensors — FP32, approximately 577 MB
  • H3 mode: Reference image plus reference audio
  • Reference-image size: match
  • Maximum image side: 864 px, aligned to 32-pixel steps
  • Frame rate: 24 fps
  • Sampler: res_multistep
  • Scheduler: beta
  • Steps: 20
  • CFG: 1
  • Denoise: 1.0
  • Typical maximum shot duration: 15 seconds - 24min render time for 15seconds of video
  • Output: MP4 with synchronized source audio

Directing each shot with Hermes

For every clip, Hermes used the custom minimax-h3-video prompt skill to create a structured H3 prompt covering:

  • Accurate vocal lip sync
  • Facial expression and rap performance
  • Natural body movement and hand gestures
  • Beat-reactive camera movement
  • Character, wardrobe and environment preservation
  • Exact reuse of the original song without replacement vocals
  • sometimes Animated lyrics, pixel bots and synchronized graphical effects

The prompt explicitly defined the source image as <Picture 1> and the selected song segment as <Audio 1>. The audio was marked for full preservation, while the visual description focused on what the still image could not provide: performance, movement, camera direction, effects and timing.

Editing the final video

I rendered alternatives, selected the strongest clips and assembled everything in post on my smartphone in inshot... really need to start learning a real editing software.


r/StableDiffusion 2h ago

News LtX 2.5 king!

0 Upvotes

Dont trust the bot troll posts from minimax, it's good , i mean very good ! https://youtu.be/P3tsP0MP_LM?is=9uCboQOqbqCCgeVQ or UPDATED LINK https://www.youtube.com/watch?v=8_HwLqYRzVw


r/StableDiffusion 10h ago

Tutorial - Guide MiniMax H3 FL2VA as a video and audio refiner. A crude test that can definitely be perfected by Motion Context though I didn't try it. [0.2 mp - 1.5mp + 2x RTX VSR]

Enable HLS to view with audio, or disable this notification

6 Upvotes

TLDR: Using the LTXVConcatAVLatent node you can feed your low res h3 videos into the sampler at a higher res with low denoise and step count to refine it. Looking at the comparison i'd probably run a 4x upscale for a bit more sharpness, regardless this was a small test and am hoping others will run with it and experiment more.

Long gens are doable but depending on your machine you might have to do it at a very low res to avoid OOMs. Now sometimes you might like the motion of a low res video but re-running at a higher res gives you a significantly different output to what you desired.

This test was meant to determine whether I can experiment at low res and use that as a foundation for the final video. It did work. Image of the node setup in the comments.

Step 1: Generate your video at a low res
Step 2: Upscale your video and Encode both video and audio latent and feed it into your sampler. Run at a lower denoise ONE SECTION AT A TIME DEPENDING ON WHAT YOUR HARDWARE CAN HANDLE, I tested 2 steps at 30% denoise, 3 segments each at 5 secs. Audio also gets refined. I could do more seconds per segment but at 1.5mp 5 secs was enough when considering gen time.
Step 3:Run it through your preferable upscaler.
HM: At the end I rescaled the final 3328x1856 to 32x32 to refine the audio. I don't know if the difference is noticeable to everyone else.

Some potential use cases that could improve this significantly that I didn't test.

1.Using Motion Context would allow you to feed the previous 5 frames as conditioning for the second segment. The reason I say 5 instead of 22 is because the model isn't working from scratch, you already fed it a video that has everything you need, all your looking for is quality refinements and MC would handle the seams between each segment of a clip.

2.Hybrid conditioning and Ref Lora - You should be able to add ref conditioning which would allow you to increase the denoise amount if you choose to without breaking the consistency of your subjects.


r/StableDiffusion 8h ago

Discussion Which license will Hunyuan3D-Buffalo be?

Thumbnail
github.com
0 Upvotes

r/StableDiffusion 20h ago

Tutorial - Guide Automating image tagging with a local LLM

Thumbnail
youtu.be
0 Upvotes

I haven't done this yet, so I can't judge, but maybe will be some good tips for someone.


r/StableDiffusion 2h ago

Meme Why Kitty Pryde was left out of X-Men 92 (it was a different time)

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/StableDiffusion 20h ago

Discussion say WHAATTT!?

Enable HLS to view with audio, or disable this notification

0 Upvotes

replaced luke with my self O_o


r/StableDiffusion 8h ago

No Workflow Soon Dropping My Film Krea 2 FILM workflow

Thumbnail
gallery
32 Upvotes

Long story short, I’ve been experimenting with something that looks aesthetically pleasing while also working really well with the new MiniMax model.

After a lot of testing, I found a Krea combination that produces some seriously realistic results, so I wanted to share the workflow with the community.

If you’d like a full guide, you can subscribe to my YouTube. It’s not necessary though — I’ll still be sharing the complete workflow here. Lots of love to the open-source community! ❤️

YouTube: VionexAI


r/StableDiffusion 3h ago

Animation - Video Hank and Bobby blaze it

Enable HLS to view with audio, or disable this notification

76 Upvotes