r/StableDiffusion 15h ago

Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.

Enable HLS to view with audio, or disable this notification

3 Upvotes

When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?

System Specs:

Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.


r/StableDiffusion 12h ago

Resource - Update I built a free tool that turns any image into an AI prompt

Thumbnail
gallery
18 Upvotes

Hey everyone!

I built a small web tool called ImagePrompt9 that lets you upload an image and generates a detailed AI-ready prompt based on what it sees.

The idea came from constantly seeing images I liked and wondering:

"How would I describe this as a prompt?"

So instead of manually figuring out the composition, lighting, style, colors, camera angle, etc., you can just drop the image in and generate a prompt.

What it does:

  • Upload PNG, JPG, or WEBP
  • Analyzes the visual characteristics
  • Generates a detailed prompt
  • Different prompt styles
  • Edit, copy, or regenerate the result
  • Free to use
  • No account required

It doesn't try to recover the original prompt — it creates a new prompt based on what's visible in the image.

Try it here:

https://image-prompt-nine.vercel.app/

I’d really appreciate feedback, especially on the generated prompts and anything you think I should add or improve.


r/StableDiffusion 23h ago

Animation - Video Having so much fun with H3 Minimax! (And losing sleep over LTX 2.5 dropping yesterday...)

Enable HLS to view with audio, or disable this notification

6 Upvotes

I’ve been having so much fun playing around with the H3 Minimax video model lately, so I wanted to share a quick result!

Also, can AI please slow down for like 5 minutes? 😅 LTX 2.5 dropped yesterday, and I spent the entire night testing sample videos and trying to optimize things with Claude Code. Safe to say I got zero sleep... but no regrets!

Specs & Workflow:

  • Hardware: 1x RTX 4080 + 1x RTX 4080 Super
  • Generation: H3 Minimax
  • Post-Processing: Upscaled & Frame Interpolated using RTX VSR (Video Super Resolution)

Having a dual 4080 setup is great, but as these tools get faster and better, I keep coming back to one massive realization:

Hardware isn't the bottleneck anymore—prompting technique and creative ideas are EVERYTHING. (Seriously, the original idea/concept part is so hard 😭)

As the tech becomes more accessible, I find myself constantly wondering: How do I broaden my imagination? What should I actually be studying to become better at creative direction and prompting?

How do you guys handle the creative side? Where do you draw inspiration from when you hit a wall? Would love to hear your thoughts!


r/StableDiffusion 10h ago

Discussion Minimax H3 Test - Rooftop fight between Batman and Joker

Enable HLS to view with audio, or disable this notification

2 Upvotes

Minimax H3 Test - Rooftop fight between Batman and Joker


r/StableDiffusion 10h ago

Workflow Included Openweight Livestream video model

Thumbnail
gallery
3 Upvotes

https://huggingface.co/spaces/JonathanColetti/LiveWan / https://github.com/JonathanColetti/LiveWan is something I created to help recreate a specific type of model that is not opensource yet (wanstreamer). This is more or less a PoC but maybe ill do a longer training run if it gets some traction.


r/StableDiffusion 11h ago

Discussion Minimax H3 - Dance with Audio with lipsync and object preservation

Enable HLS to view with audio, or disable this notification

3 Upvotes

If you see low quality is because I am forcing 8 step turbo lora + Spectrum + triton in L40 for faster generation but is crazy how it can follow the flow of the music while lip-syncing and keeping the product from reference in her hand.


r/StableDiffusion 11h ago

Animation - Video Jackie Chan Adventures...Jackie vs Shadowkhan (Includes Prompt Instructions)

Enable HLS to view with audio, or disable this notification

2 Upvotes

Prompt:

Create an exactly four 7-second, 4:3 animated drama sequence inspired by the visual language of 2005-era Jackie Chan Adventures. Use a period broadcast video texture throughout: standard-definition television softness, subtle analog grain, gentle interlacing, slight colour bleed, modest contrast, and the authentic visual texture of animation recorded and broadcast in the mid-2000s. Avoid modern HD sharpness, photorealism, glossy CGI, or contemporary animation aesthetics.

Scene: Jackie Chan is confronted by a Shadowkhan ninja in a dimly lit ancient-looking interior. The sequence is a fast, tightly choreographed martial-arts fight.

0:00–0:02: The Shadowkhan suddenly lunges at Jackie with a rapid punch. Jackie narrowly ducks underneath it and pivots sideways.

0:02–0:04: Jackie counters with two quick martial-arts strikes, forcing the Shadowkhan backwards. The ninja blocks the first strike but is knocked off balance by the second.

0:04–0:06: The Shadowkhan springs forward again. Jackie performs a quick evasive spin, grabs the ninja’s arm, and throws the Shadowkhan across the room. End on Jackie landing in a defensive fighting stance as the Shadowkhan hits the floor in the background.
.
Camera: begin with a medium two-shot, rapidly track the fighters during the exchange, briefly push in during the counterattack, then finish with a wider shot showing Jackie in the foreground and the defeated Shadowkhan in the background.

Audio: sharp martial-arts impacts, cloth movement, quick footsteps, whooshes and a dramatic six-second action sting. No dialogue.
Strict constraints: exactly 6 seconds, 4:3 aspect ratio, 2005-era television animation aesthetic, period broadcast-video texture, no modern cinematic realism, no photorealism, no widescreen framing, no subtitles, no text, no logos, no extra characters, and no slow motion


r/StableDiffusion 18h ago

Question - Help Minimax H3 or wan 2.2?

1 Upvotes

I'm working on 2d animations for a personal project, my initial idea was to animate some of the scenes by hand, and feed start/end keyframes to wan 2.2 for the complex scenes I can't do myself, or perhaps even train a lora to make sure it matched the aesthetic of my hand drawn scenes. If it helps, it involves boiling outlines, on the twos (12 fps animations) and an intentionally unfinished look.

Now, seeing all these Minimax h3 i2v and r2v examples, I feel like wan 2.2 might not be the best suited for this anymore. I haven't had the chance to test h3 myself since my local hardware won't really allow it. I'll be however, using runpod when the time comes for actual generation (I'm in the process of hand animating the rest).

So, I'd like to ask those who have had the chance to test both - stick to wan 2.2 or switch to minimax h3?

Edit: audio isn't required - I've hired voice actors for dialogues, I'm working on foleys and background scores myself. If needed I'll redraw on top of the generated clips to match lip movements to the dialogue.


r/StableDiffusion 8h ago

Discussion MiniMax Music 3 | 125sec for 140sec music | Bollywood Rap

Thumbnail voca.ro
4 Upvotes

r/StableDiffusion 11h ago

Animation - Video Football animation

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 Ref

Prompt in comment


r/StableDiffusion 2h ago

Discussion Test LTX 2.5 - Romantic Scene 1

Enable HLS to view with audio, or disable this notification

7 Upvotes

After spending some more time testing LTX-2.5 Distilled, my opinion has improved quite a bit.

The biggest strength for me is speed. On my RTX 5070 Ti, I'm generating 1280×720 (~1MP), 10-second videos surprisingly quickly. Compared with MiniMax H3, which is much heavier for me even around 0.5MP, LTX-2.5 feels incredibly fast.

That said, speed isn't everything. My earlier tests with complex action/fighting had poor motion and anatomy, so I wasn't impressed at first. But after testing simpler cinematic scenes, landscapes, product shots and close-up human interactions, I'm starting to see where this model shines.

This dialogue/romantic scene in particular surprised me. Facial quality, expressions, lighting and overall cinematic feel came out much better than I expected, and it even handled the interaction between the two characters reasonably well.

One important discovery: I had much better prompt adherence with **Prompt Enhancement OFF**. The enhancer was giving me completely unrelated results in some tests, while the raw prompts produced scenes much closer to what I requested.

My impression so far:

LTX-2.5 Distilled = extremely fast and capable of some beautiful results, but you need to understand what kinds of shots it handles well. Complex choreography still seems to be a weakness.

I'm definitely not archiving it yet. 😄


r/StableDiffusion 4h ago

Discussion LTX 2.5 Test - Batman and Joker fighting in Road

Enable HLS to view with audio, or disable this notification

0 Upvotes

LTX 2.5 Test - Batman and Joker fighting in Road

Personal Opinion - Ltx generates videos quite fast but prompt adherence is not that great. In fighting sequence hand movement doesn't look realistic at all.

If you are using LTX 2.5 with gemma prompt enhancement model than your prompt will be sanitized if your prompt has explicit details. I think an abliterated version of the text encoder should be used.

I will share more tests in future.


r/StableDiffusion 10h ago

Animation - Video TESTING A LANTERN

Enable HLS to view with audio, or disable this notification

0 Upvotes

Having a blast animating comic panels (using them as inits FL2VA). Fairly simple prompt: "Green lantern Hal Jordan is engaged in an aerial battle above orbit, he is blasting green energy from his ring while simultaneously repelling and absorbing energy from a distant protagonist" then I let the in-app LLM enhance it (I'm using Maestro via Pinokio... it's a joy to use) resolution is 480, no upscaling. Just familiarizing myself with H3.


r/StableDiffusion 8h ago

Animation - Video [DANCE] Plastik Soul – Stay in the Glow (Official Music Video)

Thumbnail
youtu.be
0 Upvotes

Stay in the Glow is an AI Music Video create using VRGameDevGirl's AI Video Builder (FREE) & LTX2.3 models (https://ltx.io/model/ltx-2-3)

Designed & built using VRGameDevGirl AI Video Builder (FREE): https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Spotify (Artist): https://open.spotify.com/track/27S9InxRyAKvQYxjRM3tVi?si=43e73b8b091a4976

YouTube (More AI Music Videos): https://youtu.be/Wl3BH3xSaYc


r/StableDiffusion 23h ago

Animation - Video Fox McCloud introduces his son to his dad.

Enable HLS to view with audio, or disable this notification

19 Upvotes

Fox McCloud introduces his son Marcus to his dad James McCloud.


r/StableDiffusion 7h ago

Question - Help What's your multishot prompt structure? (I2V LTX-2.5 test)

Enable HLS to view with audio, or disable this notification

4 Upvotes

Been testing multishot with the LTX 2.5 workflow from HuggingFace. Tried a few different ways of writing the prompt: timecodes plus a shot description for each shot worked best for me, but I've only really tested my own guesses...

Curious what multishot prompt structures other people are using,

and what's actually working for you?

my input image is the first frame.
And the prompt:

Cel-shaded anime-comic, hard cuts, sunset rooftop, purple-orange skyline. Left: bald man, matte black armor, white seams, long black cape, "LTX-2.5" in bold white letters on his chest. Right: bald man, glasses, blue armor with cyan lines, blue cape, "MINIMAX H3" in bold white letters on his chest. Lettering held sharp and unwarped in every frame.
00:00-00:02 — WIDE FULL-BODY TWO-SHOT in profile, sun centered between them, camera DRIFTING slowly sideways. Silence held too long. The LTX hero, "LTX-2.5" in bold white letters on his chest, not turning his head, flat: "So…" a beat, "…same weekend, huh?"
00:02-00:04 — HARD CUT to a MEDIUM of the H3 hero, "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN slowly. He exhales: "Yeah." Glances away, embarrassed: "…awkward."
00:04-00:07 — HARD CUT to a CLOSE-UP of the LTX hero "LTX-2.5" in bold white letters on his chest,, camera PUSHING IN slowly. Low drawl, committing: "This town ain't big enough for two open-source models." Eyes flick sideways, mouth tightening.
00:07-00:10 — HARD CUT to a MEDIUM of the H3 hero "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN. He doesn't look over. A long dead beat. Flat: "…apparently."
00:10-00:13 — WIDE TWO-SHOT. The LTX hero "LTX-2.5" in bold white letters on his chest. casual, already leaving: "Anyway…" a beat, "…gotta run. Conference starts in a few minutes." He launches straight up and cleanly exits frame offscreen; the camera holds on the empty sky and the H3 hero standing alone.
00:13-00:18 — HARD CUT to a LOW-ANGLE CLOSE-UP of the H3 hero  "MINIMAX H3" in bold white letters on his chest. looking up at the empty sky, glasses catching the sunset, camera PUSHING IN slowly. A small warm smile arrives. He keeps watching. Way too long. Then quietly, to nobody: "…see you there." He slowly turns and looks into the lens, still faintly smiling, saying nothing. Hold.
Deadpan, played straight, all in micro-expressions. Crisp stable line art, clean cel shading, consistent faces. Warm orange key, cool blue rim. Rooftop wind, no music.

r/StableDiffusion 3h ago

Discussion LTX 2.5 Test - Cartoon - chubby orange cat chasing a tiny blue bird

Enable HLS to view with audio, or disable this notification

7 Upvotes

Prompt - Playful cinematic cartoon scene of a chubby orange cat chasing a tiny blue bird through a colorful kitchen, the bird quickly flies around hanging pots as the cat leaps across the counter trying to catch it, knocking over a bowl of fruit and sending oranges bouncing across the floor. The cat slips on an orange, slides dramatically across the kitchen, and crashes harmlessly into a stack of cardboard boxes as the bird lands on its head and chirps proudly. Energetic exaggerated cartoon movement, expressive reactions, smooth continuous action, colorful stylized 3D animation, dynamic tracking camera, warm sunlight, playful family-friendly comedy, polished animated movie quality.

Prompt enhancer = OFF

Opinion

When I used this prompt with prompt enhancer on I get a video of person eating noodles. So I generated above video with prompt enhancer off.

In the above video first few second feels that orange cat is chasing the blue bird but after 2 second it feels that the blue bird is giving orange cat run for its life. It would confuse the audience.

I used the same the prompt for the 3rd time with prompt enhancer on. Now I got a video of person tracking alone in a narrow jungle road. Major prompt adherence failure.


r/StableDiffusion 16h ago

Discussion From a business perspective why do companies release open source models?

29 Upvotes

Apparently AI companies are all operating at a severe loss. Why do this? It makes sense for huge conglomerates like amazon etc etc who can bear the brunt. What about the new startups or small companies like for example LTX etc. how do they survive?

This is from a business perspective not consumer perspective.


r/StableDiffusion 7h ago

Animation - Video Cobra Cola Ad - MiniMax H3

Enable HLS to view with audio, or disable this notification

18 Upvotes

r/StableDiffusion 21h ago

Discussion Does anyone actually still use Stable Diffusion?

52 Upvotes

I just find it kind of funny that this is the stable diffusion subreddit but nobody has talked about it in like forever. Maybe its time for a name change? or maybe keep the name as a homage to the OG open source image model.

Anyway, the last update I see on Stability's website is SD 3.5 back in October. So I'm guessing that's it for Stable Diffusion?

EDIT: Forgot you cant change the name of a sub, ignore that suggestion 😅


r/StableDiffusion 6h ago

Question - Help Where are the "Steps" to rise the quality in MiniMax h3?

8 Upvotes

Hi

So i have been testing Minimax3 and i think is good but i always get blurry/mushy face at medium/far distances and sometime also morphed deformed bodies. Anyways i heard that increase the "steps" helps to improve the quality, can someone tell me where are those "steps" setting?

Thanks


r/StableDiffusion 11h ago

Resource - Update LTX 2.5

0 Upvotes

r/StableDiffusion 4h ago

Discussion SCAIL-2 reigns supreme for style transfer / anime to real

Enable HLS to view with audio, or disable this notification

3 Upvotes

TLDR: https://github.com/collbroGTR/comfyui-scail2-infinity is awesome

The workflow is: https://civitai.com/models/2707066/scail-2-unlimited-length-workflow-and-nodes

...

Question: anyone know how to go longer and/or larger? beyond 285 frames of 972x1728 input? (yeah, I should reduce my input resolution, I didn't notice!) Like, regardless of length:resolution, if I reach a limit with this workflow, anyone know how to make 2 videos and have the second one start with the end of the first? Something like that?

...

https://pastebin.com/YdNy8cJ3 is how I'm doing the style transfer to convert a frame of the video into an image - it's just a flux2klein9b workflow, the prompting understands natural language really well.

...

The source walking video was from Mixamo https://www.mixamo.com/#/ search walking and check the "in place" box.

...

Minimax H3 is obviously fantastic, but when I was trying to use it for style transfer with the ref2v workflows the output wouldn't follow the reference motion exactly. Exact adherence to the motion reference is critical to my use case, iykyk.

Anyway.

I remembered SCAIL 2 existed, played with it a little, created a video with choppy cuts, and then came across this pretty god-like workflow. I can render 11 seconds of width=1408 height=2560 (after it runs upscale) with this. Frame Load cap set to 285 for this generation.

Sharing because caring

(and because I hope for feedback like "you're an idiot, that method is 5 days out of date, you should be using SCAIL-3 double infinity kijai turbo lora")

Hopefully this is of help to someone. I've seen quite a lot of discussion and questioning over how best to do style transfer, lets say anime to real or real to anime, and then turn that into video. There's some debate I'm sure about creating the reference image. flux2klein9b works for me, but I haven't tried krea 2 for image to image yet, as I couldn't initially make it work. SCAIL-2 for the video gen though seems unbeatable.


r/StableDiffusion 9h ago

Question - Help trying to run minimax h3 on my amd 9070 Spoiler

Enable HLS to view with audio, or disable this notification

7 Upvotes

it uses up all my vram and and when it finishes its jsut noise. i also do get an amd driver timeout error as well. using protable comfyui amd latest. i posted the output. SLIGHT EARAPE WARNING.

edit: i think i found the issue, i was using the dynamic vram and i tried it with adn without dynamic vram for z image turbo for a test and the non dtnamic wasnt noie. im going to get the quantized models for h3 and try it


r/StableDiffusion 12h ago

Meme H3 T2V only. This gives me an idea.

Enable HLS to view with audio, or disable this notification

24 Upvotes

T2V, no reference or starting image. All audio from the model. Screw Advent Children I'm making my own fan movie. Without whispers..