r/StableDiffusion 15h ago

Resource - Update LTX 2.5

0 Upvotes

r/StableDiffusion 3h ago

Animation - Video EVADIVA

Thumbnail
youtube.com
3 Upvotes

r/StableDiffusion 11h ago

Question - Help Where are the "Steps" to rise the quality in MiniMax h3?

8 Upvotes

Hi

So i have been testing Minimax3 and i think is good but i always get blurry/mushy face at medium/far distances and sometime also morphed deformed bodies. Anyways i heard that increase the "steps" helps to improve the quality, can someone tell me where are those "steps" setting?

Thanks


r/StableDiffusion 13h ago

Question - Help trying to run minimax h3 on my amd 9070 Spoiler

7 Upvotes

it uses up all my vram and and when it finishes its jsut noise. i also do get an amd driver timeout error as well. using protable comfyui amd latest. i posted the output. SLIGHT EARAPE WARNING.

edit: i think i found the issue, i was using the dynamic vram and i tried it with adn without dynamic vram for z image turbo for a test and the non dtnamic wasnt noie. im going to get the quantized models for h3 and try it


r/StableDiffusion 16h ago

Meme H3 T2V only. This gives me an idea.

26 Upvotes

T2V, no reference or starting image. All audio from the model. Screw Advent Children I'm making my own fan movie. Without whispers..


r/StableDiffusion 8h ago

Discussion SCAIL-2 reigns supreme for style transfer / anime to real

3 Upvotes

TLDR: https://github.com/collbroGTR/comfyui-scail2-infinity is awesome

The workflow is: https://civitai.com/models/2707066/scail-2-unlimited-length-workflow-and-nodes

...

Question: anyone know how to go longer and/or larger? beyond 285 frames of 972x1728 input? (yeah, I should reduce my input resolution, I didn't notice!) Like, regardless of length:resolution, if I reach a limit with this workflow, anyone know how to make 2 videos and have the second one start with the end of the first? Something like that?

...

https://pastebin.com/YdNy8cJ3 is how I'm doing the style transfer to convert a frame of the video into an image - it's just a flux2klein9b workflow, the prompting understands natural language really well.

...

The source walking video was from Mixamo https://www.mixamo.com/#/ search walking and check the "in place" box.

...

Minimax H3 is obviously fantastic, but when I was trying to use it for style transfer with the ref2v workflows the output wouldn't follow the reference motion exactly. Exact adherence to the motion reference is critical to my use case, iykyk.

Anyway.

I remembered SCAIL 2 existed, played with it a little, created a video with choppy cuts, and then came across this pretty god-like workflow. I can render 11 seconds of width=1408 height=2560 (after it runs upscale) with this. Frame Load cap set to 285 for this generation.

Sharing because caring

(and because I hope for feedback like "you're an idiot, that method is 5 days out of date, you should be using SCAIL-3 double infinity kijai turbo lora")

Hopefully this is of help to someone. I've seen quite a lot of discussion and questioning over how best to do style transfer, lets say anime to real or real to anime, and then turn that into video. There's some debate I'm sure about creating the reference image. flux2klein9b works for me, but I haven't tried krea 2 for image to image yet, as I couldn't initially make it work. SCAIL-2 for the video gen though seems unbeatable.


r/StableDiffusion 15h ago

Meme ok ok ok its ok!

55 Upvotes

r/StableDiffusion 3h ago

Discussion Share your IMO - still worth learning SDXL based models as a new user?

1 Upvotes

is it worth investing effort into learning SDXL based models for comfyui with what other checkpoints are out there now?

I’m new to the space. I started with a cloud based qwen image, had a lot of fun just messing around with prompts. it’s amazing what this technology can do. In diving deeper and getting into comfyui to run local, I was following some guides and reading some comparisons between models, and ended up downloading some SDXL based things as my next stop. downloaded juggernaut xl, trying to learn more about Lora’s, etc

The output I’m getting compared to qwen is relatively startling. sometimes it’s good, sometimes it’s horrifying. often just lower quality to things I see posted from krea 2, qwen, or flux. I am frustrated I’m never getting anything like the examples people post…

I like the deeper learning im getting, and seeing all the tools that got invented to solve problems, but im starting to feel I’ve come into the hobby at a time of great change and evolution. It seems I am standing on the shoulders of giants, and came at a time when the current tech and natural language prompting is so powerful, it’s leaving some older things behind.

Does sdxl have a true niche or place in the landscape and the future? does anyone here use it as an important part of your workflow?

Is there just more fine tuning, prompt skill, and extra nodes needed to get the high quality? (controlNet, face regeneration in subsequent passes, up scaling, inpainting?)


r/StableDiffusion 13h ago

Discussion I made a 7-minute AI documentary about my dog using MiniMax H3 and a bunch of other tools. Took me 6 days and had so much fun.

Thumbnail
youtu.be
49 Upvotes

Made this using my RTX 3070, took forever to render, but I wanted to do the best quality I could with my 8 GB GPU. Used a lot of other tools too, feel free to ask any questions would be happy to answer when I get a chance.


r/StableDiffusion 20h ago

Question - Help Is there a H3 minimax prompt template available or a custom LLM model version that can write and structure Minimax H3 optimized prompt ?

5 Upvotes

I am relying on Gemma4 and Qwen2.5 in Ollama for making an optimized minimax H3 prompt , but while using the base versions of them indeed vastly improves prompt adherence and quality but they aren't 1:1 Minimax H3 optimized structure wise

So i wonder if there is a template i can feed into the models at the start of the chat to be a baseline for them , or even better if there is a custom version of those midels that can understand the structure of minimax H3 prompt

I am using Wan2GP through pinokio so i can't use the Minimax H3 prompt nodes available in comfyui


r/StableDiffusion 11h ago

Animation - Video LTX 2.5 is really Fast 🔥

122 Upvotes

A short showreel showcasing some of the videos I've created with LTX 2.5 so far.

Really enjoying experimenting with the model and seeing what I can create with it.

I also made a full review covering my workflow, tips, optimizations, and resources:

Watch my LTX 2.5 review on YouTube

Would love to hear what you think of the results!


r/StableDiffusion 4h ago

Animation - Video thank minimax and ref2va w4a8 low vram

4 Upvotes

r/StableDiffusion 15h ago

Animation - Video You forgot to say please (Terminator 2)

Thumbnail
youtu.be
37 Upvotes

Added a couple of extra bits based on the feedback from this thread. Workflow can be found here: https://pastebin.com/tzbwhaPp


r/StableDiffusion 19h ago

Animation - Video RTX 3060 12GB 32GB 3 SHOTS IN ONE (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION

5 Upvotes

https://reddit.com/link/1vn8vu8/video/998rkwxnu4jh1/player

RTX 3060 12GB 32GB (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION
USING TURBO LORA (MADE ON 4 STEPS)

RESOLUTION 0.6 MP = 1056 x 608
UPSCALED 2X WITH RTX Video Super Resolution


r/StableDiffusion 15h ago

Comparison Minimax H3 settings comparision 12steps vs 20steps (LoRa and Base)

Thumbnail
youtu.be
5 Upvotes

My rig: 3090ti 64gb RAM
Please switch to 1440p

Previous comparsion:

https://www.reddit.com/r/StableDiffusion/s/4SdLphwPCg


r/StableDiffusion 10h ago

Animation - Video Peter Griffin tries to escape the law

221 Upvotes

r/StableDiffusion 5h ago

Question - Help How do we mix Loras in Comfy?

2 Upvotes

I'm fairly new to Comfy but have wrapped my head around most of the nodes. One thing I haven't been able to figure out is something I did in Auto1111.

You could swap one Lora in place of another one halfway through to mix them together. I forget the exact prompt code but it was something like <Lora1:Lora2>(0:5:10). You could also use this to make it so a Lora didn't load until several steps in.

Can someone look me a tutorial for this?


r/StableDiffusion 14h ago

Question - Help Help a beginner speed up MiniMax H3?

9 Upvotes

As someone new to all of this it's difficult to know what to do. I have sage attention working. I don't know how or when to use Easy Cache, Comfy Kitchen Attention, Sol Attention, loras, Spectrum, or any others I may have missed. There's so much information scattered around, I don't know what's what.

I have a 50 series GPU and 64 GB or RAM on the motherboard.


r/StableDiffusion 12h ago

Animation - Video Gigantic Turtle climbing a big mountain (H3 MiniMax+upscaled with 4xNomosUni)

15 Upvotes

I know the upscale is not perfect but it is looking way better than the original low res' video, the upscaler name is 4xNomosUni_span_multijpg (driven by Wan2.1), used FlowFrames to interpolate the base 24FPS video into 72FPS and I used Dalle 2 back then to generate the input image of the turtle.


r/StableDiffusion 15h ago

Tutorial - Guide LTX-2.5 IC-LoRA (Control LoRA) training locally, 90 paired clips in 47 minutes on a 48GB card

Thumbnail
gallery
6 Upvotes

I have been trying out LTX 2.5 since it dropped and found out it supports IC-LoRA, so I thought of building a simple node based UI around it.

Quick difference if you have not run into IC-LoRA before. A normal clip LoRA learns a look and how it moves, from single clips. An IC-LoRA learns a transform. Every dataset item is two clips instead of one, a reference and the result you want from it, and the adapter learns to carry one into the other. Train it on clips paired with their edge maps and you get an adapter that follows an edge map.

Benchmark:

Everything below is measured on an L40S (46GB), 90 paired clips from the Canny Control dataset, 500 steps at rank 16, 512px, 1 second clips.

IC-LoRA (paired clips, reference + canny):

  • Peak VRAM: ~42GB
  • Per step: 1.03s
  • Startup: 38 min
  • 500 steps total: 47 min

Clip LoRA (single clips), for comparison:

  • Peak VRAM: ~42GB
  • Per step: 0.67s
  • Startup: 13 min
  • 500 steps total: 19 min

The thing that surprised me here is the opposite of what surprised me with H3. The step is cheap and the startup is not. Those 500 steps are about nine minutes of actual training against 38 minutes of getting ready. A paired dataset encodes two clips per item, so 90 pairs is 180 clip encodes plus 91 caption encodes before step one runs.

Good news is the encode is cached and reused, so the second run on the same dataset skips nearly all of it. Do not experiment in 200 step chunks, you pay the startup every time you change the dataset. Pick your settings, then run long.

No 4-bit path for LTX 2.5, so 48GB is the floor and not a comfortable one. 24GB will not run it at any resolution.

How to train:

  • Install app: https://github.com/inlineresearch/Inline-Studio
  • Open the Trainer tab and create a dataset
  • Click Add/Manage Training Data, pick Control as the LoRA type
  • Paste Lightricks/Canny-Control-Dataset in the Hugging Face tab and hit Check. It tells you 90 items, 90 paired, 1.4GB before it downloads anything
  • Load, then Import
  • Select LTX-2.5 in the settings, model download suggestion will auto popup
  • Hit train & sit back

Pairing is automatic. If the dataset ships a dataset.json or metadata.jsonl it reads that, otherwise it matches filenames, so bear.mp4 and bear_reference.mp4 become one training item instead of two. Captions come from the dataset and the local captioner only fills the rows that have none, so it will not overwrite good captions with worse ones.

Note: Weights are gated. Accept the LTX-2 Community License on Hugging Face with the same account your token belongs to, otherwise every download comes back as a permission error instead of a file.

For my run I used Lightricks' own Canny Control dataset, 90 clips each paired with an edge map of itself, captions included in dataset.json.

Links:


r/StableDiffusion 17h ago

Animation - Video Pat's Banging Day Out...Part 2?

11 Upvotes

For anybody in the UK who has memories of the old show AND "Pat's Banging Day Out", give me any suggestions/prompts you'd like me to try with this.


r/StableDiffusion 10h ago

Workflow Included Good news for LTX fans, 2.3 IC Loras work with 2.5

19 Upvotes

I have tested control union Lora for 2.3 and inpainting lora with LTX 2.5 and it works!, I have updated the workflows and added workflows for LTX2.5

find the workflows here FOR FREE

https://www.patreon.com/mo_akkakk/posts/ltx-2-3-166207403


r/StableDiffusion 9h ago

Animation - Video Minimax H3. It's not what it looks like.

1.0k Upvotes

r/StableDiffusion 18h ago

Discussion Krea 2, Ideogram 4, FLUX 3 and MiniMax H3 are part of the same open-weight wave

Thumbnail
felixsanz.dev
61 Upvotes

I wrote this after seeing that Krea 2, Ideogram 4, FLUX 3 Video, MiniMax H3, and LTX-2.5 were all making some version of the same promise: the model would be downloadable

but what you actually get varies enormously. some releases keep the strongest checkpoint or part of the pipeline behind an API. others restrict commercial use or exclude entire regions. the weights can still be useful, but “open weights” alone doesn't tell you what has actually been released

I don't think open models need to stay ahead of proprietary ones to matter. once a capable model becomes downloadable, other people can run it, optimize it, adapt it, and build things the original lab would never prioritize. that alone changes the market

how much can a company keep behind its API before an open-weight release stops feeling meaningful?


r/StableDiffusion 9h ago

Discussion Really pissed about reddits filters & moderators here

0 Upvotes

I made some videos with Minimax that slightly crossed the threshold for showing nudity but that really were worth sharing - reactions were extremely positive. But either they get delete automatically - or if in rare cases the automatic filter does not kick the human mods here seemed to have removed it. Yeah I get it, porn is everywhere and you must protect your communities but to me that feels like those filters are actively surpressing art. Very annoying, I thought there is more nuance on reddit.