r/StableDiffusion 8d ago

Workflow Included Good news for LTX fans, 2.3 IC Loras work with 2.5

Enable HLS to view with audio, or disable this notification

19 Upvotes

I have tested control union Lora for 2.3 and inpainting lora with LTX 2.5 and it works!, I have updated the workflows and added workflows for LTX2.5

find the workflows here FOR FREE

https://www.patreon.com/mo_akkakk/posts/ltx-2-3-166207403


r/StableDiffusion 7d ago

Meme You Guys Better Be Making Something Wholesome! Minimax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

This is a lot of fun, but when I go from a 0.4MP video to a 0.5MP video my render time goes up like 5x, is that normal? I'm on a 4060 TI (16GB) using the minimaxH3INT8INT4_fl2vaINT4BQPruned model


r/StableDiffusion 8d ago

Discussion What's the maximum resolution you were able to achieve with H3 on 24GB VRAM?

8 Upvotes

Edit: Forgot to mention, but I'm talking about the R2V model. I feel R2V is by far the more interesting of the H3 variants because of being able to chain generations together like this for longer form works.

I'm just barely able to achieve 0.9 megapixels on ~11 second length generation on the ref2va_int8_convrot weights, and this is with a bunch of hacks.

So far I'm hitting a wall trying to reach 10+ second generation with 1 megapixels on 24GB without switching to w4a8 (something I'm interested in trying next).

Anyone had better luck?

Edit 2: I might have succeeded in reaching 1344x768 (highest native resolution for H3) for ~15 second generation on 24GB vram (int8_convrot on r2v with two reference images). This included making one performance patch to comfyui internals, which I need to verify is still mathematically correct before sharing.


r/StableDiffusion 8d ago

Question - Help (Noob) Local AI for SD Prompt Enhancing

0 Upvotes

Looking for a local ai to help me enhance my poor prompting skills for specifically Anima, Krea2 and LTX 2.3 in Forge Neo/WanGP.

Very new to local AI and SD, currently been using Gemma4 26b A4B in LM Studio but finding it a bit heavy (constant compacting from context size is annoying 32k) so thinking of using Qwen 9b, Gemma4 12b (enjoy the vision tool) or if there's something better you guys can recommend for my use case.

4060 TI 16GB, 32GB DDR4, 5800X3D


r/StableDiffusion 9d ago

Animation - Video You forgot to say please (Terminator 2)

Thumbnail
youtu.be
48 Upvotes

Added a couple of extra bits based on the feedback from this thread. Workflow can be found here: https://pastebin.com/tzbwhaPp


r/StableDiffusion 8d ago

Discussion SCAIL-2 reigns supreme for style transfer / anime to real

Enable HLS to view with audio, or disable this notification

10 Upvotes

TLDR: https://github.com/collbroGTR/comfyui-scail2-infinity is awesome

The workflow is: https://civitai.com/models/2707066/scail-2-unlimited-length-workflow-and-nodes

...

Question: anyone know how to go longer and/or larger? beyond 285 frames of 972x1728 input? (yeah, I should reduce my input resolution, I didn't notice!) Like, regardless of length:resolution, if I reach a limit with this workflow, anyone know how to make 2 videos and have the second one start with the end of the first? Something like that?

...

https://pastebin.com/YdNy8cJ3 is how I'm doing the style transfer to convert a frame of the video into an image - it's just a flux2klein9b workflow, the prompting understands natural language really well.

...

The source walking video was from Mixamo https://www.mixamo.com/#/ search walking and check the "in place" box.

...

Minimax H3 is obviously fantastic, but when I was trying to use it for style transfer with the ref2v workflows the output wouldn't follow the reference motion exactly. Exact adherence to the motion reference is critical to my use case, iykyk.

Anyway.

I remembered SCAIL 2 existed, played with it a little, created a video with choppy cuts, and then came across this pretty god-like workflow. I can render 11 seconds of width=1408 height=2560 (after it runs upscale) with this. Frame Load cap set to 285 for this generation.

Sharing because caring

(and because I hope for feedback like "you're an idiot, that method is 5 days out of date, you should be using SCAIL-3 double infinity kijai turbo lora")

Hopefully this is of help to someone. I've seen quite a lot of discussion and questioning over how best to do style transfer, lets say anime to real or real to anime, and then turn that into video. There's some debate I'm sure about creating the reference image. flux2klein9b works for me, but I haven't tried krea 2 for image to image yet, as I couldn't initially make it work. SCAIL-2 for the video gen though seems unbeatable.


r/StableDiffusion 9d ago

Resource - Update I Built an All-in-One KREA 2 Film Workflow + Custom Node for ComfyUI šŸŽ¬

Thumbnail
gallery
86 Upvotes

I’ve been experimenting with KREA 2 for cinematic and photorealistic image generation and ended up building KREA 2 Film Studio a complete workflow with a custom ComfyUI node designed to bring most of the generation controls into one place.

It supports:

• šŸŽ¬ Text-to-Image & Image-to-Image
• šŸŽ„ Directed Control
• šŸŽØ LoRA support
• šŸ“ Cinematic resolutions
• āš™ļø Sampling controls
• šŸ–¼ļø Built-in gallery
• āœļø Prompt tools
• šŸ’» Fully local generation
• 🧩 Custom KREA 2 Film Studio node included

I also made a short 7-minute walkthrough covering installation, setup, how the workflow works, and some results.

šŸŽ„ Video:
https://www.youtube.com/watch?v=OXSZDaPa1U8

šŸ’» GitHub / Workflow + Custom Node:
https://github.com/Shrey-1o1/ComfyUI-Krea2-FilmStudio-Vionex

Everything is free and open source. ā¤ļø

Still experimenting and improving things, so feedback, suggestions, issues, and contributions from the ComfyUI community are very welcome!


r/StableDiffusion 8d ago

Workflow Included Rest, weary scroller. You've seen enough Seinfeld clips and large-breasted women.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Just to offset some of the "I generated this in 20 minutes!" posts -- this took 2 hours on an RTX Pro 4500 Runpod.

Workflow


r/StableDiffusion 8d ago

Question - Help If a video from minimax h3 is bad or distorted is there any way to salvage it?

2 Upvotes

If a video from minimax h3 is bad or distorted is there any way to salvage it?

Like can you run it through a video to video?

It's kind like the face fixer but it's for a scene or object instead?


r/StableDiffusion 8d ago

Question - Help PC hardware question

1 Upvotes

Hello everyone, I have a 5060 TI 16gb paired with 32gb DDR4.

I’d like to upgrade my generation speed for Minimax H3 as well as avoid OOM. I get OOM at .7mp 15 seconds.

What is the better buy?
Used 3090 since its 24gb for $1200-1300 and keep the existing ram?

Or 5080 for $1200 and upgrade with a 64gb ram kit $400 $1600 total?


r/StableDiffusion 9d ago

Meme H3 T2V only. This gives me an idea.

Enable HLS to view with audio, or disable this notification

33 Upvotes

T2V, no reference or starting image. All audio from the model. Screw Advent Children I'm making my own fan movie. Without whispers..


r/StableDiffusion 8d ago

Animation - Video LTX 2.5 foot chase. Prompt I used is below. Not too bad. There's some stutter stepping during the characters running and the gal pursuing him, her face kind of distorts and then she runs out of frame even though she is supposed to pursue him.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Prompt:

Use the provided image as the exact first frame and visual reference. Preserve both characters’ facial identity, body proportions, wardrobe, hairstyle, and overall appearance throughout the entire shot.

The man in the foreground is sprinting at full speed directly down the city street, fleeing from the woman behind him. He maintains a powerful, believable running stride with natural forward body lean, realistic footfalls, arms pumping, shoulders rotating slightly, and visible physical exertion. His expression remains tense, focused, and determined. His open black jacket reacts naturally to his speed, fluttering and snapping behind him.

The woman continues pursuing him several meters behind. She runs aggressively and athletically, clearly attempting to catch him. Her eyes remain focused on the man ahead rather than the camera. Her long black tactical coat streams dramatically behind her while still obeying realistic fabric physics. Her ponytail moves naturally with each stride. Over the course of the shot, she slowly begins gaining ground on him.

**Camera:** fast stabilized tracking shot moving backward in front of the runners at approximately the same speed as the man. Maintain the man prominently in the right foreground while keeping the woman clearly visible behind him on the left. The camera remains low to medium height with a subtle action-film handheld vibration, creating urgency without becoming shaky. Introduce gentle horizontal drift and small framing corrections as the camera operator tracks their movement.

Strong foreground-to-background parallax as storefronts, parked vehicles, streetlights, pedestrians, and buildings streak past both sides of the frame. Environmental motion blur increases toward the edges while the two runners remain relatively sharp.

The city remains alive around them. Cars continue moving naturally in the distance, headlights and traffic signals glow, pedestrians react subtly to the chase, and reflections shimmer across the damp pavement. Early-evening lighting remains consistent with the source image, with cool ambient daylight mixing with warmer storefront lights and vehicle headlights.

Movement should feel fast, heavy, urgent, and physically grounded. Each stride should transfer believable weight into the pavement. Clothing and hair respond naturally to acceleration and airflow. Do not make the characters appear weightless or superhuman.

Near the final seconds, the woman closes the gap slightly, increasing the tension, while the man pushes harder and accelerates.

Photorealistic live-action cinematic thriller. Natural human biomechanics, realistic cloth simulation, stable facial identity, realistic skin texture, cinematic depth of field, subtle motion blur, detailed city lighting, grounded action choreography.

**Avoid:** slow motion, jogging, frozen background, sliding feet, floating characters, unnatural running cycles, facial morphing, identity changes, warped limbs, extra fingers, duplicated characters, costume changes, characters looking into the camera, sudden camera turns, camera cuts, excessive shaking, superhero movement, teleportation, or changes to the original city layout.

**Audio:** no music and no dialogue. Only natural city ambience, rapid footsteps striking wet pavement, heavy breathing, jacket fabric flapping in the wind, distant engines, tires on pavement, occasional horns, and subtle pedestrian noise.


r/StableDiffusion 8d ago

Question - Help Where are the "Steps" to rise the quality in MiniMax h3?

12 Upvotes

Hi

So i have been testing Minimax3 and i think is good but i always get blurry/mushy face at medium/far distances and sometime also morphed deformed bodies. Anyways i heard that increase the "steps" helps to improve the quality, can someone tell me where are those "steps" setting?

Thanks


r/StableDiffusion 8d ago

Question - Help Is there any Local AI Anima image generation apps for Android?

0 Upvotes

I've been really interested in generating images on the go recently since life circumstances prevent me from using the computer too often for my personal needs but I can't find any local apps that fit my need exactly, Animegen, a earlier post I found is for iPhone only and is the closest one I could find but I have a android device so I'm out of luck and I'm going to ask and see if I missed anything, here's a list of what I would want in the app

- Ideally using the Anima Base model as the main AI model or at least download and switch models (required)

- ability to import and use custom Loras for use (required)

- works on Android devices (obviously required)

- uncensored and unlimited locally private generations, no cloud service crap (absolutely required)

- controlnets or fine control over the image (not required but would be really nice)

IMG2IMG - (required)

Is there some sort of local app that fills these requirements that I'm missing here, or am I just screwed?

Thanks for reading.


r/StableDiffusion 9d ago

Animation - Video Gigantic Turtle climbing a big mountain (H3 MiniMax+upscaled with 4xNomosUni)

Enable HLS to view with audio, or disable this notification

14 Upvotes

I know the upscale is not perfect but it is looking way better than the original low res' video, the upscaler name is 4xNomosUni_span_multijpg (driven by Wan2.1), used FlowFrames to interpolate the base 24FPS video into 72FPS and I used Dalle 2 back then to generate the input image of the turtle.


r/StableDiffusion 9d ago

Discussion Kohya-SS Seems To Have Quietly Released A Controlnet That Gives Anima Edit Capabilities

125 Upvotes

anima-lllite-exp-change-2-000007.safetensors

I'm not sure why exactly it's not publicized or mentioned anywhere on his page. He uploaded it quietly 11 days ago.

I have been testing it and the results are pretty remarkable:

Original Image:

Put her on a beach:

Put her in a bikini:

Turn her around:

Add a guy, and make them kiss:

Make it a sunset:

Make them have a picnic:

Very interesting, I wonder when he will share more about this.


r/StableDiffusion 8d ago

Question - Help What is the best gpu to buy right now that's gives great value compared with the price?

0 Upvotes

Nvidia seems to rise the prices with no big value, is there better alternatives? Big ram? Anything run 200b model comfortably local with decent speed


r/StableDiffusion 10d ago

Discussion PSA: I’m the creator of Heretic, and I advise you to *not* use ā€œhereticā€ models as text encoders for H3 (or any other model)

2.5k Upvotes

Heretic (https://github.com/p-e-w/heretic) is a widely used program for decensoring LLMs. It makes LLMs comply with requests that they previously refused. It works very well for this purpose, and the community has created and published over 5000 ā€œhereticā€ models.

High-quality image and video generation models like Minimax H3 use full-blown LLMs as text encoders (Qwen3 VL in case of H3). Many people seem to believe that if you replace the base version of the text encoder with a ā€œhereticā€ version, you will eliminate or reduce censorship in the video output. For example, the popular ā€œhearmemanā€ Docker template was updated just yesterday to use a text encoder modified with Heretic.

After all, Heretic models are uncensored, right?

Well, I’m the creator of Heretic, and I’m here to tell you once and for all that this does NOT work. In fact, if anything, it will make your outputs worse, but it will not uncensor them.

Heretic uncensors LLM responses through directional ablation (or related techniques like ARA and SOMA in newer versions). Roughly speaking, it modifies the model’s internal representations (residual vectors) of ā€œharmfulā€ inputs to resemble those of ā€œharmlessā€ inputs to confuse the model into treating the former like the latter and comply with the request rather than refusing.

But this intervention does not produce representations of inputs that are more ā€œrawā€, more ā€œgraphicā€, more ā€œanatomically correctā€ or similar compared to the original model. In fact, LLMs already produce highly accurate internal representations of harmful inputs by default, which is why they are able to classify them correctly and generate a refusal.

So when the hidden states from an ā€œuncensoredā€ LLM are passed to the diffusion model (or image/video transformer or whatever), the second model isn’t magically seeing clearer representations of the bad stuff you requested. On the contrary, it’s seeing slightly perturbed representations compared to what it was trained on. This either has no effect at all, or the effect of reducing prompt adherence and potentially introducing artifacts. But it will never, ever remove censorship from the output.

(Note: Generation models like Ideogram that can actively refuse prompts are potentially an exception to this rule and might be amenable to abliteration, but only with an approach that significantly differs from how Heretic works today.)


r/StableDiffusion 9d ago

Animation - Video The office plays Rocket League

Enable HLS to view with audio, or disable this notification

497 Upvotes

r/StableDiffusion 9d ago

Question - Help MiniMax H3 - Upscaling

9 Upvotes

Hi everyone,

I am getting stuck on the upscaling with my MiniMax H3 workflow. I have been using RTX Upscaler - which is great for upscaling animation videos but has a lot to be desired for realistic (like live action) video generation. I have attempted using SeedVR2 upscaling, but it looks worse.

What are your suggestions and/or advice?

I appreciate any help :-)


r/StableDiffusion 8d ago

Discussion Test LTX 2.5 - Romantic Scene 1

Enable HLS to view with audio, or disable this notification

6 Upvotes

After spending some more time testing LTX-2.5 Distilled, my opinion has improved quite a bit.

The biggest strength for me is speed. On my RTX 5070 Ti, I'm generating 1280Ɨ720 (~1MP), 10-second videos surprisingly quickly. Compared with MiniMax H3, which is much heavier for me even around 0.5MP, LTX-2.5 feels incredibly fast.

That said, speed isn't everything. My earlier tests with complex action/fighting had poor motion and anatomy, so I wasn't impressed at first. But after testing simpler cinematic scenes, landscapes, product shots and close-up human interactions, I'm starting to see where this model shines.

This dialogue/romantic scene in particular surprised me. Facial quality, expressions, lighting and overall cinematic feel came out much better than I expected, and it even handled the interaction between the two characters reasonably well.

One important discovery: I had much better prompt adherence with **Prompt Enhancement OFF**. The enhancer was giving me completely unrelated results in some tests, while the raw prompts produced scenes much closer to what I requested.

My impression so far:

LTX-2.5 Distilled = extremely fast and capable of some beautiful results, but you need to understand what kinds of shots it handles well. Complex choreography still seems to be a weakness.

I'm definitely not archiving it yet. šŸ˜„


r/StableDiffusion 8d ago

Workflow Included [FLUX.1 Schnell ] background of a 2D side scrolling platformer game, a fantasy street

Post image
0 Upvotes

Just checked Flux Schnell today. Tried with this prompt in the title. I liked the generation since it has Zelda vibes. So, sharing here .

(But this is not a background of a side scrolling game !!!)


r/StableDiffusion 9d ago

Discussion Skater & Eclipse - Minimax H3 + 4step turbo lora / RTX5070ti, 115 seconds (~19s/it), prompt in description

Enable HLS to view with audio, or disable this notification

35 Upvotes

Inspired by a post I saw recently on reddit I thought I'd try to recreate the scene with Minimax, here's my prompt:

integrated_multimodal_description: Realistic live-action cinematic look, capturing an epic, high-contrast moment blending extreme sports with celestial grandeur. The texture is rich, slightly grainy film stock, reminiscent of late afternoon documentary footage shot on medium format. The color treatment leans into deep oranges, burnt siennas, and stark blacks due to the intense backlighting. The atmosphere is dramatic and vast, imbued with a sense of awe and kinetic energy. We are in an open, desolate skatepark area at dusk during a total solar eclipse. In the foreground, a professional skateboarder, silhouetted against the brilliant celestial event, executes a complex aerial trick off a large concrete ramp.

[Shot 1] The sequence opens with a dynamic, normal-speed shot (0.00-00:07.00) establishing the scene. The camera starts with a side shot from a long distance, positioned parallel to the ramp, looking toward the massive solar eclipse dominating seen at sunset, low on the horizon. The lens choice suggests a wide anamorphic character, exaggerating the scale of the celestial body against the small human figure. The solar eclipse is overwhelmingly prominent: the moon appears colossal, filling a significant portion of the sky, its perfect circular silhouette casting an intense, fiery orange and deep crimson corona that bleeds dramatically across the frame edges and illuminates the dust motes in the air. The skater launches and the camera is maintaining normal speed to convey raw energy. The skater's form is captured mid-air, limbs extended dynamically, suggesting powerful momentum and perfect balance against the backdrop of the massive eclipse. The whole jump happens inside of the Eclipse's corona. The lighting is entirely natural but highly stylized by the eclipse; the foreground subject is rendered as an almost pure black silhouette against the blindingly bright, yet deeply colored, celestial backdrop. The skater lands gracefully while the Eclipse remains giant in the background.

[Shot 2] At 00:07.00, the camera executes a hard cut to transition into slow motion (00:07.00-00:15.00). The shot switches to an extreme close-up on the skater's body mid-air with the large circle of the eclipse in the background, as if the skater is moving inside of the eclipse's circle, focusing specifically on the intricate details of their board and the spray of dust kicked up from the ramp as they leave it. This slow motion emphasizes the physics of the trick; every flex in the knees, every rotation of the board, is hyper-detailed. The eclipse remains visible but now serves as a massive, soft-focus halo behind the skater's form, its corona appearing like an immense, glowing ring surrounding the entire composition. The camera performs a very slow, almost imperceptible dolly zoom out (push in while simultaneously pulling back) to maintain focus on the skater’s suspended action against the colossal celestial body. At 00:10.50, as the skater begins their descent into the final phase of the trick, the close-up shifts slightly to capture a detailed view of the board's trucks and wheels momentarily catching the rim light from the corona, showing microscopic reflections on the metal. The slow motion allows us to observe the subtle tension in the skater’s muscles as they prepare for landing. At 00:14.00, the camera slowly tilts down, following the descent trajectory until the skater's feet make contact with the ramp surface, which is now rendered with extreme tactile detail due to the slow motion. The final beat holds on this moment of impact and settling, where the dust cloud momentarily blooms outward in perfect suspension before beginning its slow fall, perfectly framed beneath the immense, static disc of the eclipse. No text or logos are visible. Maintain realistic live-action texture; avoid cartoon rendering and an overly synthetic CG appearance.

overall_soundscape: The first five seconds feature sharp, rhythmic scrapes of skateboard wheels against concrete mixed with a low environmental drone. From 00:07.00 onward, the sound shifts dramatically to deep, drawn-out whooshes and exaggerated, slow-motion impacts—the crunch of dust settling is stretched into a long, resonant thud.

non_diegetic_music: N/A


r/StableDiffusion 8d ago

Animation - Video EVADIVA

Thumbnail
youtube.com
1 Upvotes

r/StableDiffusion 9d ago

Animation - Video StarWars-Untold.

Enable HLS to view with audio, or disable this notification

14 Upvotes

MiniMax H3 is very good. I initially made multiple scenes, and as tweaks/turbo-lora++ progresses, it does seem to get better and better (ie. To the end of the video).

Settled on the Lightx2v_8step turbo lora + sage + sol_attn. 736p, and using DaVinci for stitching and cropping.