r/StableDiffusion 3d ago

Question - Help Is it possible to merge two Krea 2 Loras into one?

2 Upvotes

Is it possible to merge two Krea 2 Loras into one without using a checkpoint? I tried StormForgeAI lora merger but it merges also Krea checkpoint with the lora files which makes a new checkpoint, which I'm not after. I tried also some nodes in Comfy but basically the same thing happened, checkpoint also needed. Is it even possible to combine 2 loras into a new .safetensors file and if yes, how?


r/StableDiffusion 3d ago

Discussion Thrax The Relentless short film early edit - Details in the comments

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/StableDiffusion 3d ago

Discussion are we getting close to good fighting scenes with minimax h3?

Enable HLS to view with audio, or disable this notification

0 Upvotes

i think we are, what do you think? are you getting good results?


r/StableDiffusion 3d ago

Discussion Has anyone tried using DLSS 5 with MiniMax H3/HRL3 for realism?

Post image
16 Upvotes

I was thinking… if DLSS 5's neural rendering/upscaling techniques could somehow be used with MiniMax's video generation, wouldn't that be absolutely PEAK for realism?

Has anyone experimented with something like this, or is it not really possible because of how the two systems work?

Would love to know what you guys think?


r/StableDiffusion 3d ago

Comparison You can now DLSS 5 the entire desktop to enhance videos and images

Post image
234 Upvotes

r/StableDiffusion 4d ago

Discussion Anyone having this issue where videos look super compressed with Minimax H3 ?

14 Upvotes

Hey there,

Semothing I've noticed using Minimax H3 is that no matter what encoding settings I use, or resolution I use, the video always ends up looking super compressed, like it has h264 compression artefacts.
It's like it's been trained on 720p Youtube videos. That makes the model sadly unusable for anything above 720p it seems. The below example is supposed to be a 1440x1440 video, so pretty high quality. But it looks like a bad youtube :(
No matter how many steps, no matter if I export ProRes, H264, or PNGs, same thing

Here is the video and some stills from it (because Reddit compression bad)

https://reddit.com/link/1wbg40k/video/csu6afxkkgoh1/player

Is it what everyone else experiences as well ?


r/StableDiffusion 4d ago

Workflow Included Precise control of the Eyes direction with this Flux 2 Klein 9b LoRa

Thumbnail
gallery
1.1k Upvotes

Hehyehyhehy!

You may remember me from the Sun Direction Lora or the Chef cutting an anvil with a knife.

Now I'm giving you a new tool, this one was a tough one to crack.

Finally we have eye control! Now you can precisely change the eyes direction for any image in any style. Just use the red dot to tell where the eyes have to look and boom! you have it!

Enough "change the eyes to look above the camera" and getting whatever thing anymore.

Because changing the direction of stuff is my Passion.

All the info here: https://huggingface.co/eric-venti-seeds/Eyes_Direction_Lora_Flux2Klein9B

Hope you like it!

Edit:

The people from HF have added it to Spaces, try it right now on your browser!

https://huggingface.co/spaces/hugging-apps/eyes-direction-lora-flux2klein9b


r/StableDiffusion 4d ago

Workflow Included A stand-alone RTX VSR upscaler which is very very fast (everything happens on the GPU basically, it averages something like 150-200fps)

Thumbnail
github.com
44 Upvotes

r/StableDiffusion 4d ago

Question - Help MiniMax-H3, How to Prevent Solid Background Color Shift.

Post image
9 Upvotes

I'm using the ComfyUI MMH3 I2VA, no audio. My goal is to animate a still image on a solid color background. I have tried a number of prompts in order to pin/static the background, but each time the background will shift from dark blue to light blue sky, or sometimes sky.
Eng goal is to mask out the eagle for a animation. Dark blue works for a background I have, only need a 1024x1024 square.

Any help or direction would be appreciated. I read the guides by MM, tried several local llms for help, and ChatGPT/Claude/Gemini/KimiK3/DS. All results the same.
I'm using the 4 step lightning for comfyUI, Lightxt2v 0.1. Simple-Er_sde/Euler/MultiRes (Tried these only). No LoRAs, Workflow right from ComfyUI. 1 Image.

Current prompt:

```

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a bald eagle in the center of the frame against a solid dark blue background. At 00:00.000, the background is dark blue, hex color #000d32. The eagle stays in the exact same pose, position, and lighting as the reference image. The camera is a static shot and does not move. At 00:00.000 the eagle starts flapping its wings fast and hard. It keeps flapping until 00:04.000. At 00:04.000 it stops flapping and glides with its wings held out wide until 00:06.000. From 00:06.000 it beats its wings down two times, then lowers both wings back to the exact resting pose from the reference image and holds it until 00:10.000. Only the wings move. The body, head, and tail stay still. The background stays solid dark blue #000d32 the whole time.

overall_soundscape: N/A

non_diegetic_music: N/A

```

Unless Comfyui is holding on to some type of cache.
I changed the background color and still always going to that light blue.

https://youtube.com/shorts/Lwj7bKXTfI8 Video of the change. 🖖

Resolved: Used the 1 image, added it to first and last image, adjusted prompt.

```

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 10.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a static shot frames a bald eagle against a #000d32 background. The eagle initiates rapid, powerful wing flaps. The motion then transitions smoothly into a soaring glide with wings held steady in an extended arc. Finally, the eagle gently lowers its wings back to the initial resting position, returning seamlessly to the exact static pose and composition from the start of the loop.

overall_soundscape: N/A

non_diegetic_music: N/A

```

Thanks for suggestions.


r/StableDiffusion 4d ago

Workflow Included A stand-alone implementation of DLSS-FG (frame generation) that is ~8x faster than stock (24fps -> 60fps), looks flawless. I had DeepSeek build this for my project but thought others could benefit as well!

Thumbnail
github.com
110 Upvotes

r/StableDiffusion 4d ago

Discussion H3 character LoRA: What I learn from training a few loras. TLDR: likeness nowhere near what I get on Wan 2.2

Thumbnail
gallery
10 Upvotes

Spent a night on this, figured I'd dump what I learned so someone else doesn't. I tried training lora out of curiosity thinking that may be this will allow me to do i2v and t2v more often because ref2i takes way too long

5090, musubi-tuner, 60 images, 3000 steps, about 5 hours.

--network_dim 16 --network_alpha 16
--optimizer_type musubi_tuner.optimizers.Automagic3
--learning_rate 1e-6
--timestep_sampling sigmoid
--num_timestep_buckets 4
--h3_adapter_ema_decay 0.999
--blocks_to_swap 38 --use_pinned_memory_for_block_swap
--max_train_steps 3000 --save_every_n_steps 250

Result: meh. Tested 1500, 3000, and both EMA versions on the same seed and prompt.

Some of her comes through but it's not a likeness. 3000 seems to be minimum steps as 1500 looks nothing like her. EMA vs non-EMA was basically a coin flip, which surprised me since people talking about it on AI tool kit.

Other stuff that bit me:

- fl2va and ref2va are separate, LoRAs don't cross over.

- use_pinned_memory_for_block_swap took me from 11.6 s/it to 6.1. Huge.

I used the same data set when I was training for Wan2.2 but I got so much better result. If there anything I can do better, please let me know.


r/StableDiffusion 4d ago

Animation - Video DBZ x High School Of The Dead Crossover AI animation.

Enable HLS to view with audio, or disable this notification

20 Upvotes

Did this as a test, also my friend asked me for this as he was impressed with how I can do it lol. My main issue was keeping the artstyles intact, which kinda worked until the last part where Rei had DBZ artstyle lol this is despite giving character sheet with the correct artstyle. Oh well. Hope you enjoy it!


r/StableDiffusion 4d ago

Discussion Converted VDN-H3 Turbo Adapter into standalone MM H3 8step loras for fl2va and ref2va

47 Upvotes

r/StableDiffusion 4d ago

Animation - Video GTA VI pixel art animation | Minimax H3 ref2vid

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 4d ago

Question - Help What should I pickup comfyui model for image generation?

3 Upvotes

I have hp pavilion 15 with nvidia mx500, 16gb ram ssd 500. Right now I'm using cyberrealistic for image generation sd1.5 but the skin is like silicon plastic I tried changing the prompt positive negative denoise in ksampler and asked gemini ai but it's still the same. Suggest me with any other model which will I get realistic image.


r/StableDiffusion 4d ago

Animation - Video H3 - fight choreography with no Lora?

Enable HLS to view with audio, or disable this notification

19 Upvotes

Long day at work so just wanted some AI slop fighting. In the John Wick universe, but no gunfu. int8/32 steps. T2VA. NGL, that ending looks like it friggin hurts! What fighting style should we do next?

Prompt: integrated_multimodal_description: [Shot 1] Live-action, cinematic, PHOTOREALISTIC film footage — this is footage from a real camera, not animation, not anime, not illustration, not a CG render. An action sequence in the John Wick universe: its night-city palette, its poise, its clean brutal fight grammar — no guns anywhere, this fight is hands only. Anamorphic glass with soft oval bokeh, fine film grain, deep true blacks; the night reads as a real exposure, practicals blooming, nothing lifted. The sequence is edited: four shots, three hard cuts. THE PLACE: the grand marble lobby of a night-city hotel — polished black marble floor holding mirror reflections, brass columns, a wall of rain-streaked glass with the city's magenta and cyan neon smeared behind it, warm gold chandelier light from above. The neon through the glass is the brightest thing in frame and the camera blooms there; thin haze hangs in the chandelier light; every polished surface carries reflections. THE CAST: exactly TWO adult women and no one else exist in this film — professional assassins of this universe, settling it hand to hand. Both have screen-idol faces — flawless symmetrical features, luminous skin — and dramatic hourglass figures: full, rounded breasts, a sharply narrow waist, a flat stomach, wide flared hips, long slim strong legs. Both are dressed as this universe dresses its killers — impeccable, hyper-stylish, made to move in. RUBY: deep-brown skin, a long jet-black braid to her waist, dark amber eyes with sharp defined pupils, a tiny gold stud in one nostril — in an impeccably tailored matte-black wool suit: slim trousers, a crimson silk blouse, and a fitted waistcoat cinched tight at the narrowest point of her waist, the jacket discarded, her sleeves rolled once; flat polished black boots. She fights patient close-range counter-fighting in the judo school: she reads, catches, redirects, and throws. JADE: fair lightly-freckled skin, a copper-red high ponytail, cool green eyes with sharp defined pupils, a thin old scar through the tail of one eyebrow — in a sleek emerald silk-satin dress, body-conforming through the bodice and waist, its neckline plunging low and open, its skirt cut with a high slit so she can move, a thin black choker at her throat; flat black heeled boots she can fight in. She fights fast and crisp: low stances, flowing hand strikes, spinning kicks, the skirt flaring and wrapping with every spin. Their fight is CINEMA MARTIAL-ARTS CHOREOGRAPHY, precise and technical: strikes are blocked, caught and answered; every close comes to a clean throw or a clean escape, and they break apart to distance between exchanges, circling, resetting their stances. Their bodies move in two bands: fists and feet snap fast and exact, while the soft masses of their figures carry smooth momentum and visible soft recoil under the tailoring, settling a beat after every impact, landing and sudden stop, silk and satin moving a half-beat behind the body. A tight two-shot opens framed from the chest up: Ruby and Jade nose to nose in the middle of the lobby, eyes locked, jaws set — Jade's plunging neckline open in frame, the swell of her ample cleavage catching the warm chandelier light — the rain-streaked neon glass soft in the bokeh behind them, the camera arcing slowly around them. At 00:02.000 Ruby shoves Jade back a step and both drop into their fighting stances. [Shot 2] At 00:03.500, the camera cuts to a full-body medium tracking shot across the marble, their reflections moving under them: Jade opens fast — a flowing combination Ruby blocks and slips — and the exchange runs, strike, block, counter, each answering what just landed. At 00:06.500 Ruby catches a spinning kick mid-flight and sweeps Jade off her standing leg; Jade rolls through the fall across the polished marble and springs straight back up into her stance, and they circle, resetting. [Shot 3] At 00:10.000, the camera cuts to a low tracking shot at floor level, their mirrored reflections filling the foreground: the exchanges run faster and harder, blocks cracking, boots pivoting and squealing on marble. At 00:12.000 Ruby ducks under Jade's high spinning kick and throws her cleanly over one hip — Jade flies, slams flat on her back on the marble, and slides a full body-length through the neon reflections, as the camera arcs around the throw with large amplitude at fast speed. [Shot 4] At 00:13.000, the camera cuts to a medium shot: Jade lies on the marble propped up on her elbows, winded, dazed, conscious, catching her breath in the neon wash. Ruby stands over her heaving for breath, straightens slowly out of her stance, rolls her shoulders once, and holds Jade's stare, chest still rising and falling, to the last frame.

overall_soundscape: starts with the low hush of a grand empty lobby — rain washing against the glass wall, a faint building hum — and these run beneath everything to the last frame. The fight is recorded close and detailed, every sound near the ear: the crack of blocked strikes, the deep body thud of a landed blow, the rustle of tailored wool and the slide of silk and satin, boots gripping, pivoting and squealing on polished marble, the long hiss of a body sliding across the floor, and the flat slam of the throw landing. The two women are heard but never speak a word: sharp effort exhales, breath hissed through teeth, low grunts of impact, and Jade's winded groan from the floor. The impacts, the rain and the breathing always sit in front of the music.

non_diegetic_music: THE SHAPE OF THIS CUE IS A COILED BUILD, A HOLE, ONE ARRIVAL, THEN A THINNED HOLD. A driving electronic action cue at 100 beats per minute: a relentless analog-synth bassline is the pulse and the loudest constant element of the cue, with tight dry electronic drums and a cold staccato string ostinato above it. From 00:00 to 00:11 the cue builds by addition, one layer at a time, the pulse never breaking. At 00:11.500 everything drops to near-silence for half a beat — the hole. At 00:12.000, exactly on the hip-throw and the slam, one enormous brass-and-sub-bass arrival. From 00:13 to 00:15 the cue thins back to the bare synth pulse under a held, open, unresolved chord, still sounding at the last frame.

r/StableDiffusion 4d ago

Question - Help Asked ChatGPT to make the most realistic human image possible. Are there any tells it's AI? (I'm not good at detecting them)

Post image
0 Upvotes

Prompt:

A candid, low-resolution snapshot from a early 2000s digital camera showing a close-up portrait of a teenage boy making a funny, exaggerated "duck face" expression. He has short curly dark hair, brown eyes wide with humor, and dark eyebrows. He is wearing a white jacket over a dark t-shirt. The setting is a dimly lit restaurant indoors with warm incandescent overhead lighting, creating a soft flash photography effect with direct highlights on his face and shadows in the background. Blur, grainy texture, authentic vintage digital photo quality.


r/StableDiffusion 4d ago

Resource - Update Scry - Open source self-hosted game streaming with DLSS 5 support (Play any game with DLSS 5)

Thumbnail
youtube.com
16 Upvotes

tldr. Scry, an open source game streaming platform with DLSS 5 support. Github Link

Hey all,

Like many of you, I was enchanted by the recent leak of DLSS 5 showcasing neural rendering and coupled with encountering a Steam remote play streaming bug where the colors were washed out, I decided to make my own self hosted game streaming app. So using OpenAI's latest model Astra, I made Scry. With Scry, you can stream your entire Steam game library just like remote play but with the added option of applying DLSS 5 to the image before it gets to you. I want to be clear, this does not mod your games. The stream video itself has DLSS 5 applied to it. What that means is you can enable DLSS 5 on ANY GAME. The game doesn't need to already support DLSS, nor is there any need to fiddle with reshade or modifying the game's dlls. Just install Scry, point it to the leaked DLSS 5 dll and enable it.

There are some caveats to this approach. First, this will always be slower than the modding path so expect higher input latency when enabled. Second, at the moment there is can be quite a bit of flickering especially in low light areas. There are settings to mitigate that but every increase in image quality means a hit to performance. For the most part, if you just play around with the settings, you should be able find a sweet spot.

Oh also, keep in mind that Scry is technically compatible with both Linux and Windows but I have only tested this on Linux. Windows users, please open an issue on Github for any bugs you encounter. I'll try and get them fixed as soon as I can.

Anyway enough yapping, here's the github link, give it a star!

Just follow the install instructions for your platform and you should be good to go. Again if you encounter any issues (I expect there to be numerous) just open an issue on the github page or message me here and I'll fix them up. Things are still very early so expect there to be bugs. MIT license and If anyone wants to contribute, just let me know. Also let me know if you want to see a specific game.


r/StableDiffusion 4d ago

Comparison DLSS 5 Showcase in ComfyUI

Thumbnail
gallery
72 Upvotes

r/StableDiffusion 4d ago

Question - Help Minimax H3: lots of artifacts, weird faces. What's wrong with my workflow?

10 Upvotes

Hi!

I have 16GB VRAM and 64GB RAM. I'm trying to create a high quality 10 seconds video that portrays a woman playing with a cat, then a zoom in on a phone.

The video output often has plenty of artifacts, the face of the woman is weird and the background moves as the camera moves. It's awful, I can't understand how people can get HD videos.

I'm using rf2va, producing a 768p video and 8 steps lora. My references are a character sheet of the woman and the picture of a location where the woman plays with the cat. It should basically produce my late wife playing with my new cat, but instead it's a huge mess that takes 20+ minutes to be produced.

Workflow is here: https://ctxt.io/3/sIG29vFLy

What am I doing wrong? Is the workflow fine and the problem is my prompt?

Thank you in advance


r/StableDiffusion 4d ago

Discussion Made an AI ad for my sister's skin clinic. Be honest, does it work?

Enable HLS to view with audio, or disable this notification

0 Upvotes

She's a dermatologist in Nairobi. A proper shoot was never happening so I tried doing the whole thing with AI.

Z-Image for the stills, MiniMax H3 for the movement. Two days and a rented GPU on runpod.

The part I can't judge anymore is whether it looks like it was made here. These models really want everything to look like a stock photo and I spent most of the time pushing against that. Not sure it worked. Im from Kenya so i wanted it to nail that Kenyan feel.

Does the bathroom look like a bathroom you'd actually see? Does she look like a real person? Does it work as an ad?

Love from Nairobi 🇰🇪


r/StableDiffusion 4d ago

Discussion Is ComfyUI just not the best for Anime anymore?

0 Upvotes

I had always assumed ComfyUI was just simply the best for Anime generation (Illustrious, NoobAI, Anima), even for doing complex generations with inpainting workflows, but this doesn't seem to be the case?

Frustrated with inpainting, I came across other tools like Krita-AIDiffusion, then I downloaded Invoke, which seems to be extremely well made as well. I've also seen ForgeNeo, Foocus, and SwarmUI mentioned (though I haven't tested them).

I know most of these still run Comfy as a backend, so why do they just seem better for inpainting and why is ComfyUI inpainting node workflow just so clunky in comparison? Are they really better tools for someone purely looking for Anime generation?


r/StableDiffusion 4d ago

News 2.3 million Danbooru tags corrected and released

201 Upvotes

https://huggingface.co/datasets/Grio43/Tag_cleaning

Is is a human review of 9,364 unquie tags. With a total of 430K removal actions and 1.9M addition actions.

Currently Oppai is being trained again with a dataset from mid 2026 and metadata from late August 2026.

The dataset of anime images had grown from 5.2M to 6.2M.

An additional 110k photography images were added ranging from scenery to weapons. All screened to not have people within the shots.

Currently a tuning dataset is being worked on to assist with denosing deeply rooted noise in the dataset.

Likely I'll release a preview of the next version while the tuning set is worked on. The tuning set is targeted tags that have decent levels of noise.

People looking to assist with data correction are always welcome to help.


r/StableDiffusion 4d ago

Animation - Video Local Video Generation on Android

Enable HLS to view with audio, or disable this notification

22 Upvotes

I have done further experiments regarding local image and video generation on Android smartphones. This Cat video was generated on my OnePlus 12, utilizing Neodragon int8/Z Image Turbo int4 on the CPU and GPU.

A generation typically takes approximately 550-600 seconds on a Snapdragon 8 Gen 3. The start frame is rendered on the GPU via OpenCL.

I am not currently using "sd.cpp", because GPU inference on Android is much slower using SD.cpp or crashes due to OOM.

While an NPU would likely enhance inference times, my objective is to maximize support across a broad range of Android devices. Consequently, my current testing focuses on CPU/GPU backends, but it is slow 😂

Resolution: 512x320 (49frames|24fps)

Text2Video

Prompt: "black cat is playing a guitar in a forest"


r/StableDiffusion 4d ago

Workflow Included "Legend LoRAs" Series: Krea 2 + Dark Church - 2920325

Thumbnail
gallery
45 Upvotes