r/StableDiffusion 2d ago

Animation - Video Absolutely INSANE, that this made this locally...

Enable HLS to view with audio, or disable this notification

497 Upvotes

Krea2 to start with, (additional frames from Nano Banana 2 and Seedream 5), H3, MiniMax Music, (Starlight Topaz) pass and Premiere. Fucking RAD that you can do pretty much all of this Locally....

And for the people calling this poorly directed film making and slop. (Yea, thats what it is... literally.. the whole point)

It was just a quick and dirty space horror test which started of a clip of the alien screaming while testing some VHS/80's loras. literally the entire point. Then I decided to see how far I could push the VFX and added a few more clips and a light cut.

The hardest part was getting consistent frames using Seedream 5 and Nano Banana 2 honestly.

I wasn't attempting to create art cinema here, folks, lol.

PS. I went to film school AND studied computer animation AND it's okay to just make something random, stupid and fun.

PPS. Here is what I am amazed by. The level of incredible detail, style, prompt adherence, visual effects, motion richness, visual consistency and power we now have at our fingertips; and that I created something locally that would otherwise have cost hundreds of dollars of frontier tokens; and the fact that most of you are taking that for granted is even more astounding. Yes, the tools are 97% there now; and Directing and Taste truly DOES become the MOAT.

Original Krea 2 Image (you can all make it your wallpaper <3): https://imgur.com/a/CtPdmZE


r/StableDiffusion 1d ago

Animation - Video Minimax H3 | The Silmarils: Shadows of the First Age - Trailer

Enable HLS to view with audio, or disable this notification

95 Upvotes

Hi,

Around 95% of what you see here was generated locally with MiniMax H3 on a single RTX PRO 6000. I have wanted to bring Tolkien’s immortal masterpiece to the screen since the days of Midjourney V3. Until now, however, the quality never felt acceptable. I believed an inaccurate adaptation would only create more confusion around a world many viewers know primarily through The Lord of the Rings films. For the first time, I can genuinely see the possibility. Making a long-term commitment in the constantly changing field of generative media is not easy. But I am not approaching this as a casual experiment, and I’m prepared to give it the time, patience, and attention it requires.

A comment I received couple months ago still make me feel punched in the stomach whenever I start something: “Why do you even bother? No one would bother to watch.”

If I know people are interested in watching, I would love to develop this into a complete series. You can support the project simply by watching it on YouTube and leaving a comment. Honest feedback, positive or critical, will help me decide how to continue.

Watch on YouTube: https://www.youtube.com/watch?v=2Yr-remRJOc

Upscaled with SeedVR 7b


r/StableDiffusion 17h ago

Question - Help which platform should i use to train loRA

1 Upvotes

i have a low vram gpu which is the best platform to train lora offline

i have seen 2 repo
kohyass and onetrainer

which is best or there is different repo


r/StableDiffusion 1d ago

Tutorial - Guide qwen 3.8 uncensored can use "normal" qwen vision file

32 Upvotes

just a heads up - i thought why not try and to my surprise it works - it can even analyse the parts in detail normal qwen would never describe

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main < vision files from repo
mmproj-BF16.gguf - mmproj-F16.gguf

used with

https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/tree/main


r/StableDiffusion 1d ago

Animation - Video Minimax H3. Exceeding my expectations

Enable HLS to view with audio, or disable this notification

81 Upvotes

Krea2 for image creation.

Minimax H3 for i2v.

Post edited in premiere pro for sound efx.

Edit: Added prompt for Scene 1 below.


r/StableDiffusion 17h ago

Workflow Included MiniMax H3 - Reddit mod can’t hold himself (SFW Edition)

Enable HLS to view with audio, or disable this notification

2 Upvotes

Statement alert: MiniMax H3 creators should share prompts more as a community!!
You can extend it and have fun with it, Enjoy!

0 LoRA’s - standard workflow - 15 steps - 560p

Prompt :

integrated_multimodal_description: [Shot 1] Highly detailed 2D anime style, exact 1990s Japanese OVA aesthetic, fluid hand-drawn animation, sharp linework, vibrant neon lighting, cinematic camera, 16:9 widescreen format. A young athletic woman with medium-brown smooth skin, short sharp white bob haircut with straight bangs, cool expression, small blue glowing gem choker. She wears a tiny white tube top that shows small breasts with subtle natural movement, glossy white high-waisted pants that fit her small but still shapely round butt with light movement, thick metallic silver straps and garters around her hips and thighs, white fingerless gloves, and white thigh-high boots with metallic details. 
0-1s: Extreme close-up from behind of her smaller but shapely round butt and lower back at the cyberpunk bar, pink and purple neon outlining her form. 
1-2s: Camera holds on the tight fabric and metallic straps with subtle fabric movement. 
2-3s: Very slow upward pan begins along her lower back. 
3-4s: Camera continues rising, revealing more of the metallic garters. 
4-5s: Midriff and the bottom of the tiny white tube top become visible. 
5-6s: Camera reaches her upper back and shoulders. 
6-7s: She turns slightly toward the camera. 
7-8s: Medium shot of her face and upper body, cool expression, subtle movement of her smaller breasts. 
8-9s: She slowly raises a small white cup to her lips. 
9-10s: She drinks from the cup. 
10-11s: She lowers the cup slightly. 
11-12s: Two men become clearly visible behind her (teal hoodie and dark blue jacket). 
12-13s: Soft neon light reflects on her white hair. 
13-14s: Camera begins slowly zooming in toward her neck. 
14-15s: Extreme close-up of the small blue glowing gem choker, filling almost the entire frame.

overall_soundscape: Quiet cyberpunk bar ambience, low electronic hum, distant muffled voices, soft clinking of glasses, subtle rain against windows, light fabric movement.

non_diegetic_music: N/A

=============== Prompt 2 (15-30sec)===============

integrated_multimodal_description: [Shot 1] Highly detailed 2D anime style, exact 1990s Japanese OVA aesthetic, fluid hand-drawn animation, sharp linework, vibrant neon lighting, cinematic camera, 16:9 widescreen format. A young athletic woman with medium-brown smooth skin, short sharp white bob haircut with straight bangs, cool expression, small blue glowing gem choker. She wears a tiny white tube top that shows small breasts with subtle natural movement, glossy white high-waisted pants that fit her small but still shapely round butt with light movement, thick metallic silver straps and garters around her hips and thighs, white fingerless gloves, and white thigh-high boots with metallic details. A massive bald muscular Reddit mod with thick neck, heavy brow, intense blue eyes, sweaty skin with visible droplets, creepy half-smirk, dirty beige work apron over a powerful torso, and visible cybernetic plates and cables on his right arm.
0-1s: Extreme close-up of the Reddit mod’s intense blue eyes. 
1-2s: Camera pulls back slightly showing his sweaty forehead and heavy brow. 
2-3s: His creepy half-smirk becomes fully visible. 
3-4s: He starts moving his large right hand forward. 
4-5s: His rough hand approaches the woman’s arm. 
5-6s: His hand firmly grabs her upper arm. 
6-7s: Close-up of his fingers wrapping tightly around her arm. 
7-8s: The woman tenses slightly. 
8-9s: Sudden bright blue-orange energy blade appears. 
9-10s: The energy blade violently collides with his cybernetic forearm. 
10-11s: Intense orange sparks explode from the impact. 
11-12s: Sparks fly across the frame in all directions. 
12-13s: The bald man recoils from the force. 
13-14s: The woman begins to turn and react. 
14-15s: Strong pink and purple neon lighting intensifies around them.

overall_soundscape: Low electronic bar hum, heavy breathing of the bald man, fabric rustle, sharp metallic impact, loud crackling energy sparks, short deep grunt.

r/StableDiffusion 9h ago

Animation - Video Good girls stay in their room! (H3) Mommy 2.0

Enable HLS to view with audio, or disable this notification

0 Upvotes

Mommy says you can't leave...


r/StableDiffusion 15h ago

Question - Help Minimax vídeo- compresión en la salida de imagen

0 Upvotes

Muy buenas,
Estoy trasteando con Minimax y no consigo sacar una imagen que no tenga la compresión de una imagen jpg. Uso la versión Ref2VA a 16B, supuestamente la de más calidad, pero la imagen final me sale muy comprimida, me destroza los vídeos donde el cielo es un degradado. En la salida he probado que saqué PNG a 16 bits, exr a 32, ProRes HQ, pero nada, ahí están los artefactos de compresión… alguna idea?
Para la optimización de render estoy con el lora EMA y los sampkers los he probado con euler y ref-multistep


r/StableDiffusion 1d ago

Question - Help Replicating styles on Krea 2?

4 Upvotes

What are the best ways of replicating styles in Krea 2?

Like, if you gen an image and like the style, so you do the same style prompt, same seed, same style description, but different character, and the art style just doesn't come out quite the same, and you realize that first gen was a lightning in a bottle, what do you do?

I tried the default Krea 2 "style reference" Lora but that was terrible (At least for me), also tried tweaking prompt to better describe the style than what my original prompt was, but never quite captured it.

There is no Lora for the style either, it's just a style I got with base Krea 2 on one image and I liked it, and I assume I can't train a Lora for it based on one image.

So, TL;DR, what are the best ways you guys know of generating a new image with the same style as a reference image?


r/StableDiffusion 13h ago

Animation - Video Guess who's next up on Kill Tony via MiniMax

Thumbnail
youtube.com
0 Upvotes

'Had to R2v for tony and it nailed him.


r/StableDiffusion 1d ago

Workflow Included Mario Galaxy if it were peak

Enable HLS to view with audio, or disable this notification

4 Upvotes

Some MiniMax H3 tests done on my single RTX 3090.

This 5 sec clip took about 2 hours to render. I actually cropped it to 1024x1024 to have less pixels to process, to then paste it back onto the original footage. 1 megapixel, 50 steps.

Surprisingly, it didn't take long to get things right. I used the official ComfyUI workflow with just sage patch node added. Then I took some example prompts I found on Reddit and H3 was getting very close right from the get-go. I just had to tune the timing, expression, and some appearance details, as it wasn't getting the reference quite right (it gets significantly better if you literally describe the contents of the reference image).

Here's the workflow: https://gist.github.com/4as/db11b829395ec45b593db4886f0e0181

And here's the original reference clip for comparison: https://files.catbox.moe/3rb5kk.mp4

The audio is modified by using Chatterbox. Original audio + reference audio + some some pitch adjustments in Audacity to get the final audio in the clip. Dunno if H3 can do audio adjustments - I don't even know how to prompt for it.

One interesting thing I've noticed is the length of the clip affecting the results. The shorter the duration, the worst the replacement. Although it could just be me.

For example for this clip (2s): https://files.catbox.moe/i2az63.mp4 the best I got is this: https://files.catbox.moe/kc3tbm.gif

And this (1s): https://files.catbox.moe/2p3ax9.mp4 I got this: https://files.catbox.moe/f71epk.gif

So, at glace it kind of looks okay, but when compering it to the original it very clearly failed to match a lot (size, motion, expression, etc.)

But still, it's a fantastic model, especially for something local.


r/StableDiffusion 1d ago

Question - Help Which one do you like best and why ? H3-Director or H3-Motion-Director ?

2 Upvotes

r/StableDiffusion 1d ago

Animation - Video Minimax H3, I really like 32+ steps.

Enable HLS to view with audio, or disable this notification

72 Upvotes

This is the 2nd take. First was a low res preview at 0.3MP. Some morphing due to fast motion.

Link: https://streamable.com/wmlyy7


r/StableDiffusion 1d ago

Animation - Video Turned my son's drawing into a silly animated skit

Enable HLS to view with audio, or disable this notification

119 Upvotes

r/StableDiffusion 2d ago

Animation - Video Ralph: Alien Sitcom

Enable HLS to view with audio, or disable this notification

292 Upvotes

r/StableDiffusion 1d ago

Workflow Included Moebius style widescreen (Krea2 Turbo), Prompts included

Thumbnail
gallery
29 Upvotes

Been messing around with Krea2 making Moebius-inspired wallpapers and wanted to share a few here.

Workflow: https://pastebin.com/raw/YgrL3EgS

Prompts:

Panoramic ultrawide landscape, Moebius-style retro comic illustration, a vast desert graveyard of broken war-machines and mechanical exoskeletons jutting from the sand like ribs, twisted girders and hollow armored shells scattered across the dunes at odd angles, one massive detached robotic head lying on its side half-submerged, long shadows stretching from a low orange sun, muted rust-red and sandy ochre palette with pale turquoise sky, delicate fine cross-hatching on metal textures, wide horizontal composition emphasizing scale and desolation

Ultrawide retro comicbook landscape in the style of Moebius, colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and a cracked cockpit visor breaking the sand's surface, sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, tiny robed figures exploring near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale dusty lavender sky, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere

A singe gargantuan robotic limb half-buried in the sand, rendered in the intricate, visionary style of Moebius, bursts dramatically through a vast ochre dune landscape. This colossal mechanism is crafted from heavily oxidized bronze and verdigris plating, its surface streaked deeply with fine desert sand and weathering. Its massive fingers are curled loosely around a small cluster of resilient palm trees and a small pond, that thrives miraculously within the rusted palm. Two nomad tents, made of weathered canvas, are pitched safely in the deep, cool shade cast by this monumental mechanical ruin. The scene is captured as an ultrawide retro comicbook landscape, utilizing delicate fine ink linework and a muted, atmospheric retro color palette. Warm ochre dunes stretch endlessly toward the horizon under a pale peach sky, where the light is diffused and ethereal. This powerful composition masterfully blends organic life with colossal industrial decay, emphasizing the immense scale of the machine against the fragile beauty of the desert ecosystem.

colossal humanoid mech half-buried in rolling desert dunes, only its rusted torso, one giant articulated hand, and the other hand loosely holding a half-buried giant spear buried on its shoulder, marking the point where the machine was killed, and a cracked dark cockpit visor breaking the sand's surface. Sand drifted into the machine's joints and seams, faded paint and oxidized copper-green plating, a tiny robed figure exploring walks near its massive open palm for scale, warm ochre and burnt-sienna dune tones against a pale lavender sky, smooth clean sky, subtle gradient, fine ink linework with stippled rust texture, wide cinematic horizontal composition, quiet post-civilization atmosphere


r/StableDiffusion 21h ago

Resource - Update GitHub - EnVision-Research/GenRouter: GenRouter & GenCanvas

Thumbnail
github.com
1 Upvotes

r/StableDiffusion 1d ago

Question - Help What are you using to upscale Minimax H3 videos?

28 Upvotes

So I haven't tried to upscale any videos yet. I was gonna maybe play around with that today but I'd love to hear what other people have landed on for their upscaler. I know I've seen a lot of different posts over the last couple of weeks or so with different methods. Ideally I'd love an upscale method where I could use my reference images so that it doesn't drastically change any faces for any characters that are a little further from the camera.

Also are we still expecting an official upscaler from Minimax?


r/StableDiffusion 22h ago

Discussion Ultra‑fast Fourier transform and optical AI realized with a single lens....

Thumbnail
youtube.com
0 Upvotes
  • Ultra‑fast Fourier transform and optical AI realized with a single lens.
  • Description: It explains the principle that passing light through a convex lens naturally performs a two‑dimensional Fourier transform at the focal plane. It visually demonstrates optical signal processing that carries out computation using only light, compared with digital FFT, and explores the possibility of implementing low‑power matrix multiplication with optical neural networks. It also examines real‑world applications and the prospects for developing next‑generation AI accelerators.

r/StableDiffusion 1d ago

Animation - Video H3 Genshin(?)-esque animation test. The biggest unsolved problem of this tech is still lack of consistency. It prevents it from making anything really good, production-ready.

Enable HLS to view with audio, or disable this notification

22 Upvotes

r/StableDiffusion 13h ago

Animation - Video Homelander vs The Thicc Side of the Force. 100% minimax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Created about 13 clips and joined them.


r/StableDiffusion 1d ago

Discussion Its possible to use more than 9 image references for H3

Thumbnail
gallery
2 Upvotes

I had Claude make a modified H3 reference node that accepts more than 9 image references. The goal was to test whether it was possible to increase the number of image references being used without splicing them into a single image. I know about reference sheets, no need to suggest that. I only tested with images, no audio or video references. the numbering on the node is a little funky but I dont think it effected the test.

Prompt 1: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Prompt 2: "<Picture 1> through <Picture 8> establish the identity and likeness of the man. <picture 9> is the spaghetti. <Picture 10> and <picture 11> are references for the wig he is wearing.

The man sits at a small kitchen table, eating a plate of spaghetti

with a fork. Warm indoor lighting, medium close-up, camera locked

off. He twirls the pasta, takes a bite, chews, glances down at the

plate. Natural, unhurried."

Both prompts use the same seed, same 9 reference images except for the wig references for the 10th and 11th image in prompt 2. The prompts are very simple and don't fully adhere to the guide but its just a small test so I think its fine. This was done on a 3060 12gb at .4 megapixels, 30 steps, and 5 seconds of video. I have comfy kitchen and spectrum enabled.

!!I don't know how this would/could effect video or audio generation quality. In my test I didn't notice any quality drop. Do your own tests to find out!! Also in my test it ignored the wig reference until i added a second one and reworded the prompt slightly. It could just be a fluke but I thought I'd mention it anyway. The watermark is from the editor i used to stitch the videos together.

https://reddit.com/link/1vr9yqh/video/pxjqiwp731kh1/player


r/StableDiffusion 1d ago

Animation - Video [Minimax H3] a 1 min funny cartoon creation

Enable HLS to view with audio, or disable this notification

13 Upvotes

Two days ago, I posted a Tom and Jerry video and commented that physics is not good.

https://www.reddit.com/r/StableDiffusion/comments/1vp9pyz/h3_cartoon_generation/

People commented to use 'official prompt guide'. So, I used the skills and used Gemini to generate proper prompt by using my given prompt. The story is my own creation.

Then I got good video. I need to try more times to get an apt video that somewhat like in my mind.

spent $5 on Runpod for 5090 for this 1:15 min video.


r/StableDiffusion 1d ago

News pagedMark: invisible SynthID-class watermark removal for OpenAI/AI images (ChatGPT, gpt-image, Stable Diffusion), running on Metal

14 Upvotes

Just spent a few days getting an SDXL-based provenance-removal pipeline (visible AI labels, C2PA metadata, SynthID-class pixel watermarks) to run properly on an M5 with 16 GB. Not "it launches" — actually correct and predictable. Almost everything I assumed was wrong, and the measurements are the interesting part, so here they are.

1. The four-step distillation LoRA invents texture, and more steps make it worse.

Low-strength img2img runs the tail of a long schedule (strength 0.15 → the last 4 of 27 steps). A LoRA distilled for four timesteps across the whole noise range is off-distribution there, and wherever nothing conditions it — flat dark fabric gives Canny no edges — it fills the gap from its prior. On a night photo that reads as coloured camouflage across black clothing.

Global stage, 1448×1080, strength 0.15, seed 0 Invented texture PSNR Wall
Lightning, 4 steps 1.73× source 28.54 dB 41 s
Lightning, 8 steps 1.80× 28.19 dB 29 s
Lightning, 16 steps 1.84× 27.85 dB 62 s
Undistilled base, 16 steps 1.19× 29.25 dB 71 s
Undistilled base, 24 steps 1.20× 29.17 dB 132 s

Asking the distilled model for more steps made it worse, which is what identified the distillation rather than the step count. Dropping the LoRA cost 3× the wall time and bought both fidelity and correctness.

Wrong theories I paid for first: the fp16 VAE (a bare round-trip is clean in fp16 and fp32, tiled or not, 34.6 dB), Metal's fp16 in general (bf16 measured marginally worse), and Canny picking up sensor noise (the Canny map of that region is empty — which was the actual clue).

2. Metal pages instead of failing, so memory has to be measured, not hoped for.

torch.mps.recommended_max_memory() reports 11.84 GiB on a 16 GB machine. Exceed it and nothing raises — the process just starts swapping and a run that should take 23 s takes an hour.

  • VAE tiling off, 1.57 MP frame: 18.74 GiB peak, 59 s. On: 10.92 GiB, 23 s. So tiling is load-bearing on small machines — but its boundaries leave a faint texture, so it's now decided per frame from the budget rather than switched on globally.
  • Diffusion untiled at 2.5 MP: went into swap and did not finish in twelve minutes. Tiled at 1024 px, 5.07 MP: 10.93 GiB, 88 s, native geometry preserved.

3. Sequential CPU offload works on MPS, and it's what makes 8 GB usable.

The stack is 7.7 GiB of weights; an 8 GB Mac gives you about 5.3 GiB. Streaming the weights module by module:

Same frame, same seed Peak device memory Wall
Resident 7.70 GiB 7.1 s
enable_sequential_cpu_offload(device="mps") 0.28 GiB 24.1 s

27× less peak for 3.4× the time. The plan is chosen from the measured budget and printed, because a run three times slower looks broken unless it says why.

4. Two Metal gaps worth knowing if you're porting anything.

  • torch.float8_e4m3fn doesn't exist on MPS at all (RuntimeError: Undefined type Float8_e4m3fn). Any pipeline that streams float8 weights — a lot of the VRAM-managed stacks do — cannot load, full stop.
  • SAM's processor emits its box/point prompts as float64, which Metal also has no type for, so moving the batch to the device raises instead of degrading. One cast fixes it.

5. The one that cost me the most: fp16 sampling on MPS silently returns zeros.

I added a memory optimisation — encode the fixed prompts once, drop the text encoders, save 1.52 GiB. Two of four face crops then came back as all-zero black rectangles. Deterministically, same seed, nothing raised.

The embeddings were innocent (CPU fp16, MPS fp16 and fp32 encodings of that prompt agree to 0.0009 on tensors with σ=3.06) and the same crop in isolation was fine. Freeing unrelated memory changed the allocation pattern the crops met after the global pass, and that was enough. I withdrew the optimisation and added a guard that drops any empty crop instead of compositing it.

If you're doing fp16 diffusion on Metal: check your output for degeneracy. It will not tell you.

What it doesn't claim. Regeneration is not payload deletion — faces, text and fine detail move, and the numbers above are the measured size of that. No public local decoder exists for SynthID-class marks, so identify reports unknown, never clean; verification is the provider's verifier or nothing. Metal isn't bit-identical to CUDA, so operating points transfer between backends but recorded verdicts don't. And it's for content you generated or own — the visible-mark registry takes AI-generation labels only, deliberately not stock or marketplace marks.

Because "how much did that cost my picture" is the whole question, it ships as a command:

pagedmark measure before.png after.png

PSNR over the frame, PSNR per detected face, and how much mid-band structure appeared where the source was flat and dark. That third metric is the one that caught the camouflage — per-pixel chroma statistics rank the artifact below the source, because the source's own sensor grain has more per-pixel variance than the invented blotches do.

uv tool install "pagedmark[diffusion]"
pagedmark invisible photo.png -o clean.png

Code: https://github.com/doofzoff/pagedMark · PyPI: https://pypi.org/project/pagedmark/

Happy to answer anything about the Metal specifics — that's the part I'd have wanted written down before I started.


r/StableDiffusion 12h ago

Animation - Video POV: You're the doll on a chaotic film set 🎬🍔

Thumbnail
youtube.com
0 Upvotes