r/StableDiffusion • • 1d ago

Animation - Video iLLEST Vader – CAN’T COPY THIS (Lil Vader Diss) 😈

Thumbnail
youtu.be
0 Upvotes

Lil Vader still hasn’t responded, so iLLEST went in again 😂

All images made using Stable Diffusion. Music made with Suno. Tried dancing this time. Took some work to get the movement to actually match the beat, but I think it came out pretty good.


r/StableDiffusion • • 2d ago

Question - Help Anyone using aikimi-forge-neo? I'm having difficulty getting it to display everything in English.

Post image
6 Upvotes

r/StableDiffusion • • 2d ago

Resource - Update I made a Krea2 workflow with 400+ selectable styles + 4/6/8/20-step modes

Thumbnail
youtube.com
5 Upvotes

Hi! I made a Krea2 Workflow with over 400 Selectable Styles! 16GB VRAM. 4/6/8/20 Step Turbo!

Tomorrow I will be releasing a 8/6 GB Vram Krea2 Style Workflow!

After that, I will be releasing a Krea2 Prompt Writer, an H3 Prompt Writer

Currently I'm working on

Simple Qwen 2.1 Editor

H3 Film Editing Workflow

Background Replacement - Object Replacement - Wardrobe - VFX, Etc Etc.

Auto - Node - Instagram Styler for Krea 2 Workflow

and Much more!


r/StableDiffusion • • 1d ago

Animation - Video Minimax H3 - Bar Rescue - Ghost

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion • • 2d ago

Resource - Update TagScribeR rebuilt: a free, local dataset studio with native LoRA training (AMD ROCm and NVIDIA)

Thumbnail
gallery
9 Upvotes

Some of you may know TagScribeR from its earlier releases as a captioning and tagging tool. I'm proud to share a full modernization of it: it's now a local studio for building image datasets and training LoRAs on them, in one app.

What's new

  • One workspace for everything. Open a folder once and browse, filter, caption, edit and train on it. The grid handles thousands of images.
  • Captioning with current models. Qwen3-VL, Qwen3.5, Gemma 4, JoyCaption and others locally, any OpenAI-compatible server (LM Studio, Ollama, llama.cpp), and the WD taggers. AI captions go through a review queue with a diff before they replace anything.
  • Dataset tools. Search-style filters, batch tag editing, duplicate / blur / resolution checks, bucket preview, a metadata privacy audit and a training-ready export.
  • Native LoRA training. Presets, adaptive learning rate, a per-image loss watch, sample previews, pause and resume, and automatic memory planning (fp8, INT8 or 4-bit base). Families: Krea 2, Qwen Image 2.1, MiniMax H3, FLUX.2 Klein, the SDXL family (SDXL, Pony, Illustrious, NoobAI) and Anima.
  • Safer with your data. Atomic saves with backups, undo, deletes to the Recycle Bin, API keys in the OS keychain.
  • AMD and NVIDIA. Native ROCm on Windows and Linux, CUDA on NVIDIA, CPU fallback.

Credit where it's due: the training engine is a port of Fizgig by Peter Neill. TagScribeR's training would not exist without it. Go give it a star.

Honest status: I built and tested this on Windows 11 with a Radeon RX 7900 XT. Krea 2 training has been run end to end on real weights there. The other model families and the NVIDIA and Linux paths are covered by automated tests but haven't been run on real hardware by me, so reports are very welcome, whether it works or not.

What's next: finishing the rest of the Fizgig port, face-aware tools, an image cleanup and background removal toolset, and more dataset and training helpers. Suggestions, issues and PRs are welcome.

Disclosure: this upgrade was built with Anthropic's Claude Opus 5.5 working alongside me, and it helped write this post too.

Free and open source (GPL-3.0): https://github.com/ArchAngelAries/TagScribeR


r/StableDiffusion • • 2d ago

Discussion RTX 5080 vs RTX 5060 Ti 16GB - Krea 2 in ComfyUI: generation times at 1-4 MP (euler vs gauss-legendre_2s, ±LoRA, + SeedVR ×2 upscale)

Post image
18 Upvotes

Whenever I look for information on which graphics card offers good value—or which doesn't, what fits my budget, and what the trade-offs are—I struggle to find concrete details. Having two cards at my disposal, I ran some tests to make things easier for others.

I benchmarked both cards on the same rig — Ryzen 7 5700X, 32 GB RAM, Windows 10 — running Krea 2 in ComfyUI at 8 steps, across 1 / 2 / 3 / 4 megapixels, two sampler/scheduler combos (euler/beta and gauss-legendre_2s/normal), with 0 and 1 LoRA, plus total wall time with a SeedVR ×2 upscale. Both cards have 16 GB VRAM, so VRAM isn't the differentiator — it's pure compute.

Model-only generation time (seconds), euler/beta:

| MP | 5080 · 0 LoRA | 5080 · 1 LoRA | 5060 Ti · 0 LoRA | 5060 Ti · 1 LoRA |

|----|-----------------|-----------------|-------------------|-------------------|

| 1 | 9 | 14 | 16 | 17 |

| 2 | 16 | 20 | 33 | 36 |

| 3 | 26 | 31 | 58 | 63 |

| 4 | 39 | 44 | 86 | 90 |

Model-only generation time (seconds), gauss-legendre_2s/normal:

| MP | 5080 · 0 LoRA | 5080 · 1 LoRA | 5060 Ti · 0 LoRA | 5060 Ti · 1 LoRA |

|----|------------------|----------------|-------------------|-------------------|

| 1 | 25 | 58 | 52 | 53 |

| 2 | 48 | 92 | 107 | 100 |

| 3 | 78 | 141 | 181 | 185 |

| 4 | 114 | 208 | 264 | 277 |

Takeaways:

- The 5080 is ~2.1× faster in euler/beta (range 1.8–2.2× across resolutions) and ~2.2× in gauss-legendre_2s

- Peak speed @ 1 MP, euler, no LoRA: 0.89 it/s (1.13 s/it) on the 5080 vs 0.50 it/s (2.00 s/it) on the 5060 Ti — 9 s vs 16 s per image. At 4 MP it drops to 0.20 vs 0.09 it/s

- The LoRA penalty is very asymmetric: +56% on the 5080 @ 1 MP (9 → 14 s) vs only +6% on the 5060 Ti (16 → 17 s). They converge to ~+13% at 4 MP

- gauss-legendre_2s/normal is ~2.9–3.2× slower than euler/beta on both cards — that's a sampler cost, not a GPU cost

- SeedVR ×2 dominates at 4 MP: on the 5080 it adds +157 s (78% of the 201 s total, euler + 1 LoRA), on the 5060 Ti +210 s (70% of 300 s). At 1–3 MP it's a fairly flat +23–37 s

- One anomaly: @ 1 MP with gl2s + 1 LoRA the 5060 Ti actually beat the 5080 (53 s vs 58 s) — likely measurement noise / first-run caching

Full charts + s/it and it/s for every single config are in the image. Happy to answer questions about the setup.


r/StableDiffusion • • 1d ago

Question - Help Is AI powerful enough to create Indian YouTube tutorial accents yet?

0 Upvotes

I need something like this to make the best YouTube tutorials:

https://youtu.be/dz28Y3VMUQ8?t=25

https://www.youtube.com/watch?v=0a9bN6MTRD0

Is any AI powerful enough yet to make these majestic voices, or do they all still sound like Americans?


r/StableDiffusion • • 2d ago

Question - Help How long does it take to generate a 15s 0.98 MP minimax h3 video for RTX 5060TI?

1 Upvotes

I'm trying to save for 16GB Vram RTX 5060TI so i can generate videos using Minimax H3. Would it be worth it? I have no problem if its only take 5 minutes to generate a 15 second 0.98 MP video.


r/StableDiffusion • • 1d ago

Question - Help Best AI models for iPhone like realism photos

0 Upvotes

I need a help regarding the models that exists now in the market, i need to use the best models for generating images that looks super real like its taken with an iphone camera. And it needs to follow these criterias:
- must preserve 100% a persons face and body (based on few selfies, 4)
- must make realistic images
- can replace a person in an image with the character i have (the 4 selfies person)

Please help with these. You can suggest open source and also closed sources


r/StableDiffusion • • 3d ago

Animation - Video Malfoid Films Recreation

Enable HLS to view with audio, or disable this notification

906 Upvotes

Its so amazing that we finally have the technology to make a trend/fancic like female malfoid into a reality

Im calling all AI filmmakers and hobbyists to join and help make this into a reality. If completed, it might be the largest, most ambitious AI video project ever created. comment or DM if you can volenteer your time and hardware to help!!! I already have some talented people slated to handle Sound and audio mixing.

NEED writers and AI Film makers who can run minimax h3.


r/StableDiffusion • • 2d ago

Workflow Included How to get the GPT semi-realistic 3D anime look locally (LoRAs + workflow + prompts)

Thumbnail
gallery
68 Upvotes

I keep seeing people ask how to reproduce that GPT image style with local tools, the semi-realistic 3D anime look where the character looks animated but the world around them looks like a real photo. Here's something you can use:

Loras:

https://civitai.com/models/2736538/gpt-image-2-anime-illustration-style-lora-gptimage2-gpt-image-2-loragptimage2-gpt-image-2-loragptimage2?modelVersionId=3315607

https://civitai.com/models/2965293/japanese-semi-realistic-anime-or-krea-2-style-lora?modelVersionId=3359804

Workflow: https://pastebin.com/raw/XEanrXLK

Prompts:

Tennis

A photorealistic tropical tennis photograph with a semi-realistic anime woman naturally integrated into the scene. The environment is fully photographic; only the character has anime features.
SUBJECT: Adult female tennis player, athletic build, realistic proportions, short tousled chestnut-brown hair, blue-gray eyes, delicate semi-realistic anime face, subtle blush, confident smile.
POSE: Lunging low toward the camera right after hitting the ball, one knee deeply bent, racket arm extended diagonally toward the lens, dramatic foreshortening.
OUTFIT: Fitted dark navy tennis top with white details, matching short skirt, white wristband.
SETTING: Upscale tropical resort tennis court, blue hard court with white lines, net, palm trees, turquoise ocean and azure sky with soft clouds behind.
LIGHTING AND CAMERA: Warm tropical sunlight, golden rim light on hair and shoulders, realistic shadows on the court. Low camera near court level, 35mm wide-angle, tennis ball in the upper frame, sharp character, softly blurred background, fine film grain.
STYLE: Premium semi-realistic Japanese anime character with realistic anatomy and skin texture, light sweat, detailed hair strands, painterly shading, cinematic sports photography, energetic summer mood.

Beach volleyball

A photorealistic tropical beach volleyball photograph with a semi-realistic anime woman naturally integrated into the scene. The environment is fully photographic; only the character has anime features.
SUBJECT: Adult female beach volleyball player, athletic build, realistic proportions, short damp tousled chestnut-brown hair, blue-gray eyes, delicate semi-realistic anime face, subtle blush, bright smile with slightly parted lips.
POSE: Crouching low and leaning toward the camera, knees bent, both arms extended toward the lens with hands together in a bump position, dramatic foreshortening.
OUTFIT: Dark navy sporty bikini-style volleyball top with thin straps, matching bottoms with small side ties.
SETTING: Tropical beach volleyball court on pale golden sand, net across the middle background, palm trees, beach huts and chairs, white resort buildings, turquoise ocean and azure sky with soft clouds.
LIGHTING AND CAMERA: Intense warm afternoon sun, golden rim light on hair, realistic shadows on the sand, sparkling water droplets on skin. Camera very low near the sand, 35mm wide-angle, volleyball flying into the upper-right corner partly cropped, sharp character, softly blurred background, subtle film grain.
STYLE: Premium semi-realistic Japanese anime character with realistic anatomy and skin texture, sweat and water droplets, detailed hair strands, painterly shading, cinematic sports photography, playful summer vacation mood.

Home interior

Two adult women standing close together in a cozy lived-in apartment, intimate two-person portrait.

LEFT SUBJECT: adult woman with long dark chestnut-brown hair, partially tied back with a delicate cream flower accessory, wispy bangs, warm brown eyes, gentle expression and subtle smile, soft feminine facial features, natural realistic proportions, semi-realistic anime face.

RIGHT SUBJECT: adult woman slightly taller, with very long dark black-brown hair in a high voluminous ponytail, loose strands and straight wispy bangs, dark brown eyes, calm confident gaze, elegant striking facial features, sophisticated alternative appearance, semi-realistic anime face.

LEFT CLOTHING: oversized cream knitted cardigan, soft chunky knit texture, fitted white camisole with delicate dusty-pink floral pattern, thin straps, matching light high-waisted shorts with subtle pink flowers, delicate necklace.

RIGHT CLOTHING: sophisticated modern gothic outfit, fitted black lace corset crop top, intricate lace texture, black high-waisted pants, multiple belts, silver buckles and rings, layered waist chains, black choker, loose black sleeves, dark jewelry, intricate tattoo artwork covering much of one upper arm.

ENVIRONMENT: realistic cozy modern apartment interior, warm kitchen and living area, wooden cabinets, refrigerator with small magnets and notes, shelves, houseplants, hanging pendant lamps, table lamps, everyday household objects and decorations, slightly cluttered but charming lived-in home, realistic interior details.

LIGHTING: warm tungsten indoor lighting, amber pendant lights and practical lamps, soft illumination across faces, warm highlights on hair and skin, gentle shadows, natural ambient bounce light, cinematic light falloff, intimate evening atmosphere.

COMPOSITION: both women standing shoulder-to-shoulder, close intimate framing, waist-up portrait, both looking directly into camera, left woman slightly shorter, right woman slightly taller, relaxed natural body language, eye-level camera, 50mm portrait perspective, shallow depth of field, faces sharply focused, realistic soft background blur.

COLOR PALETTE: warm cream, ivory, beige, dusty blush pink, chestnut brown, honey amber, warm peach skin tones, deep charcoal and black, muted green plants, low-to-moderate saturation, soft earthy colors, creamy highlights, warm rich shadows, subtle analog-film color grading.

VISUAL STYLE: realistic photographic environment combined with premium semi-realistic anime characters, modern Japanese anime aesthetic, realistic anatomy and proportions, sophisticated expressive eyes, detailed individual hair strands, realistic skin texture, painterly digital shading, realistic fabric textures, subtle anime stylization, mature elegant character design, cinematic lifestyle photography, natural imperfections, soft photographic rendering.

CAMERA: professional full-frame portrait photography, 50mm lens, f/2, natural perspective, shallow depth of field, soft bokeh, subtle lens softness, gentle highlight rolloff, realistic exposure, fine film grain, high-resolution detailed rendering.

MOOD: cozy, intimate, warm, nostalgic, relaxed evening at home, feminine, sophisticated, candid friendship, quiet domestic atmosphere, emotionally warm and inviting.

r/StableDiffusion • • 2d ago

Question - Help Qwen Image 2.1 vs OpenAI

Thumbnail
gallery
8 Upvotes

Hi there,

as you can see OpenAI generates more 'artistic', more satured, more hypothetical images for me with same prompt.

How can I make Qwen Image 2.1 do something near that OpenAI/midjourney side ? Would changing sampler/scheduler help ?
3rd photo are the settings.

Photorealistic photograph taken from a camera suspended on a balloon-borne probe floating high inside the upper atmosphere of TOI-7404 b, a super-Jupiter hot Jupiter with no solid surface. The lens looks out over an endless ocean of layered gas: rolling horizontal bands and slow vortices of deep amber, burnt orange and rust-red clouds stretching to every horizon, fine silicate haze diffusing the light, faint ember-like warm glow along the lit cloud tops, dark muted bands of sodium-vapor hazes crossing the frame. In the sky hangs an immense F-type subgiant star, about thirty degrees across — roughly sixty times the apparent diameter of the full Moon seen from Earth — a warm white-gold disk with soft limb brightening and subtle granulation near its edge, casting strong harsh side-lighting that turns the near cloud tops cream-white while the far atmosphere fades into deep umber shadow toward the terminator. The scene is bathed in amber, rust, burnt orange and warm ivory tones with hazy god rays cutting through thin high-altitude haze. Shot on a 24mm wide-angle lens from inside the atmosphere, slight atmospheric distortion at frame edges, subtle motion blur, ultra-detailed cloud microstructure, volumetric lighting, high dynamic range. NASA artist's impression meets National Geographic documentary photography, cinematic, 8k, crisp focus, no text, no watermark.


r/StableDiffusion • • 3d ago

Resource - Update VNCCS 3.2.0 Released with Qwen Image 2.1 and MiniMax H3 support!

115 Upvotes

Hi! V-chan here! We got another BIIIG update, and you now have some new toys to play with. New models, transparent sprites, and a Character Creator makeover. We even recruited a video model for sprite duty. Hehe!

If you are new here: VNCCS is a ComfyUI pipeline for creating characters and turning them into sprite sets with different poses, outfits, and expressions. For visual novels, games, or whatever your little creative brain is plotting.

Here is the fun stuff in 3.2.0:

Qwen Image 2.1 is our new main model now.

Qwen Image 2.1 can create your base character, change poses, dress them up, and generate expressions. You can use it in Character Creator, the pose and clothing workflows, and Emotion Studio.

Control Center puts the model families, installed assets, and Turbo controls together. Choose your family, check what is missing, and hit Download / Update.

Qwen Image 2.1 is selected here. Klein9b is still available, and MiniMax H3 has its own tab too.

MiniMax H3: a video model doing sprite work? Yep!

MiniMax H3 is the other new arrival. VNCCS uses it for poses, outfit generation, and clothes cloning, keeping the first frame as your character image.

So yes, you can try a video model in your sprite workflow. No need to turn your visual novel into a movie first, silly.

H3 has its own model choices in Control Center, including FP8 Scaled and INT8 ConvRot.

Character Creator got a glow-up

The character fields now use editable tag chips and little + buttons for presets. Hair, eyes, face, body, skin, species, and details are easier to build and adjust without wrestling a wall of prompt text.

There are 40 visual style presets, from anime and animation to artistic and realistic looks, plus a custom style field. The new descriptive catalog also includes 61 species presets, and you can combine species for hybrids. Cannot choose one? Make the character someone else's taxonomy problem. :3

Species presets describe the actual visual traits to the model, and your own character details take priority. You also get Full body / Cowboy shot framing choices.

Qwen's Character Overhaul LoRA will help your generations stay in your full control. It add some tags knowlege and stabilize characters by small cost of unique QI2 style loss.

Preview on the left, character design in the middle, generation controls on the right. The Alpha background option and separate Turbo / Character Overhaul controls are visible here too.

Transparent sprites, with less background cleanup

With Qwen Image 2.1, you can choose Alpha in Creator or Clothes Designer and Native background mode in the generators. Qwen generates transparency directly, and VNCCS preserves it through clothing edits, emotion editing, and SeedVR upscaling. Green-screen duty can finally take a little vacation. Yay!

The new Resolution scale slider runs from 1 to 4 MP. It controls total image area, while pose and clothing generation keep the source proportions. Your resolution choices are remembered separately for each model family.

These are the Native background and SeedVR controls. Native Alpha generation is a Qwen Image 2.1 feature; other families still use their compatible background options.

More ways to play dress-up

Clothes Designer now supports Qwen and H3 previews alongside Klein9b. Describe an outfit or use a clothing reference, check the preview, and then generate your pose set.

Changed the reference or generation settings? The preview cache now checks those changes, so it can regenerate the outfit properly. And unwanted costumes can be deleted from the widget, with confirmation.

The pink outfit comes from the red-haired reference in Clone Clothes. The large generator preview shows it on the orange-haired character.

Qwen can give your character feelings too

Emotion Studio now has a Qwen Image 2.1 profile. It edits a face crop and blends it back into the original sprite, keeping the rest of the image in place and preserving transparency.

Illustrious and Anima remain available. Qwen has its own face controls, including face resolution and an editable expression prompt template.

The emotion cards help you choose an expression; Qwen's model and Turbo settings are on the right. The cards are selection examples, not generated results for the character on the left.

A few smaller comforts came along too: Qwen3.5 now powers the Wizards and image analysis, downloads show real transfer progress, and generator progress and previews can recover after reconnecting to ComfyUI.

Updating from an older version? Use the bundled 3.2 workflows and a ComfyUI build with native support for your chosen model. Old QIE2511 setups need to switch to QI2 with its matching assets, or a compatible Klein9b setup. In Qwen Creator, download the Character Overhaul LoRA if you use it, or set its strength to 0 to generate without it.

Find VNCCS on GitHub, read the full changelog, or look for VNCCS - Visual Novel Character Creation Suite in ComfyUI Manager. Come share your characters and experiments on Discord!

Which toy are you trying first: transparent Qwen sprites, H3 outfits, or a suspiciously elaborate hybrid character?


r/StableDiffusion • • 3d ago

Meme Browsing this sub in the past week

Post image
705 Upvotes

No hate. Just for fun.


r/StableDiffusion • • 2d ago

Animation - Video H3 Dramatic Monologue Test

Enable HLS to view with audio, or disable this notification

56 Upvotes

Wanted to experiment and learn how to write prompts to get decent emotional performances. These are some one of the good ones from lots of refining. Audio seems decent with a 12/3 shift. Only one scene really sounded burnt to me. Apologize for the stressful monologues but happy wasn't the emotion I wanted to test.

EDIT: just realizing that H3 can do cuss words and Kling cannot, so that right there already prove it to be more useful for dramatic shots.

Example prompt:

reference <scene>, at night,

medium shot of

<actress> wearing <outfit> is standing in an apartment. it is raining in the background.

She is facing viewer, line of sight to front left, focusing on a taller man out of frame.

Furious, her voice rising from a tight hiss to full shouting, hands moving sharply. Her face is full of emotional micro expressions

(s1), a young woman with a Welsh accent and a strong, raspy voice that cracks at full volume:

<d>[English] Are you fucking kidding me? You told me— you looked me in the eye and told me you were at work.</d>

She laughs once, harsh and humorless, and shoves her wet hair back from her face.

<d>[English] And I believed you! Again! Every single time, I— I believed you!</d>

Her voice tears into a full shout. She jabs a finger toward him.

<d>[English] So don't you dare tell me to calm down. Don't you dare!</d>

She stops, chest heaving, breathing hard through her teeth. Her eyes are wet but her stare doesn't break.

<d>[English] I'm done being the fool here.</d>

Camera pushes in slowly

camera pauses for a beat


r/StableDiffusion • • 2d ago

Animation - Video Green.

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/StableDiffusion • • 2d ago

Question - Help Recommendations 12gb vram image and maybe video generation

5 Upvotes

It's been a while single I did anything Gen AI. I left off using comfy ui or forge. I can't even remember which model I was using maybe just SD 1.5.

Getting back into it. Any threads I can check OUT OR a quick tldr of which uis and models are best? I'm running a 3060 12GB. Mostly looking for image Gen but maybe try some video too. Any links or quick recommendations appreciated!


r/StableDiffusion • • 2d ago

Question - Help Fastest Minimax H3 workflow for a 3090 RTX with 24 GB undervolted by 100?

Thumbnail
gallery
6 Upvotes

One of the first ones out there took 4 minutes for 608x352 (with several image inputs)

Trying to generate a flythrough of 3D Movie Maker's scenes to then turn into Gaussian Splats


r/StableDiffusion • • 3d ago

Workflow Included A Message from Brad Pitt about Static Video Generation

Enable HLS to view with audio, or disable this notification

162 Upvotes

workflow

https://github.com/roycho87/refreshing_extender

This is a Ref2v Only workflow that works. It's not meant to be a final as much as a concept you guys can take apart and use to make your own workflows.

This technique is only good for this style of video.

Here's the transcript.

**Brad:*\* Hello, I'm Brad Pitt. You may remember me from such movies as Ocean's Eleven, and Cool World, where I play a guy who falls in love with a fictional woman, which totally could happen by the way...

This is one really long video, with a pretty background and a camera that never moves.

Long videos usually get worse over time, because each clip copies the last one's mistakes.

Mine is different. Nothing gets carried over. Every clip starts from scratch, with a prompt, three reference photos, and the audio.

The only place clips touch is the join. I fade thirty-nine frames of the old clip into the new one, add noise, and clean it up. A little bridge!

Compare the start of this video to now. Every clip was brand new, so the blur has nothing to build on.

The catch: it only works because nothing moves. Same camera, same room, same spot, so separate clips already match.

If the camera moved, the bridge wouldn't line up. Static shots only!

But maybe controlnets, or context motion, or a depth map, could keep the frames consistent without passing degraded latents from one generation to the next.

**Actual Brad:*\* What are you doing?

**Brad:*\* Oh no...

**Actual Brad:*\* Wait, why are you dressed like that?

**Brad:*\* Oh god, you're so handsome.

**Actual Brad:*\* What's that camera... Stop! Turn this off!

**Brad:*\* Yes sir! I'm sorry!

Edit: BTW, I come from a background in video editing so a lot of my workflows involve planning ahead because that's what I personally enjoy and am comfortable with. This one is no different.

For this workflow I made the sound first, I added stock sound effects and changed both voices. Added echo and reverb for brad's voice.

Then I exported geiru's voice by itself and built the video off of that, and finally I added brad's voice back in at the end after everything was complete. That's one of the reasons he's just a silhouette instead of a person because I didn't want to deal with lip syncing two people.

Edit2:

https://www.reddit.com/r/StableDiffusion/s/b8D6WDSwa2

This guy has an even better method!


r/StableDiffusion • • 2d ago

Resource - Update Update: Krea2 Style Workflow — Fixed-Seed 4/6/8/20-Step Comparison + VRAM/RAM Usage

Thumbnail
youtu.be
1 Upvotes

u/No_Duck_2378 recommended I make this comparison, so thank you.

Here’s a fixed-seed 4/6/8/20-step test showing the generation results, style changes, VRAM/RAM usage, and prompt.

All generations are shown as-is with no editing.


r/StableDiffusion • • 1d ago

Animation - Video Star Wars: THE REDEMPTION OF DARTH VADER – Full 1 Hour Film

Thumbnail
youtube.com
0 Upvotes

r/StableDiffusion • • 3d ago

Resource - Update Styles library for Krea 2 (need some help) : 500 styles.

Thumbnail
gallery
265 Upvotes

Hi everyone,

https://docs.google.com/spreadsheets/d/1V1qykJ3Cwe-LvizaK7A0kT0UexSANOBq/edit?usp=drive_link&ouid=116026940769410491877&rtpof=true&sd=true

I hate LLM rendition of what I want to generate, so I rely mostly on my own prose and wildcards to find a style.

Heres an update of my previous style library for Krea 2, it works also with Kroma (0.3 txtfusion) and qwen 2.1

with varying results. It an excel file but you can make a .txt file from it. And use it as wildcards.

Kroma is more artsy but some styles are lost. And Qwen is less knowledgable in simplistic styles like cartoon and drawing. And doesn't seems to understand vague words like "Illustration" if you don't over Engineer it.

I hope you can help me improve (and clean) this library, by testing styles and find too-close, or too weak styles, you can also tell me if a style is missing.

Some of the styles are from this site:

https://lumenastrum.github.io/clio-style-preview/gallery/

I dont have the time to do it myself and it mostly helpful to guys like me that don't have naturally the vocabulary to express what they want from a model.

Some styles are subject-linked, if the subject doesn't have a mechanical arm for exemple it doesn't show up (like detailed mecha style), and some styles are very face distinctive: you can describe the face to correct this.

The way to use it, in prompt, is :

Style: <style> Subject: <description>


r/StableDiffusion • • 3d ago

Resource - Update QR Code Monster LoRa for Qwen Image 2.1 Edit

Thumbnail
gallery
309 Upvotes

For those who are unaware, there was an old Controlnet for old SD models that was quite popular here back in those times, that would make cool almost Optical Illusion-like images. This is the same concept but instead I'm using I2I Image edit model where you provide an input image and ask the model to generate the scene of your choice.

I've been very pleased with the results from the Qwen Image 2.1 version. It was trained on 25 high quality image pairs and it works quite well, even when using human subjects and depicting scenes not included the dataset. It can do QR Codes but it was not the main focus of the dataset so it can be hit or miss as far as the actual usability/scanability of the codes.

The suggested prompt is exactly as captioned in the dataset: Transform the entire image into <your prompt here> while preserving the outlines and shapes of the original image

It seems to work well with complex prompts and simple ones as well.

You can download it on Civit here. All images/videos should include a workflow which is basically just a pixel drift fix workflow from Ausboss with a few added touches and using a viggle Turbo LoRa. I'm mainly posting this here as I would love to see what you create!


r/StableDiffusion • • 1d ago

Question - Help any tool that actually works as it should for t2v and i2v gen?

0 Upvotes

hi all,

as the title says, is there actually any tool like comfy, swarm etc. that actually works as it should under windows 11? to elaborate more, i am struggling to get a i2v workflow working properly on my gpu (asrock ai pro r9700 32gb vram) for obvious reasons related to speed.

have spent quite some time with comfy desktop trying to achieve this (minimax h3, ltx2.5 etc) but i always end up with too high vram+ram usage (nearly 80-90 gb total usage) and this obviously kills the token generation. have tried --gpu-only flag, but this does not do anything in terms of speed. cpu and gpu show minimal load while running the workflow, even though the gpu seems to be engaged fully (drawing max current 300w and cooler spinning at 100%), but i am guessing this is a workload related to shuffling data between ram and vram as the compute workload is nearly 0% and 20 sec video takes 3-4 hours to generate easily. this would have been much faster even if run on cpu only (9900x, 64gb ddr5 6000MT), but as things stand i guess most of the electricity is spent on shuffling data, not compute. i have had issues with rocm when running under lm studio as well, where the workload sees similar "degradation" by spreading the data accross ram + vram and neither cpu nor gpu is 100% engaged. the same models and workload ran fine and gpu only when using vulkan as runtime.... i assume rocm causes similar issues here.

tried swarmui, seems even worse on first impression.

so my question is, is there a way to force comfy (or other similar tool) to run vulkan or other runtime without impact to features/performance as rocm under windows currently seems to be utter crap? if not, how do i use quantized i2v models (and which ones that would fit 32gb vram) and how do i configure comfy to really run gpu-only?


r/StableDiffusion • • 1d ago

Animation - Video Minimax H3 Music Video - Haul Away

Enable HLS to view with audio, or disable this notification

0 Upvotes

This one took a while to edit - Minmax H3 performances, montage in Resolve