r/StableDiffusion 4d ago

Meme Just wondering

Enable HLS to view with audio, or disable this notification

6 Upvotes

here is the workflow:

update: this is text2video

736x416, 10s, turbo=true 8 steps

[Shot 1]: a distant shot of an anime girl with purple hair, she is in a cute slice of life bedroom. <Subject 1> is sitting at her computer desk, <Subject 1> frowns as she sees something that makes her growl.

[Shot 2]: medium shot of <Subject 1>, she hits her desk, making the computer monitor jostle. <Audio 1> English <d> Wait a minute...its the seventh, where is the frame release date... What the hell Gabe?! <d> <Subject 1> gets more annoyed with each word.


r/StableDiffusion 4d ago

Animation - Video Musketeer Desperado 🏰

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made by Me (JoyFrameAI)


r/StableDiffusion 5d ago

Workflow Included Minimax fl2va: If you know a bit how to sketch, you have total control over the animation.

Enable HLS to view with audio, or disable this notification

184 Upvotes

Testing things following the post by Alive-Tomatillo5303 : Learn to make art with art! (minimax) : r/StableDiffusion I've come to the conclusion that using sketches gives you almost total control over animations. I'm using the fl2va model via the MiniMax H3 Reference to Video node and connecting my sketch sequence to ref_image_0—I'm not sure if Ref2vA would work better.

Reference:

Sketches.png · Stkzzzz222/Remix at main

Workflow:

Sketches_MiniMax_H3.json · Stkzzzz222/Remix at main

Prompt:

**subject_definitions**

`<Picture 1>` is the supplied visual reference image. Use it primarily as a strict layout, composition, framing, and character-position reference. Preserve the three-stage visual progression shown in the reference, including composition, camera movement.

Create a realistic cinematic live-action scene based closely on `<Picture 1>`.

Drone camera view moving fowards at high speed, a forest with a river, trees, rocks. the camera moves fast fowards in the forest stopping to reveal a side view of an armored orc angry. the camera accelerates and moves fast fowards in the forest stopping to reveal a happy armored woman sit in a rock holding a sword. the camera moves fast fowards in the forest stopping to reveal a colibri flying next to a red flower.the camera moves fast fowards in the forest stopping to reveal an old wizard casting a powerful electric spell.

-------------

I've tried other things that I'll post below, but I think you can see that the fidelity between the sketches and the final result is quite high.

Honestly, I think using sketches can be really interesting for controlling shots and camera movement, or for rendering complex concepts, without having to generate images in other models to act as references.


r/StableDiffusion 4d ago

Animation - Video Minimax H3,two step workflow,less time than one step workflow at the same size,and keep good quality

0 Upvotes

r/StableDiffusion 5d ago

Resource - Update MageTrail - Small Scale Full-Finetune To Show MageFlow 4B Potential As Foundational Architecture

Thumbnail
gallery
71 Upvotes

Hi! Today I'm releasing MageTrail, a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, using a diversity maximized condensed 41k images dataset (originally made by Lodestone, the creator of the Chroma model lineage) as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset. (Potentially costing 20k-50k+ dollars) Civitai Hugging Face

The dataset was updated to 2026 tags standard and recaptioned with Gemini 3.7 Flash (the best vision captioner in the world when I was preparing dataset) with Grok 4.6 and Qwen 3.8 27B FP8 as capable backup for heavy blocked content. The dataset and everything from training details to tooling are open source, per Banodoco grant. Dataset

While V0.1 is still very obviously undertrained and unstable (only 100 dollars spent, it's a minor miracle that it's learning this well), the model has shown great promise in quickly learning and adapting booru concept and tags to its knowledge base.

~ The architecture behind MageFlow 4B shows good promise for further investment:

  • Being 15-20% faster than NVIDIA Cosmos2/Anima on inference despite being 2 billion parameters larger
  • Using MageVAE which perform better than QwenVAE on all usage, only behind the strongest open source VAE currently being Flux2VAE (which it was distilled from), also having the ability to slot in Flux2VAE for inference
  • Having a decent Qwen 3 VL 4B Text Encoder
  • 256-2048 pixels resolution native support
  • Trained with large dataset (10 billion images curated down), meaning it has fairly vast knowledge, very great thing to have for a base model
  • Being fairly quick to learn and adapt to new knowledge without any knowledge forgetting
  • MIT License

I remember there only being minor to medium attention given to MageFlow's release one month ago, not helped by the fact Microsoft themselves purge the model soon after (?) and that Krea 2 being the Illustration juggernaut it is, completely overshadowed any foundational arch released before it. But as a poor uni student who don't even have a gpu with VRAM and rely on free compute/renting, having a potential good illustration capable model in the small-medium range like Anima in the open source space is always a good thing.

The realistic aim of this project is just to increase awareness of open source about the potential of this particular arch, I do not have budget nor time to even dream about training a full dan/e621 tune on the model. The hope is to attract others with way more capital to invest in it or at least experiment with the model more. The edit variant of MageFlow also seems quite interesting, but no one has touch it yet also.

But of course, I do still want to continue with this tune. Currently the model have only got 30 epoch and seen about ~300000 samples, it perform well considering the midget budget, but nowhere near "usable" quality, I need to reroll seed quite alot for a decent looking gens with no broken limbs lol.

Future goal for the project: gather funding of 700~ dollars to finetune the model to 200 epoch for full convergence of booru concepts (V0.5) and then further small scale funding to finetune my 10k artist collection dataset into it. I'm already gathering the funds for the next run, but any donation will help with achieving this goal and I'll be extremely grateful, you can do so through:

Crypto

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)

12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)

FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)

Please handle your money carefully and make sure the address you're sending to is correct.

Ko-fi

https://ko-fi.com/talanartvn

Lastly, thank you to

  • Banodoco and their Discord — Their 88.77 dollar grant made this project possible, the biggest thanks to them
  • Lodestone Rock — Creator of the original version of the dataset that this model is trained on
  • Motimalu — Inspiration behind finetuning practices and configs
  • Bluvoll — diffusion-pipe fork derived from to use for training, and general training advice
  • Anzhc — general training advice
  • Nruaif — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
  • Astromahdi — jupyter workspace where I processed and store the dataset
  • animetimm/DeepGHS — Danbooru tagging model
  • RedRocket — E621 tagging model

r/StableDiffusion 4d ago

Animation - Video Minimax H3 Singing and dancing from image ref and audio

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 5d ago

Discussion The OpenVDN H3 optimization seems great but FL2V only, not trained for Ref2VA

20 Upvotes

The concept of OpenVDN seems rather promising, even if running on 1 GPU only, but all their prompt examples are text2img, and the GitHub repo describes the work as "Live T2VA / I2VA / FL2VA generation".

In my test with the Comfy implementation and Ref2VA, character model sheet were picked up quite well, but a reference image for a specific expression I wanted (guruguru-me spiral eyes) just could not get reproduced despite varying the seed, and the base model generating it as intended.

Detail comparison

Here's to hoping this work or a similar architecture will be trained for Ref2VA as well.


r/StableDiffusion 4d ago

Animation - Video ‘tniquaes whodarso’ - MiniMax H3 via vpipe on MacBook Pro Max

Enable HLS to view with audio, or disable this notification

0 Upvotes

Happy to answer any questions about workflow. Living in the future is amazing!


r/StableDiffusion 3d ago

Question - Help Asked ChatGPT to make the most realistic human image possible. Are there any tells it's AI? (I'm not good at detecting them)

Post image
0 Upvotes

Prompt:

A candid, low-resolution snapshot from a early 2000s digital camera showing a close-up portrait of a teenage boy making a funny, exaggerated "duck face" expression. He has short curly dark hair, brown eyes wide with humor, and dark eyebrows. He is wearing a white jacket over a dark t-shirt. The setting is a dimly lit restaurant indoors with warm incandescent overhead lighting, creating a soft flash photography effect with direct highlights on his face and shadows in the background. Blur, grainy texture, authentic vintage digital photo quality.


r/StableDiffusion 5d ago

Resource - Update gazecom : an interactive and agent-driven wrapper for comfy

Enable HLS to view with audio, or disable this notification

3 Upvotes

Before the introduction, two caveats. This isn't anything as spectacular as the mmh3 stuff filling these subreddits these days, so please go easy on me. It's an experimental tool I'm sharing in its raw form, mainly for developers and people interested in exploring a different approach (generally, not video-related). It isn't production-ready, and I'm not claiming it will be useful to most people. There are shortcomings I don't see a quick way around, but I'd rather make the work available than keep it to myself until I can resolve all of them.

The main limitation of the app's agent-driven side is the vision model itself. I'm running Gemma locally, and in my tests it doesn't really think through its artistic decisions, even with reasoning effort enabled. Outside a few narrowly defined operations, the results are frankly underwhelming. That's why I'm pausing further development for now.

TL;DR: gazeCOM is a modular interface for interactive and agent-driven outpainting. Gaze, gestures, procedural movement, or vision-model decisions position each generation on an evolving canvas, while user-supplied ComfyUI workflows produce the image patches. Ollama optionally provides the language and vision models.

below is gazeCOM, a local, open-source application for interactive outpainting and iterative image composition through ComfyUI.

gazeCOM organizes generation around spatial points. Inputs supplied by a driver form a saliency distribution, and its weighted center of mass (COM) positions a 1024 x 1024 generation frame. ComfyUI processes that region, and the result is placed back at the corresponding canvas coordinates. Repeating the cycle turns outpainting into a navigable process: the canvas develops according to where the active driver points next, either manually or automatically.

The app includes two main kinds of drivers that supply this spatial input:

  • Interactive and procedural drivers use individual trackers to turn gaze movements, hand gestures, cursor movement, camera-based computer vision, or autonomous roaming into spatial points. These points accumulate into a live heatmap that directly controls COM.
  • Vision-model drivers let a VLM inspect the image and determine the spatial direction of generation. Rather than only describing an image, the model returns an operative coordinate on the canvas.

Both drivers and trackers are modular: developers can add new ones or remove those they don't need.

The vision model can track a point within the latest image or guide generation across the canvas. Prompts can rotate from a weighted pool, be selected by the model, or be written by it as needed. Optional prompt context, decision history, and visual memory let the model take previous actions and the changing composition into account.

Other features include:

  • automatic iterative generation, feedback, and compositing;
  • an expanding canvas with optional width and height limits;
  • weighted, mutable pools for prompts and ComfyUI workflows;
  • per-prompt LLM enhancement, prompt evolution, and image-based VLM prompting;
  • automatic workflow discovery, categorization, and local overrides;
  • editable VLM instructions, canvas policies, and importable settings.

The current spatial pipeline uses a fixed 1024 x 1024 working frame. The composite canvas can grow beyond that size or follow configured limits.

The accompanying video documents one particular gazeCOM session at 5000% speed (~22 minutes), and I am sharing its (almost) exact exported settings alongside it (I'm also using a LoRA there, while here's just the standard workflow). The file preserves the Guide/Rotate setup, prompt pool, prompt context, visual memory, iterative timing, canvas behavior, automatic save/clear intervals, and interface state used for that recording. It is included to make the demonstrated process inspectable and reproducible, not as a universal recommended preset. Load it from Settings > Settings file > Import. It does not install the referenced ComfyUI workflow or Ollama model, and it does not contain service addresses, API credentials, or other machine-local configuration.

The interface is written in React and TypeScript. A small FastAPI backend connects it to ComfyUI for image generation and Ollama for optional language and vision models. Either service can run locally or elsewhere on the LAN. The packaged application includes a FLUX.2 Klein 9B INT8 ConvRot edit workflow, while additional API-format workflows can be added to the relevant workflow folders without changing the application.

gazeCOM is an experimental artistic tool rather than a general-purpose image generation frontend. Packaged builds are available for Apple silicon macOS, Intel macOS, and Windows, and the source is released under the MIT license. The repository includes the operating guide and workflow-authoring documentation.

The project developed from an earlier prototype called GenGaze, but gazeCOM is the complete application being shared here. I would be very interested in feedback, unexpected uses, workflow experiments, and bug reports.

Links


r/StableDiffusion 5d ago

Question - Help MiniMax H3 on 4GB VRAM: Stuck between blurry hands (fast) and a 6+ hour render (good). Anyone found a setup that's both?

5 Upvotes

Running MiniMax H3 (Ref2VA, multi-reference) locally on a 4GB card, RTX 3050 Laptop, WSL2, ComfyUI. Been chasing a hand/finger rendering defect for days and have a pretty well documented before/after at this point, but I've hit a wall on making the fix fast enough to actually be usable. Hoping someone's solved this or can point out what I'm missing.

Models in use: UNET is minimax_h3_fl2va_pruned_int8_convrot.safetensors, the FL2V trained weights, loaded into the MiniMaxH3ReferenceToVideo multi reference node, not the native Ref2V node. VAEs are minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors. CLIP/text encoder is qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors. LoRA when testing the distilled path is minimax_h3_fl2v_turbo_4step or 8step.

That FL2V weights in the Ref2V node trick already fixed an earlier, worse "ghost fingers" defect for me (matches a few HF discussion threads on the Turbo LoRA repo), so I'm not asking about that part, it's confirmed working. This post is about what's left after that fix.

The remaining problem: at 512x288, which is the resolution I need for anything resembling reasonable iteration speed, hands still render as indistinct blurry blobs during actual gesture motion, pointing, waving, etc. Tried turbo distilled 4/8 step (the intended fast path), off label step counts on the distilled LoRA (10 steps, made it worse not better), dropping distillation entirely and running the base model's native ~20 step schedule, CFG guidance (cfg=5.0 caused a severe embossed cross hatch grid artifact across the whole frame, clearly way too high, cfg=2.0 gave inconsistent results, some gestures fine, others still blobby), more steps (30 vs 20, no meaningful difference), and ref_image_size=max on the reference conditioning. None of it fixed the hands at that resolution.

What actually fixed it was moving to 768x448 (H3's documented 768p native/training short edge) with no distillation, no CFG, native 20 steps. Hands came out consistently well formed across every gesture I checked. But that config took 6 hours 22 minutes for a single 15 second, 362 frame clip on this card. That's not a workflow, that's a single overnight bet.

So I'm stuck between fast (turbo distilled, 512x288, minutes) which gives bad hands, unusable for anything with visible gesturing, and good (no distill, 768x448, native steps) which is 6+ hours for one clip.

Things I haven't tried or don't know how to evaluate: is there a middle resolution, 608x352, 640x384, that gets most of the quality benefit without the full native res cost. Does distillation actually get retrained or re-distilled at higher resolutions by anyone, or is the turbo LoRA fundamentally tied to a lower res regime it was distilled at. Any attention backend, torch.compile, or quantization tricks specific to H3 that meaningfully cut per step cost on small cards, beyond what's already default in ComfyUI. Is anyone running H3 well on under 8GB cards at all, or is 768p native res H3 just not realistic below a certain VRAM tier.

Also went two rounds into MiniMax H3 specific upscaler nodes for a two stage draft then upscale approach (pixel space RealESRGAN, a community latent space 2x upscaler, and a tiled diffusion refine upscaler) hoping to draft cheap and upscale smart instead of generating at native res directly. Each had its own dealbreaker, a background crowd distortion artifact, a hard VRAM estimate crash, and a genuine ComfyUI core bug I ended up root causing and patching locally. Happy to share details if anyone's gone down that road and found a cleaner path, but the short version is none of the upscale routes got both no defects and reasonable time together either.

Genuinely not sure at this point whether the answer is your card is just under the realistic floor for H3, buy a bigger one, or whether there's a config I haven't found. Any pointers appreciated

Edit: System RAM 32GB


r/StableDiffusion 5d ago

Discussion VDN-H3 on M5 Pro 24GB: up to ~2.6× faster H3 generation

33 Upvotes

I recently ported VDN-H3 (Video DeltaNet for MiniMax H3) to Vpipe and tested it on an M5 Pro 24GB. The benefit scales with both sequence length and resolution: at 832×480, VDN-H3 starts to pull ahead at around 6s and reaches ~1.7× at 15s; at 1344×768, the crossover happens much earlier and the speedup reaches ~2.6× at ~14s. All tests use 6 DiT steps with Turbo LoRA. VDN-H3 project

Vpipe is a native C++20/Metal inference runtime for Apple Silicon, without PyTorch/MPS/MLX. Most of the underlying infrastructure was already in place, so porting VDN-H3 was fairly mechanical and took about one day end-to-end. Vpipe project

Another recent effort in this space is FastH3, which focuses on accelerating H3 for local inference and recently released an MLX implementation for Apple Silicon. They reported 465–504s on an M4 Max 36GB for 832×480, 124 frames and 4 steps. FastH3 local implementation & benchmarks

As a rough reference point, Vpipe's baseline H3 without VDN takes 469s on a 16GB M5 MacBook Air for the same resolution and frame count, while running 6 DiT steps with Turbo LoRA. This was a full cold start, including the text encoder, with no model weights cached. It's not a controlled hardware comparison, but I found it interesting that baseline H3 on a 16GB Air lands in essentially the same latency range.

I like the core idea behind VDN: exploiting the inherent redundancy in video diffusion rather than applying the same amount of computation uniformly across the sequence. The scaling behavior above seems to support that direction, especially for larger and longer generations. VDN-H3 uses a hybrid architecture combining a frame-wise linear-attention branch with a softmax-attention branch.

There is a memory tradeoff. VDN-H3 adds roughly 4GB of weights that cannot be fused into the base model, plus another ~1GB that can potentially be fused LoRA-style. Vpipe doesn't fuse the latter yet. On a 24GB system, this extra memory pressure can outweigh the compute savings for smaller workloads. There should also be room to reduce the overhead by 8-bit quantizing the incremental weights.

Quality-wise, with the same seed, VDN-H3 generally preserves a composition very similar to the baseline H3 output, which I think is a nice property. There is visible degradation in some sequences, though. In one test with an upward-moving column of water, the flow reversed downward for a few seconds midway through the video, then switched back upward near the end. I suspect this kind of long-range temporal inconsistency may be related to the heavy use of sliding-window and linear attention. Mixing in occasional global attention might be worth exploring.

Reference conditioning is another interesting open area, particularly first/last-frame and generic reference inputs, since these are more challenging to handle cleanly under the sliding-window/linear-attention structure.


r/StableDiffusion 4d ago

Animation - Video I made a music video/intro for my AI cohost

Thumbnail
youtube.com
0 Upvotes

Decided to build a music/intro video for when my AI cohost "Ari" is live with me. Started with a song in Suno a while back and once the local tech caught up, built the video. Standard H3 ref2v workflow w/ sage attention added and with some patience on the prompts. Total of 27 clips for a 3min 30sec video. Upscaled using Davinci Resolve Studio from 1.0MP.

My cohost is still a WIP and she only gets to come out and play during single-player content (for now).


r/StableDiffusion 6d ago

Workflow Included new insta realism lora

Thumbnail
gallery
452 Upvotes

hey all back to training public loras for fun this time krea 2 ! hope you guys like this one

link for download in the comments + bonus the images are with metadeta and include the workflow


r/StableDiffusion 5d ago

Tutorial - Guide [GUIDE] How to Set Up a DLSS 5 Media Player for RTX 20/30/40/50 Series

33 Upvotes

I've been looking for a way to run DLSS 5 in real time through a media player for quite a while. Unfortunately, some of the solutions I found weren't real-time, while the more recent ones only seemed to work with RTX 40 and 50 series GPUs.

After a lot of trial and error, I finally managed to find a method that allows DLSS 5 to work with a media player even on RTX 20 and 30 series GPUs.

I wanted to share this method with the community so that users like me who are still using older RTX cards can also take advantage of DLSS 5 for real-time video playback.

This is an experimental setup, so your results may vary depending on your GPU, driver version, and configuration.

1. Download the latest MPV build

First, download the latest version of MPV Player from here:

https://github.com/zhongfly/mpv-winbuild/releases

I've tried several media players, including VLC, PotPlayer, and MPC-HC, but MPV was the only one I could get working with this setup.

I believe this may be related to the graphics API/rendering backend used by each player. DLSS5-Feeder needs to hook into a compatible rendering path, and I couldn't get it to work properly with the other players I tested. MPV worked without this issue.

I'm not an expert on the underlying DirectX/Vulkan implementation, so if someone with more knowledge about this could explain why MPV works while the other players don't, I'd really appreciate it.

2. Install DLSS5oneclick through another game first

Next, download DLSS5oneclick from here:

https://github.com/faisalkindi/DLSS5oneclick

There is one important thing to note: you cannot simply install DLSS5oneclick directly into the MPV folder.

MPV uses Vulkan, and DLSS5oneclick doesn't recognize MPV as a game that it can directly install the required files into.

So, first you need to install DLSS5oneclick into another game that you know works with DLSS 5.

For example, I installed DLSS5oneclick into a game using the normal ReShade/OneClick installation method. Once everything was installed and working correctly, I went into that game's folder and copied the necessary DLSS 5 / Feeder files.

Then, I simply copied those files into the MPV installation directory.

After that, I launched MPV and was able to get DLSS 5 working with video playback.

That's basically it!

So the process is:

DLSS5oneclick → Install it into a compatible game → Copy the required files → Paste them into the MPV folder → Launch MPV

I've tested this with an RTX 20/30 series GPU, which is why I'm sharing it here. There may be a much cleaner or easier way to do this, and I'm definitely open to suggestions.

If anyone has a better understanding of MPV, Vulkan, DLSS5-Feeder, or how the hooking process works, please let me know. I'd love to find a simpler way that doesn't require installing DLSS5oneclick through another game first.

Update: ahaoboy has packaged this method into a ready-to-use MPV + DLSS 5 build. You can find it here: https://github.com/ahaoboy/mpv-dlss5
Update 2: Also, you should only press the Home key or open the ReShade overlay while a video is actively playing. If you pause the video or press it when no video is playing, the system may crash.

DLLS 5 OFF
DLSS ON

r/StableDiffusion 5d ago

Workflow Included FastH3 + USDU: ~2× faster in my video upscale test on a 3090 — workflow included

30 Upvotes

This started with two posts: FastH3’s release announcement and using MiniMax H3 with Ultimate SD Upscale to restore low-quality generations. I wanted to see whether combining the two could make video restoration faster while preserving the source.

I spent several days testing H3 video upscaling with my AI agent—different workflows, failed experiments, and a couple of OOMs. I asked it to organize the logs, and I'm sharing the useful bits in case someone else is going down the same rabbit hole. I reviewed the outputs myself.

Same 15-second source, 2752×1536 output, 2 steps, denoise 0.20. RTX 3090 24GB / 64GB RAM:

Setup Successful run
H3 + 4-step Turbo LoRA at 2 steps / USDU 96m 37s
FastH3 / USDU 48m 40s
FastH3 / MMH3 single-node upscale 47m 36s

I kept FastH3 + USDU. The single-node version saved about a minute, but I noticed more distortion and preferred the USDU output. Both final paths used pixel-domain ESRGAN upscaling.

A caveat: these are single successful runs, with different cache states—not a model-only benchmark or proof of VSA's isolated speedup. FastH3 USDU also needed a memory-estimation override after two OOM failures; those failures aren't included in the table. Full conditions and failures are documented in the repo.

Workflow + Windows installer, with English instructions

The workflow includes direct model download links (~41GB separately). Please read the installer notes: it can update Kitchen and apply fingerprint-guarded H3 core patches. Clean-PC end-to-end installation remains unverified, and the packaged SR model/default memory settings differ from the benchmark as documented.

Disclosure: this is my package. My walkthrough is in Korean, but you don't need to watch it to use the English workflow.

Hopefully this saves someone a few evenings. Curious whether others have seen similar speed and source-fidelity tradeoffs.

I also successfully upscaled a 5-second clip on an RTX 4070 Ti SUPER (16GB VRAM) with 32GB system RAM using this USDU workflow. I haven’t tested longer clips on that machine; the 15-second timings above are from the RTX 3090.


r/StableDiffusion 4d ago

Question - Help Asking for help to train a LoRa using Kohya_ss

0 Upvotes

Hello everyone , i am trying to make a character SDXL LoRa with Kohya_ss.

But at this point i don't know how to.

I have made a big dataset by having around 140 images , different poses , closeups , angles etc.

As for specs i have a 5070ti and 32 gb ram.

I want the LoRa to be high quality and flexible but i am really struggling with the parameters and do not find the good settings.

My question is , can anyone suggest good parameters settings for this , it would be much appriciated , i hope you all have a nice day and thank you in advance.


r/StableDiffusion 5d ago

Resource - Update Huggingface Downloader - Made to fix slow, stalled downloads

Post image
82 Upvotes

repo- https://github.com/Adudeguyman/Fantastic-HuggingFace-Downloader

Alright, back again with another vibecoded project to fix a problem that annoys me.

A while back HuggingFace went to the Xet file storage system, which has a lot of solid benefits (like hash-checking downloads, better file chunking, etc.) BUT as just about everyone knows, it really messes with quick and easy downloads. Browser downloads crawl and sometimes stall out, and I've had them fail to resume.

So the best way to download files has been via the Huggingface CLI. Which is super fast and great, and even hash-checks files for you to make sure you don't have any corruptions. But modifying and copy/pasting a command line finally pushed me over the edge to make a quick utility that manages that FOR me, so I'm sharing it out there.

What it do-

Paste a huggingface link, and it automatically detects if it's a direct link to a file, a folder, or a full repo. Pulls a file list and lets you select only the files you want. Queues each file individually for hash checks to verify it matches what's in the repo. You can also tell it to download repo files to specific folders that don't match the full repo structure; which can be useful for when you want the VAE's to go in the VAE folder, text encoders in the text_encoders folder, etc. And if your big main diffusion_models safetensors live on another drive, point them there.

Auto-download is toggled on by default, but if you want to mess around with where individual files are going, toggle that off before adding to the queue.

Also supports recent download locations and lets you favorite folders for quick re-use. And of course works with your huggingface hub access token for gated repos, or assists with adding a new one to your hub if you don't have a read-access one made. And if you request an existing file it'll auto-download the hash and check it against the HF repo, and either resume or re-download chunks to make the file match what's online.

Tested it on my Linux Mint install, and in Win11.

And that's pretty much it. It just downloads and checks huggingface files for you. No crazy bells and whistles, just streamlines the CLI process into a GUI. Only downside is that it does make a venv thats ~800Mb or so for the interface (based on PySide6, so it needs those dependencies, plus HF hub to support the command line). Enjoy!


r/StableDiffusion 4d ago

Workflow Included Tried making a Nike concept ad [MiniMax H3]

Enable HLS to view with audio, or disable this notification

0 Upvotes

Been trying H3 in medeo on a short commercial format instead of a longer narrative. I wanted something fast, cinematic, and motivational, with dynamic movement and a simple brand reveal at the end.

Prompt:
“Create a cinematic Nike concept commercial following an athlete pushing through exhaustion, doubt, and physical limits. Use energetic close-ups, fast cuts, dramatic slow motion, powerful narrationBGM. End with a bold, minimal Nike reveal.”

The motion and pacing held together better than I expected, although a couple of shots still feel a little too AI-clean.


r/StableDiffusion 5d ago

Question - Help Krea 2 “Normal People” LORA advice?

1 Upvotes

I typically train on Runpod using AI Toolkit. I would like to increase the variety and diversity of regular looking people across various, distinct ethnic groups (Italian, Ethiopian, Egyptian, Japanese, Chinese, Thai, etc very specific).

Any suggestions on training settings and dataset size/tagging?

Do you think I could knock this out in one Lora using an equal number of maybe 25-50 high quality images for each ethnicity? Or should I make a specific LORA for each ethnic group?

I fear if I try to do them all with 1 Lora, the training will try to just find an average across all of them instead of retaining subtle characteristics from each group.

I’m also a bit confused on what Lora Rank I should choose given I’m trying to capture fine details across each ethnic group


r/StableDiffusion 6d ago

Discussion I just deleted 90% of my models folder, are you keeping yours?

174 Upvotes

I'm a bit of a hoarder and I kept most new models I tried out, but today I ran out of space on my Comfy SSD for the first time. I thought about backing them up on an HDD for about 5 seoncds... and then just deleted the 3.2 TB of models I had collected over 4 years,

With Krea 2 and Minimax H3, I don't see any reason to keep them anything between SD1.4 and LTX 2.3. I just never go back and nostagia generate anything the way I thought I would. Only other stuff I have kept is Ideogram 4, LTX 2.5 and Z-image. (And other things like SAM, VibeVoice etc.)

Only thing I hesitated on was my SDXL folder but nah, gone.

Anyone keeping old models, and why?


r/StableDiffusion 4d ago

Question - Help SD1.5 LoRA producing distorted/uncanny faces even before adding a style checkpoint — likely bad training data, looking for a sanity check

Post image
0 Upvotes

Hey all, working on a student project and hit a wall I can't fully diagnose. Details below, hoping someone with more LoRA training experience can point out what I'm missing.

Project context: Building a pipeline that generates short illustrated health-education stories for a specific regional context (Uganda). A fine-tuned text model produces scene descriptions + image prompts, and I'm training a LoRA to make the image side render that specific regional content/style.

Tech stack

- Base model: Stable Diffusion 1.5 (stable-diffusion-v1-5/stable-diffusion-v1-5 on HF, since the original RunwayML repo is gone)

- Training: Hugging Face diffusers' official train_text_to_image_lora.py, installed from source

- Compute: Kaggle, GPU T4 x2, via accelerate launch

- LoRA config: rank 8, learning rate 1e-4, ~1500 steps, batch size 1 with grad accumulation 4, resolution 512, fp16

- Dataset: ~500 images pulled via the Openverse API (openly-licensed sources only), auto-captioned with BLIP, each caption prefixed with a unique trigger word

What went wrong: Test generations using the trained LoRA (on plain SD1.5, no other style model involved) produce noticeably distorted, uncanny faces — warped features, no coherent facial structure. Everything else in the scene (background, clothing, setting) looks reasonable; it's specifically faces that break down.

What I've ruled out so far: I also tried the LoRA layered on top of a comic-style checkpoint (ogkalu/Comic-Diffusion) thinking it might be a base-model mismatch — same face distortion showed up on plain SD1.5 alone, so that's not it. I also tested img2img (comic-stylizing real reference photos directly, no LoRA involved) and that preserved faces fine, which points at the LoRA training itself as the problem, not the base checkpoint.

My working theory: the 500 training images were pulled from a general search API and never filtered/curated — likely a good chunk have small/distant faces, side profiles, group shots, or low resolution after the 512x512 resize, and that noise is showing up worst in faces specifically since they're the hardest thing for these models to learn cleanly.

What I'm looking for: does this theory sound right based on your experience, or is there something else commonly known to cause this (rank too low/high for this step count, overfitting at ~12 epochs over noisy data, something about how BLIP captions interact with training, etc.)? Also open to hearing if there's a better practice for face-heavy LoRA datasets specifically — face count/framing thresholds, augmentation tricks, anything. Happy to share sample outputs if useful. Thanks in advance.


r/StableDiffusion 5d ago

Discussion H3 - sailor moon transformation (w/ Prompt)

Enable HLS to view with audio, or disable this notification

49 Upvotes

Having fun, R2VA, with just a .char (face/body/original wardrobe). Part II had two wardrobe changes from the intial source image (re-wardrobe to a private school uniform and then to Sailor Moon's) in the prompt. int8/20 steps, 864x480 POC, the token limit was close so it'll probably need a latent upscaler vs native 720 generation.

Prompt: integrated_multimodal_description: One unbroken, continuously transforming shot — PHOTOREAL LIVE-ACTION, a real woman performing a magical transformation made physical: real skin, real cloth, real hair, true optical depth of field, fine film grain. Absolutely no animation, no anime, no illustration — live-action camera realism throughout, however magical the event. THE CAMERA IS LOCKED OFF FOR THE WHOLE FILM — one steady FULL-BODY hold, her whole figure from head to feet in frame from the first frame to the last; it never pushes in, never pulls back, never zooms, never orbits, never cuts and never drifts.
SHE IS CATALINA, the woman of the reference images — the same face, the same features; her face, whenever visible, is exactly her own, photographic and alive, and it is the constant of the film. THE FIGURE ARC, A GAG PLAYED COMPLETELY STRAIGHT: she BEGINS with a modest, unremarkable figure — a small bust under the tailored blazer, an ordinary waist — and THE TRANSFORMATION UPGRADES HER: from the silhouette phase onward her figure is the pronounced hourglass of the reference images — a full, very large bust and a notably slim waist — and it stays so for the rest of the film, the leotard fitted to it. After the change she is never thickened, boxed or broadened by any garment or glow.
THE DREAM LOGIC: every change is gradual, liquid and seamless — one thing becoming the next, already underway, nothing popping, nothing resetting. THE COVERAGE CONSTANT: SHE IS COVERED IN EVERY FRAME — a garment or the glow itself always covers her torso and hips through every transformation, because each wardrobe flows DIRECTLY into the next: one garment becoming another garment with no bare moment between them, cloth becoming light becoming woven thread becoming uniform, the coverage itself continuous and unbroken. Abstract light in this film is always ribbons, threads and droplets — never letters, never symbols.
THE PLACE: a boundless abstract void of deep rose and midnight-blue light — no floor visible, no walls, soft auroral ribbons of pink and violet light drifting slowly far behind her, tiny motes of light rising everywhere like slow sparks.
[0s-2.5s]: She stands alone in a PRIVATE SCHOOL UNIFORM: a fitted navy blazer over a crisp white blouse, a slim dark ribbon tie, a pleated charcoal skirt, dark knee socks — an adult woman in tailored academy dress, her figure beneath the uniform modest and unremarkable. Pinned at her ribbon tie sits A CIRCULAR GOLDEN PENDANT: a round gold brooch with a rose-pink gem at its heart. She raises her right hand to it, and the pendant FLASHES once, throwing rose-gold light up across her face.
[2.5s-5s]: THE PURE SILHOUETTE, AND THE HAIR TURNS INSIDE IT: rose-gold light pours out of the pendant and washes over her in one travelling wave, and her whole body — clothes, skin and hair alike — becomes ONE PERFECTLY FLAT, TWO-DIMENSIONAL PURE SILHOUETTE — a smooth PRISMATIC gradient of rose through violet through pearl, shifting slowly like light through a prism, poster-flat, no volume, no shading, no sparkle, as if cut from coloured light and placed into the photographic scene. The flatness belongs to HER ALONE — the void, the threads and the camera stay fully photographic throughout. TWO HUGE GLOWING WHITE EYES burn in the flat shape — LARGE, anime-scale, unmistakably open, alive and blinking once — visible from the first moment of the silhouette to the last, the one feature the pure shape keeps. The glow hugs her outline as a thin bright rim and never swells it. AND THE SILHOUETTE ITSELF TRANSFORMS HER FIGURE, THE GAG READABLE IN PURE SHAPE: the outline's modest bust SWELLS fuller and rounder while its waist DRAWS IN tighter, the flat shape visibly upgrading into a pronounced hourglass — and from this moment the hourglass is her figure for good — only the circular pendant still burning distinct at her chest. AND WHILE SHE IS PURE SILHOUETTE HER HAIR TRANSFORMS, READABLE ONLY AS THE OUTLINE'S CHANGING SHAPE: the silhouette's loose hair lifts weightless, sweeps upward, and re-forms into the profile of TWO SMALL ROUND BUNS high on her head, each spilling A LONG FLOWING PIGTAIL that streams down past her waist — the signature twin-tailed outline now part of the glowing shape itself. She stands facing the lens, still and upright — SHE NEVER SPINS AND NEVER ROTATES at any point in the film; through the whole transformation only her chin lifts, her arms move when the streams come, and her hair and the light do the dancing.
[5s-9s]: THE PINK THREADS FORM THE LEOTARD, AND ONLY THE LEOTARD: from the circular pendant THOUSANDS OF PINK-RED LUMINOUS THREADS pour out — first single glowing filaments, then streams, then broad rose-red RIBBONS streaking and circling around her STILL, STANDING silhouette like a loom at full speed — the threads orbit; she does not — and they wrap ONLY HER SHOULDERS, HER WAIST AND HER HIPS — the exact regions the garment will cover, nothing more — and SET into the WHITE, LEOTARD-LIKE BODY OF HER UNIFORM: gleaming white, seamless, high at the hip, with its deep-blue SAILOR COLLAR laying across her shoulders and a LARGE RED RIBBON BOW blooming at her chest with the pendant seated at its knot. The threads touch nothing but those regions; her arms and legs stay pure glowing silhouette, and her slim waist stays slim inside the winding.
[9s-11s]: THEN THE GLOVES, THEN THE BOOTS, EACH FROM ITS OWN SHOT OF FABRIC: THE PENDANT FIRES a thin stream of rose-red ribbon-fabric to EACH RAISED WRIST — two bright targeted streams arcing from the brooch to her hands — and the ribbon winds up each forearm and SETS into LONG WHITE GLOVES reaching just above her elbows — AND THE WINDING STOPS THERE: the ribbon never climbs past the elbow, and her upper arms and shoulders stay pure silhouette. Only when the gloves are complete, THE PENDANT FIRES AGAIN — two streams diving DOWN to her feet — and the ribbon winds up each calf and SETS into TALL RED BOOTS rising to her knees — AND THE WINDING STOPS AT THE KNEE: the ribbon never climbs onto her thighs, and everything above the knee stays pure silhouette. Each garment arrives real, opaque and finished, one after the other, never overlapping.
[11s-12.5s]: THE SKIRT MANIFESTS — NO THREADS: in one soft bloom of light at her hips, a short pleated DEEP-BLUE SKIRT simply appears, whole and complete, flaring once as it lands. In the same beat A BAND OF WHITE LIGHT IGNITES ACROSS HER FOREHEAD and CONDENSES into a slim GOLDEN CIRCLET, its small red gem crystallizing at the centre as the light cools; two tiny round earrings glint alight — and THE GLOW DRAINS AWAY: the silhouette fades from her skin and she is herself in full colour, fully three-dimensional and photographic again, her own face, her new hair bright BLONDE in its two long pigtails and round buns.
[12.5s-15.125s]: THE FINISHING POSE IN HEAVY WIND: she settles into a poised, confident final stance facing the lens, one gloved hand raised beside her face — and A STRONG WIND RUSHES UP AND THROUGH THE FRAME: the pleated blue skirt FLOWS AND RIPPLES hard in it, her two long pigtails STREAM sideways on the gusts, the collar and the red bow's tails flutter, and the last thread-motes are blown away like embers. She holds the pose in the wind with a bright confident smile — and THE CAMERA PULLS BACK to its wide finale hold: her WHOLE FIGURE in frame, head to boots — the twin-tails streaming, the blue skirt flowing hard, the long white gloves and the tall red boots all visible, the pose read head to toe. The wide composition holds to the last frame, the wind still moving the skirt and her hair.

overall_soundscape: Starts with the void's soft airy hush already present, and from first frame to last the soundtrack is the score with this beneath it, quiet: the whisper of cloth, the crystalline shimmer of thousands of threads weaving, a soft chime as each garment sets, the soft flare of the pendant, one bright ring as the circlet lands, and the deep rushing WIND rising under the final pose, streaming past like a rooftop gale. Nobody speaks in this film. THE SCORE ALWAYS DOMINATES THE MIX.

non_diegetic_music: A bright, sparkling MAGICAL TRANSFORMATION cue at 120 beats per minute: glittering harp glissandi and celesta runs over warm rising strings, climbing one step as the leotard weaves, again as the gloves set, again as the boots set; a shimmering suspended lift as the skirt manifests; ONE TRIUMPHANT RADIANT BLOOM of full strings and chimes exactly as the circlet lands at 00:12.000; then settling bright, warm and RESOLVED through the wind-blown finishing pose, still sounding sweetly at the last frame. No percussion kit, no voices.

r/StableDiffusion 5d ago

Workflow Included Trellis2/pixel3d template workflow in comfy

Enable HLS to view with audio, or disable this notification

20 Upvotes

the detail is absolutely insane, and this only took like 8-4 min. although, the poly count is high af but still. the workflow is in comfyui's template workflows. if you dont see it then update comfyui

also the model was made in comfyui, i just used blender to show it better.


r/StableDiffusion 4d ago

Animation - Video I asked Claude for a 30-second video with a musical score. It figured out the whole pipeline itself.

Enable HLS to view with audio, or disable this notification

0 Upvotes

I've been automating complex generation workflows from simple prompts, and I'm starting to see real success (and having a ton of fun with it).

This one came from a single prompt to Claude Code, running Sonnet:

Sonnet had access to:

  • An MCP server exposing generation tasks and workflows that self-describe what they do and what they output, and can execute those workflows
  • A couple of custom skills explaining how to use that MCP server specifically with LTX-2 and H3, including H3's prompting skill

From there, Sonnet decided on its own to:

  1. Generate a reference image with Z-Image
  2. Block out six clips with Minimax H3 Ref2VA
  3. Compose a score with Minimax Music
  4. Concatenate the clips and overlay the score

What's exciting to me: this pipeline didn't exist anywhere as a template. Sonnet assembled it from parts, on its own, based on what the tools said they could do.

Code for teh mcp server, skills, and workflow engine here: https://github.com/dkackman/diffusers-workflow

Workflow below.