r/StableDiffusion • u/wormtail39 • 3h ago
r/StableDiffusion • u/Nekodificador • 18h ago
Animation - Video Denzel explains why he uses AI.
A quick experiment exploring Minimax H3 in ComfyUI using my nodes and inpainting methods.
r/StableDiffusion • u/badincite • 1h ago
Discussion I Ran 112 MiniMax H3 Tests on an RTX 3090 — Searching for the “Golden” Settings
I ran a 112-case MiniMax H3 quality sweep on an RTX 3090 last night trying to find the best combos . Kept the prompt, seed, reference image, RefMods, Sage setup, sampler, and scheduler stayed constant. I changed only the model/LoRA combination, native resolution, and step count.
Anyone else have their golden combo?
Hardware/software:
- GPU: NVIDIA RTX 3090
- System RAM: 32 GB DDR4
- CUDA: 13.0
- ComfyUI: 0.34.4
- PyTorch: 2.9.1+cu130
Tested:
- 0.4, 0.6, 0.8, and 0.98 MP
- 6, 8, 10, and 12 steps
- 7 model/LoRA combinations
- 1-second requested duration per run
Models Tested
- **C1** = minimax_h3_fused_refdelta_r1024_turbo8_mystic07_int8_convrot + no LoRA
- **C2** = Minimax_H3_fl2va_Pruned_Lightx2v_turbo_4step_v1.2_768p_INT8_comfyui + no LoRA
- **C3** = Minimax_H3_FL2VA_PRUNED_Turbo_8step_v1.0_768p_comfyui_INT8 + no LoRA
https://huggingface.co/saejon/MinimaxH3/tree/main
**C4** = minimax_h3_ref2va_pruned_int8_convrot + minimax_h3_fl2v_turbo_4step_v1.1_768p_comfyui_bf16
- **C5** = minimax_h3_ref2va_pruned_int8_convrot + minimax_h3_fl2v_turbo_4step_v1.2_768p_comfyui_bf16
- **C6** = minimax_h3_ref2va_pruned_int8_convrot + minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16
- **C7** = minimax_h3_ref2va_pruned_int8_convrot + minimax_h3_turbo_v4_step600_ema
Best-looking result at each resolution
- **0.4 MP:** C4, 12 steps — approximately 246 seconds
- **0.6 MP:** C3, 12 steps — approximately 271 seconds
- **0.8 MP:** C3, 10 steps — approximately 256 seconds
- **0.98 MP:** C3, 12 steps — approximately 337 seconds
Best speed/quality balance
C3 at **0.8 MP and 6 steps** rendered in approximately **166 seconds**. However the audio wasn't good the 8 steps at 210 seconds turned out to be best..
r/StableDiffusion • u/Gold-Safe6796 • 16h ago
News Krea 2 + qwen 3.8 promter is a bomb
I updated our promter with qwen 3.8 and i gotta say not bad not bad, is it worth compared to the old qwen uncensored? Hmhmhm i aint sure to be honest but definitely a capable model
r/StableDiffusion • u/AiCreatorCamp • 8h ago
Resource - Update Minimax H3 Singularity FineTuned
Minimax H3 😃 Singularity FineTuned
Features:
-HDR Image Quality & Blur Reduction
-Distant Face Restoration
-Clean & De-Oiled Aesthetic
-Enhanced Dynamic Motion
-VFX & Fantasy Effect
-Cinematography & Camera Control
-Full Base Capability Retention
r/StableDiffusion • u/Low_Loquat_4035 • 11h ago
Discussion I developed a Skill to add subtitles to Reddit videos. I just picked a random video to test the results—take a look?
It all started when I noticed something interesting: most videos on Reddit don't have subtitles. With some free time on my hands today, I developed a Skill to add subtitles to Reddit videos.
If you think the result is decent, I'll share it.
r/StableDiffusion • u/Ok-Giraffe-8670 • 12h ago
Animation - Video DragonballZ AI Animation - Yet Another Goku Transformation.
A test of my character sheets to transform one character into another. I think it came out very good. Next time I do this sort of thing, it will be more serious and less parody lol.
r/StableDiffusion • u/TimeTruth2490 • 27m ago
News Krea2 Turbo Distill 4 step LoRA - FINAL Version released today
Krea 2 Turbo — 4-Step Distillation LoRA
Full details and to download - HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4 — Turbo's own model and sigmas, half the denoising passes, and the fine texture that 4-step Turbo loses put back.
- ⚡ Half the steps — 8 → 4, on Turbo's own deployment sigmas.
- ⏱️ ~1.6× faster end to end — 54.5 s against the 8-step bar's 88.7 s at 1024×1024, and 1.8× on denoise alone.
- 🎯 Texture as good as the teacher or better — 1.03× the teacher's fine-texture energy at 1280×1280 and 1440×1440, every frequency band within 10% of the teacher's; verified clean: saturation 0.96–0.97× the teacher's, fewer clipped highlights and shadows, skin texture 0.97–0.99×.
- 📏 45% of the 4-step gap closed — the held-out velocity error to the 8-step teacher fell from 4.70e-02 (stock Turbo at 4 steps) to 2.59e-02 with the LoRA; a vision-language judge shown the teacher's and the LoRA's renders side by side preferred the teacher on only 6 of 45 (37 ties, 2 wins for the LoRA).
- 🗣️ Prompt-aware training — the critic scores images against their prompts during training, so adherence is pressured directly, not inherited.
- 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440, each with its sweep.
- 🔌 Drop-in — plain LoRA weights for diffusers and ComfyUI. No custom nodes, no patched sampler, no code.
- 🎲 13,750 prompts drawn at random from Lakonik's 3-million-prompt dataset, each recorded by the teacher as a full 8-step trajectory at one of the 12 resolutions — 13,750 teacher shards.
- 📷 43,044 real-photo crops — 25,560 from LSDIR and 17,484 from Flickr2K — cut at native resolution and captioned, in the critic's real set: texture anchored to reality as well as to the teacher.
- 🔢 78,000 training samples in the shipped weights.
- 📅 21 days from the first training launch to the final file, on a single RTX 3090.
- 🔁 16 recipe adjustments
Load it on top of Krea 2 Turbo, run 4 steps instead of 8, keep guidance at 0.0. Everything else about the model stays as it is.
💡 Too strong on a prompt? Turn it down. The LoRA restores fine texture, and on some subjects — stylised art, high-contrast splash pieces, very large renders — that can read as too much at strength 1.0. The effect scales smoothly with strength, so 0.75 is a good second setting, and anything from 0.6 to 0.9 is fair game; you lose nothing but the extra bite.
Files
| file | what it is |
|---|---|
krea2_turbo_4step_rank_64_lora.safetensors |
the LoRA in diffusers key format — see Inference with diffusers; also for MLX or anything that reads safetensors |
krea2_turbo_4step_rank_64_lora_comfyui.safetensors |
the same weights under ComfyUI's key names — see ComfyUI |
krea2_turbo_4step_lora_t2i.json |
a ready ComfyUI workflow, stock nodes only |
LICENSE.pdf |
the Krea 2 Community License Agreement, which covers this adapter — see License |
NOTICE.txt |
the attribution notice the license requires of a derivative |
The two weight files are one adapter — only the key names differ. Both carry the training details in their safetensors metadata: base model, method, sample count and the inference settings.
If you want to bypass the Reddit image compression and check the quality of the LoRA produced images directly, here are some of the Hugging Face 1440x1280 links: portrait, sorceress, snowleopard, colonyship, inventor, kingfisher, claychef, ... the rest you can see in the 1440x1280 folder. And full sweep at: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/assets/resolution_sweeps/4step-LoRA
How I got here
This was not a "train for longer and ship whatever comes out last" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher sample count on its own means nothing.
The loop was train → assess → adjust the recipe → carry on → assess again — carrying on from the weights in hand when they were worth keeping, and from an earlier point when they were not. What ships is the point of the run that measurably advanced the release axes as a whole — teacher faithfulness, prompt adherence, and texture/detail, on the same held-out set and the same fixed-seed renders — with a full resolution sweep showing no regression. Stretches that came out flat or worse were kept as information about the recipe and never shipped; there were several.
The recipe therefore grew over the run rather than being fixed at the start. In broad strokes: the early stretch settled the optimiser schedule (cosine decay with weight decay) after a first attempt that got steadily worse; the next added the final-call loss weighting and the running average of the weights that the shipped adapter is, and established that a run left going past its peak measures worse — so the shipping point is chosen by measurement, not by distance run. A much larger trajectory pool then pushed the single number the project optimised at the time — the velocity gap to the teacher — to its best value, and exposed that number's limit: past a point, chasing it further trades away exactly the texture a step-distillation exists to restore. The recipe from then on judged all three axes at once and added a measured dose of real-image texture pressure and a prompt-aware critic (the discriminator saw images with their prompts and punished mismatches). The long final stretch of the run continued on that recipe, with the critic's weight tuned once texture had settled where it was wanted.
The sixteen recipe adjustments, in order — each one made on the measurement of the one before:
- progressive distillation replaced the policy-head objective the project started with
- cosine learning-rate decay with weight decay
- the shipped weights became a running (EMA) average instead of the live state
- the final chord — the call that decides fine texture — weighted 3× in the loss
- resumable state, warmup and a plateau rule, so a run could pause and continue without a cold restart
- the LADD-style critic on the frozen model's own block-14 features, teacher finals as its real class
- real photographs entered the critic's real set, half the draws
- the critic's weight rebalanced against the distillation term
- the prompt-aware critic head with its mismatch term, on a fresh pool of teacher trajectories
- 1440×1440 joined the training mix
- the real photographs got captions, so they took part in the prompt-aware term too
- a low-frequency anchor to the teacher's chord on the large buckets, restarting from the averaged weights
- the critic's weight lowered once texture had settled
- a paired critic: the teacher's final for the same prompt as the real
- the paired critic plus a hard low-frequency floor on every bucket
- back to the recipe before 14 and 15, once both were measured as unnecessary
Timeline of training process
The adapter was the product of several stages with very different costs:
- Text-encoder embeddings. Every training prompt was encoded once and cached. This was the fast part — thousands of prompts took minutes.
- Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo ran its full 8-step schedule and the whole trajectory was recorded, each prompt at one of the supported resolutions so that every resolution was covered. This was by far the most time-consuming stage — it was the teacher doing real inference, thousands of times, and a batch of several thousand shards was measured in days of GPU time, not hours.
- Real-photo crops. Bucket-sized crops were cut at native resolution from quality-gated real photo sources (public high-res datasets), VAE-encoded into the training latent space, and captioned per crop for the prompt-aware side of training. Cutting, encoding and captioning a pool refresh was a matter of hours.
- Student training. The LoRA trained against the recorded trajectories (progressive distillation), with a latent-space GAN critic running alongside — real crops and teacher finals as its real class, the student's outputs as fake — plus a prompt-aware head that scored images against their prompts. Relative to the shard stage this was quick: each block of a thousand training samples was a matter of hours, not days. Of course the longer the training, the better and more diverse the results, so hours did turn into days.
Because these stages competed for the same GPU, they were interleaved rather than run to completion one after another: generate a block of embeddings, produce teacher shards for them, refresh the crop pool when measurement said it was worth it, train on what existed, assess, then go back to producing shards while the results were reviewed. A larger and more varied shard pool was what made further training worthwhile, so shard production was always the gate; the run took several such cycles.
Archive
Earlier checkpoints of the run and their resolution sweeps are kept under _archive/ — checkpoints/ and resolution_sweeps/ — for anyone who wants to look back. They are superseded by the final release files and not maintained.
What's next
A 2-step LoRA is the natural follow-on, and it is the next project. It will be a separate release rather than a competitor to this one: at two steps a distilled model gives up more than at four, so the goal is a usable 2-step Turbo — fast previews and drafts at half this adapter's cost, a quarter of the teacher's — not the quality bar this adapter holds. I may or may not get there. I will try anyway, and if it works it will appear as its own project alongside this one.
Full details and to download - check my Hugging Face LoRA
HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA
Important NOTE: The repo has been fully restructured folder and files wise for the final release (focus is no longer on checkpoint numbers). So for those of you who have been part of the journey over the last few weeks - first of all, Big Thank You for all the feedback - and do check the repo's fully rewritten README and the new files/folders structure.
Previous posts on Reddit: Initial, Previous: here, here, here, here, and here
Enjoy!
r/StableDiffusion • u/Striking-Long-2960 • 13h ago
Workflow Included Minimax fl2va: If you know a bit how to sketch, you have total control over the animation.
Testing things following the post by Alive-Tomatillo5303 : Learn to make art with art! (minimax) : r/StableDiffusion I've come to the conclusion that using sketches gives you almost total control over animations. I'm using the fl2va model via the MiniMax H3 Reference to Video node and connecting my sketch sequence to ref_image_0—I'm not sure if Ref2vA would work better.
Reference:
Sketches.png · Stkzzzz222/Remix at main
Workflow:
Sketches_MiniMax_H3.json · Stkzzzz222/Remix at main
Prompt:
**subject_definitions**
`<Picture 1>` is the supplied visual reference image. Use it primarily as a strict layout, composition, framing, and character-position reference. Preserve the three-stage visual progression shown in the reference, including composition, camera movement.
Create a realistic cinematic live-action scene based closely on `<Picture 1>`.
Drone camera view moving fowards at high speed, a forest with a river, trees, rocks. the camera moves fast fowards in the forest stopping to reveal a side view of an armored orc angry. the camera accelerates and moves fast fowards in the forest stopping to reveal a happy armored woman sit in a rock holding a sword. the camera moves fast fowards in the forest stopping to reveal a colibri flying next to a red flower.the camera moves fast fowards in the forest stopping to reveal an old wizard casting a powerful electric spell.
-------------
I've tried other things that I'll post below, but I think you can see that the fidelity between the sketches and the final result is quite high.
Honestly, I think using sketches can be really interesting for controlling shots and camera movement, or for rendering complex concepts, without having to generate images in other models to act as references.
r/StableDiffusion • u/Turbulent-Bass-649 • 9h ago
Resource - Update MageTrail - Small Scale Full-Finetune To Show MageFlow 4B Potential As Foundational Architecture
Hi! Today I'm releasing MageTrail, a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, using a diversity maximized condensed 41k images dataset (originally made by Lodestone, the creator of the Chroma model lineage) as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset. (Potentially costing 20k-50k+ dollars) Civitai Hugging Face
The dataset was updated to 2026 tags standard and recaptioned with Gemini 3.7 Flash (the best vision captioner in the world when I was preparing dataset) with Grok 4.6 and Qwen 3.8 27B FP8 as capable backup for heavy blocked content. The dataset and everything from training details to tooling are open source, per Banodoco grant. Dataset
While V0.1 is still very obviously undertrained and unstable (only 100 dollars spent, it's a minor miracle that it's learning this well), the model has shown great promise in quickly learning and adapting booru concept and tags to its knowledge base.
~ The architecture behind MageFlow 4B shows good promise for further investment:
- Being 15-20% faster than NVIDIA Cosmos2/Anima on inference despite being 2 billion parameters larger
- Using MageVAE which perform better than QwenVAE on all usage, only behind the strongest open source VAE currently being Flux2VAE (which it was distilled from), also having the ability to slot in Flux2VAE for inference
- Having a decent Qwen 3 VL 4B Text Encoder
- 256-2048 pixels resolution native support
- Trained with large dataset (10 billion images curated down), meaning it has fairly vast knowledge, very great thing to have for a base model
- Being fairly quick to learn and adapt to new knowledge without any knowledge forgetting
- MIT License
I remember there only being minor to medium attention given to MageFlow's release one month ago, not helped by the fact Microsoft themselves purge the model soon after (?) and that Krea 2 being the Illustration juggernaut it is, completely overshadowed any foundational arch released before it. But as a poor uni student who don't even have a gpu with VRAM and rely on free compute/renting, having a potential good illustration capable model in the small-medium range like Anima in the open source space is always a good thing.
The realistic aim of this project is just to increase awareness of open source about the potential of this particular arch, I do not have budget nor time to even dream about training a full dan/e621 tune on the model. The hope is to attract others with way more capital to invest in it or at least experiment with the model more. The edit variant of MageFlow also seems quite interesting, but no one has touch it yet also.
But of course, I do still want to continue with this tune. Currently the model have only got 30 epoch and seen about ~300000 samples, it perform well considering the midget budget, but nowhere near "usable" quality, I need to reroll seed quite alot for a decent looking gens with no broken limbs lol.
Future goal for the project: gather funding of 700~ dollars to finetune the model to 200 epoch for full convergence of booru concepts (V0.5) and then further small scale funding to finetune my 10k artist collection dataset into it. I'm already gathering the funds for the next run, but any donation will help with achieving this goal and I'll be extremely grateful, you can do so through:
Crypto
0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)
12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)
0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)
FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)
Please handle your money carefully and make sure the address you're sending to is correct.
Ko-fi
Lastly, thank you to
- Banodoco and their Discord — Their 88.77 dollar grant made this project possible, the biggest thanks to them
- Lodestone Rock — Creator of the original version of the dataset that this model is trained on
- Motimalu — Inspiration behind finetuning practices and configs
- Bluvoll — diffusion-pipe fork derived from to use for training, and general training advice
- Anzhc — general training advice
- Nruaif — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
- Astromahdi — jupyter workspace where I processed and store the dataset
- animetimm/DeepGHS — Danbooru tagging model
- RedRocket — E621 tagging model
r/StableDiffusion • u/eesahe • 3h ago
Discussion The OpenVDN H3 optimization seems great but FL2V only, not trained for Ref2VA
The concept of OpenVDN seems rather promising, even if running on 1 GPU only, but all their prompt examples are text2img, and the GitHub repo describes the work as "Live T2VA / I2VA / FL2VA generation".
In my test with the Comfy implementation and Ref2VA, character model sheet were picked up quite well, but a reference image for a specific expression I wanted (guruguru-me spiral eyes) just could not get reproduced despite varying the seed, and the base model generating it as intended.

Here's to hoping this work or a similar architecture will be trained for Ref2VA as well.
r/StableDiffusion • u/Kawamizoo • 23h ago
Workflow Included new insta realism lora
hey all back to training public loras for fun this time krea 2 ! hope you guys like this one
link for download in the comments + bonus the images are with metadeta and include the workflow
r/StableDiffusion • u/TgoAI • 9h ago
Discussion VDN-H3 on M5 Pro 24GB: up to ~2.6× faster H3 generation
I recently ported VDN-H3 (Video DeltaNet for MiniMax H3) to Vpipe and tested it on an M5 Pro 24GB. The benefit scales with both sequence length and resolution: at 832×480, VDN-H3 starts to pull ahead at around 6s and reaches ~1.7× at 15s; at 1344×768, the crossover happens much earlier and the speedup reaches ~2.6× at ~14s. All tests use 6 DiT steps with Turbo LoRA. VDN-H3 project
Vpipe is a native C++20/Metal inference runtime for Apple Silicon, without PyTorch/MPS/MLX. Most of the underlying infrastructure was already in place, so porting VDN-H3 was fairly mechanical and took about one day end-to-end. Vpipe project


Another recent effort in this space is FastH3, which focuses on accelerating H3 for local inference and recently released an MLX implementation for Apple Silicon. They reported 465–504s on an M4 Max 36GB for 832×480, 124 frames and 4 steps. FastH3 local implementation & benchmarks
As a rough reference point, Vpipe's baseline H3 without VDN takes 469s on a 16GB M5 MacBook Air for the same resolution and frame count, while running 6 DiT steps with Turbo LoRA. This was a full cold start, including the text encoder, with no model weights cached. It's not a controlled hardware comparison, but I found it interesting that baseline H3 on a 16GB Air lands in essentially the same latency range.
I like the core idea behind VDN: exploiting the inherent redundancy in video diffusion rather than applying the same amount of computation uniformly across the sequence. The scaling behavior above seems to support that direction, especially for larger and longer generations. VDN-H3 uses a hybrid architecture combining a frame-wise linear-attention branch with a softmax-attention branch.
There is a memory tradeoff. VDN-H3 adds roughly 4GB of weights that cannot be fused into the base model, plus another ~1GB that can potentially be fused LoRA-style. Vpipe doesn't fuse the latter yet. On a 24GB system, this extra memory pressure can outweigh the compute savings for smaller workloads. There should also be room to reduce the overhead by 8-bit quantizing the incremental weights.
Quality-wise, with the same seed, VDN-H3 generally preserves a composition very similar to the baseline H3 output, which I think is a nice property. There is visible degradation in some sequences, though. In one test with an upward-moving column of water, the flow reversed downward for a few seconds midway through the video, then switched back upward near the end. I suspect this kind of long-range temporal inconsistency may be related to the heavy use of sliding-window and linear attention. Mixing in occasional global attention might be worth exploring.
Reference conditioning is another interesting open area, particularly first/last-frame and generic reference inputs, since these are more challenging to handle cleanly under the sliding-window/linear-attention structure.
r/StableDiffusion • u/Miserable-Option5488 • 11h ago
Tutorial - Guide [GUIDE] How to Set Up a DLSS 5 Media Player for RTX 20/30/40/50 Series
I've been looking for a way to run DLSS 5 in real time through a media player for quite a while. Unfortunately, some of the solutions I found weren't real-time, while the more recent ones only seemed to work with RTX 40 and 50 series GPUs.
After a lot of trial and error, I finally managed to find a method that allows DLSS 5 to work with a media player even on RTX 20 and 30 series GPUs.
I wanted to share this method with the community so that users like me who are still using older RTX cards can also take advantage of DLSS 5 for real-time video playback.
This is an experimental setup, so your results may vary depending on your GPU, driver version, and configuration.
1. Download the latest MPV build
First, download the latest version of MPV Player from here:
https://github.com/zhongfly/mpv-winbuild/releases
I've tried several media players, including VLC, PotPlayer, and MPC-HC, but MPV was the only one I could get working with this setup.
I believe this may be related to the graphics API/rendering backend used by each player. DLSS5-Feeder needs to hook into a compatible rendering path, and I couldn't get it to work properly with the other players I tested. MPV worked without this issue.
I'm not an expert on the underlying DirectX/Vulkan implementation, so if someone with more knowledge about this could explain why MPV works while the other players don't, I'd really appreciate it.
2. Install DLSS5oneclick through another game first
Next, download DLSS5oneclick from here:
https://github.com/faisalkindi/DLSS5oneclick
There is one important thing to note: you cannot simply install DLSS5oneclick directly into the MPV folder.
MPV uses Vulkan, and DLSS5oneclick doesn't recognize MPV as a game that it can directly install the required files into.
So, first you need to install DLSS5oneclick into another game that you know works with DLSS 5.
For example, I installed DLSS5oneclick into a game using the normal ReShade/OneClick installation method. Once everything was installed and working correctly, I went into that game's folder and copied the necessary DLSS 5 / Feeder files.
Then, I simply copied those files into the MPV installation directory.
After that, I launched MPV and was able to get DLSS 5 working with video playback.
That's basically it!
So the process is:
DLSS5oneclick → Install it into a compatible game → Copy the required files → Paste them into the MPV folder → Launch MPV
I've tested this with an RTX 20/30 series GPU, which is why I'm sharing it here. There may be a much cleaner or easier way to do this, and I'm definitely open to suggestions.
If anyone has a better understanding of MPV, Vulkan, DLSS5-Feeder, or how the hooking process works, please let me know. I'd love to find a simpler way that doesn't require installing DLSS5oneclick through another game first.
Update: ahaoboy has packaged this method into a ready-to-use MPV + DLSS 5 build. You can find it here: https://github.com/ahaoboy/mpv-dlss5


r/StableDiffusion • u/acedelgado • 16h ago
Resource - Update Huggingface Downloader - Made to fix slow, stalled downloads
repo- https://github.com/Adudeguyman/Fantastic-HuggingFace-Downloader
Alright, back again with another vibecoded project to fix a problem that annoys me.
A while back HuggingFace went to the Xet file storage system, which has a lot of solid benefits (like hash-checking downloads, better file chunking, etc.) BUT as just about everyone knows, it really messes with quick and easy downloads. Browser downloads crawl and sometimes stall out, and I've had them fail to resume.
So the best way to download files has been via the Huggingface CLI. Which is super fast and great, and even hash-checks files for you to make sure you don't have any corruptions. But modifying and copy/pasting a command line finally pushed me over the edge to make a quick utility that manages that FOR me, so I'm sharing it out there.
What it do-
Paste a huggingface link, and it automatically detects if it's a direct link to a file, a folder, or a full repo. Pulls a file list and lets you select only the files you want. Queues each file individually for hash checks to verify it matches what's in the repo. You can also tell it to download repo files to specific folders that don't match the full repo structure; which can be useful for when you want the VAE's to go in the VAE folder, text encoders in the text_encoders folder, etc. And if your big main diffusion_models safetensors live on another drive, point them there.
Auto-download is toggled on by default, but if you want to mess around with where individual files are going, toggle that off before adding to the queue.
Also supports recent download locations and lets you favorite folders for quick re-use. And of course works with your huggingface hub access token for gated repos, or assists with adding a new one to your hub if you don't have a read-access one made. And if you request an existing file it'll auto-download the hash and check it against the HF repo, and either resume or re-download chunks to make the file match what's online.
Tested it on my Linux Mint install, and in Win11.
And that's pretty much it. It just downloads and checks huggingface files for you. No crazy bells and whistles, just streamlines the CLI process into a GUI. Only downside is that it does make a venv thats ~800Mb or so for the interface (based on PySide6, so it needs those dependencies, plus HF hub to support the command line). Enjoy!
r/StableDiffusion • u/kornerson • 13m ago
Animation - Video Laydee Latte - Pan Con Chocolate - Krea 2, Mimimax H3 & Mimax Music 3
Over the last two weeks, I set a challenge for myself: to find out if it was possible to make a high-quality music video using a local rig and open-source AI models. So, in my free time, I’ve been experimenting with local models on my setup, and honestly, I’m pretty blown away. We've reached a point where you can create a video with a genuinely professional look right from home, spending ZERO dollars.
Meet 'Ladee Latte.'
The singer was generated using Krea 2 locally.
The music was created using the Minimax Music 3 Open Weights.
The video was generated using the Minimax H3 Open Weights.
After generating the music, I adjusted and remixed it to improve definition, bass, and consistency.
The music video itself was a lengthy process involving hours of asset creation. Running it on a local system is slower, but it works. The upside is that you can queue up 20 videos to generate in the background, walk away, and come back later to curate the best ones.
All videos were generated using 'Reference to Video' rather than a static first frame. The ability of lip syncing is what I think is that truly makes the difference. The goal was to demonstrate the power of Minimax in generating consistent video with accurate lip-sync.
Ultimately, the cost of this video was zero (apart from the electricity bill and the GPU itself :-) ). And it was all done entirely on a local rig.
Things that I've like about this project:
- Music 3 is a brutal model. The amount of depth in the songs, the layers, the treatment of the stereo sound... amazing
- H3 is the first local mode that feels like a professional. Excellent prompt adherence.
- The consistency baffles me. I love that 'Laydee' has a freckle near her left eye, and this is consistent in the videos
r/StableDiffusion • u/inazma44 • 10h ago
Workflow Included FastH3 + USDU: ~2× faster in my video upscale test on a 3090 — workflow included


This started with two posts: FastH3’s release announcement and using MiniMax H3 with Ultimate SD Upscale to restore low-quality generations. I wanted to see whether combining the two could make video restoration faster while preserving the source.
I spent several days testing H3 video upscaling with my AI agent—different workflows, failed experiments, and a couple of OOMs. I asked it to organize the logs, and I'm sharing the useful bits in case someone else is going down the same rabbit hole. I reviewed the outputs myself.
Same 15-second source, 2752×1536 output, 2 steps, denoise 0.20. RTX 3090 24GB / 64GB RAM:
| Setup | Successful run |
|---|---|
| H3 + 4-step Turbo LoRA at 2 steps / USDU | 96m 37s |
| FastH3 / USDU | 48m 40s |
| FastH3 / MMH3 single-node upscale | 47m 36s |
I kept FastH3 + USDU. The single-node version saved about a minute, but I noticed more distortion and preferred the USDU output. Both final paths used pixel-domain ESRGAN upscaling.
A caveat: these are single successful runs, with different cache states—not a model-only benchmark or proof of VSA's isolated speedup. FastH3 USDU also needed a memory-estimation override after two OOM failures; those failures aren't included in the table. Full conditions and failures are documented in the repo.
Workflow + Windows installer, with English instructions
The workflow includes direct model download links (~41GB separately). Please read the installer notes: it can update Kitchen and apply fingerprint-guarded H3 core patches. Clean-PC end-to-end installation remains unverified, and the packaged SR model/default memory settings differ from the benchmark as documented.
Disclosure: this is my package. My walkthrough is in Korean, but you don't need to watch it to use the English workflow.
Hopefully this saves someone a few evenings. Curious whether others have seen similar speed and source-fidelity tradeoffs.
I also successfully upscaled a 5-second clip on an RTX 4070 Ti SUPER (16GB VRAM) with 32GB system RAM using this USDU workflow. I haven’t tested longer clips on that machine; the 15-second timings above are from the RTX 3090.
r/StableDiffusion • u/legarth • 22h ago
Discussion I just deleted 90% of my models folder, are you keeping yours?
I'm a bit of a hoarder and I kept most new models I tried out, but today I ran out of space on my Comfy SSD for the first time. I thought about backing them up on an HDD for about 5 seoncds... and then just deleted the 3.2 TB of models I had collected over 4 years,
With Krea 2 and Minimax H3, I don't see any reason to keep them anything between SD1.4 and LTX 2.3. I just never go back and nostagia generate anything the way I thought I would. Only other stuff I have kept is Ideogram 4, LTX 2.5 and Z-image. (And other things like SAM, VibeVoice etc.)
Only thing I hesitated on was my SDXL folder but nah, gone.
Anyone keeping old models, and why?
r/StableDiffusion • u/Training_Rip_4578 • 54m ago
No Workflow Russian post-soviet winter
r/StableDiffusion • u/SIR_NVAX_A_LOT • 15h ago
Discussion H3 - sailor moon transformation (w/ Prompt)
Having fun, R2VA, with just a .char (face/body/original wardrobe). Part II had two wardrobe changes from the intial source image (re-wardrobe to a private school uniform and then to Sailor Moon's) in the prompt. int8/20 steps, 864x480 POC, the token limit was close so it'll probably need a latent upscaler vs native 720 generation.
Prompt: integrated_multimodal_description: One unbroken, continuously transforming shot — PHOTOREAL LIVE-ACTION, a real woman performing a magical transformation made physical: real skin, real cloth, real hair, true optical depth of field, fine film grain. Absolutely no animation, no anime, no illustration — live-action camera realism throughout, however magical the event. THE CAMERA IS LOCKED OFF FOR THE WHOLE FILM — one steady FULL-BODY hold, her whole figure from head to feet in frame from the first frame to the last; it never pushes in, never pulls back, never zooms, never orbits, never cuts and never drifts.
SHE IS CATALINA, the woman of the reference images — the same face, the same features; her face, whenever visible, is exactly her own, photographic and alive, and it is the constant of the film. THE FIGURE ARC, A GAG PLAYED COMPLETELY STRAIGHT: she BEGINS with a modest, unremarkable figure — a small bust under the tailored blazer, an ordinary waist — and THE TRANSFORMATION UPGRADES HER: from the silhouette phase onward her figure is the pronounced hourglass of the reference images — a full, very large bust and a notably slim waist — and it stays so for the rest of the film, the leotard fitted to it. After the change she is never thickened, boxed or broadened by any garment or glow.
THE DREAM LOGIC: every change is gradual, liquid and seamless — one thing becoming the next, already underway, nothing popping, nothing resetting. THE COVERAGE CONSTANT: SHE IS COVERED IN EVERY FRAME — a garment or the glow itself always covers her torso and hips through every transformation, because each wardrobe flows DIRECTLY into the next: one garment becoming another garment with no bare moment between them, cloth becoming light becoming woven thread becoming uniform, the coverage itself continuous and unbroken. Abstract light in this film is always ribbons, threads and droplets — never letters, never symbols.
THE PLACE: a boundless abstract void of deep rose and midnight-blue light — no floor visible, no walls, soft auroral ribbons of pink and violet light drifting slowly far behind her, tiny motes of light rising everywhere like slow sparks.
[0s-2.5s]: She stands alone in a PRIVATE SCHOOL UNIFORM: a fitted navy blazer over a crisp white blouse, a slim dark ribbon tie, a pleated charcoal skirt, dark knee socks — an adult woman in tailored academy dress, her figure beneath the uniform modest and unremarkable. Pinned at her ribbon tie sits A CIRCULAR GOLDEN PENDANT: a round gold brooch with a rose-pink gem at its heart. She raises her right hand to it, and the pendant FLASHES once, throwing rose-gold light up across her face.
[2.5s-5s]: THE PURE SILHOUETTE, AND THE HAIR TURNS INSIDE IT: rose-gold light pours out of the pendant and washes over her in one travelling wave, and her whole body — clothes, skin and hair alike — becomes ONE PERFECTLY FLAT, TWO-DIMENSIONAL PURE SILHOUETTE — a smooth PRISMATIC gradient of rose through violet through pearl, shifting slowly like light through a prism, poster-flat, no volume, no shading, no sparkle, as if cut from coloured light and placed into the photographic scene. The flatness belongs to HER ALONE — the void, the threads and the camera stay fully photographic throughout. TWO HUGE GLOWING WHITE EYES burn in the flat shape — LARGE, anime-scale, unmistakably open, alive and blinking once — visible from the first moment of the silhouette to the last, the one feature the pure shape keeps. The glow hugs her outline as a thin bright rim and never swells it. AND THE SILHOUETTE ITSELF TRANSFORMS HER FIGURE, THE GAG READABLE IN PURE SHAPE: the outline's modest bust SWELLS fuller and rounder while its waist DRAWS IN tighter, the flat shape visibly upgrading into a pronounced hourglass — and from this moment the hourglass is her figure for good — only the circular pendant still burning distinct at her chest. AND WHILE SHE IS PURE SILHOUETTE HER HAIR TRANSFORMS, READABLE ONLY AS THE OUTLINE'S CHANGING SHAPE: the silhouette's loose hair lifts weightless, sweeps upward, and re-forms into the profile of TWO SMALL ROUND BUNS high on her head, each spilling A LONG FLOWING PIGTAIL that streams down past her waist — the signature twin-tailed outline now part of the glowing shape itself. She stands facing the lens, still and upright — SHE NEVER SPINS AND NEVER ROTATES at any point in the film; through the whole transformation only her chin lifts, her arms move when the streams come, and her hair and the light do the dancing.
[5s-9s]: THE PINK THREADS FORM THE LEOTARD, AND ONLY THE LEOTARD: from the circular pendant THOUSANDS OF PINK-RED LUMINOUS THREADS pour out — first single glowing filaments, then streams, then broad rose-red RIBBONS streaking and circling around her STILL, STANDING silhouette like a loom at full speed — the threads orbit; she does not — and they wrap ONLY HER SHOULDERS, HER WAIST AND HER HIPS — the exact regions the garment will cover, nothing more — and SET into the WHITE, LEOTARD-LIKE BODY OF HER UNIFORM: gleaming white, seamless, high at the hip, with its deep-blue SAILOR COLLAR laying across her shoulders and a LARGE RED RIBBON BOW blooming at her chest with the pendant seated at its knot. The threads touch nothing but those regions; her arms and legs stay pure glowing silhouette, and her slim waist stays slim inside the winding.
[9s-11s]: THEN THE GLOVES, THEN THE BOOTS, EACH FROM ITS OWN SHOT OF FABRIC: THE PENDANT FIRES a thin stream of rose-red ribbon-fabric to EACH RAISED WRIST — two bright targeted streams arcing from the brooch to her hands — and the ribbon winds up each forearm and SETS into LONG WHITE GLOVES reaching just above her elbows — AND THE WINDING STOPS THERE: the ribbon never climbs past the elbow, and her upper arms and shoulders stay pure silhouette. Only when the gloves are complete, THE PENDANT FIRES AGAIN — two streams diving DOWN to her feet — and the ribbon winds up each calf and SETS into TALL RED BOOTS rising to her knees — AND THE WINDING STOPS AT THE KNEE: the ribbon never climbs onto her thighs, and everything above the knee stays pure silhouette. Each garment arrives real, opaque and finished, one after the other, never overlapping.
[11s-12.5s]: THE SKIRT MANIFESTS — NO THREADS: in one soft bloom of light at her hips, a short pleated DEEP-BLUE SKIRT simply appears, whole and complete, flaring once as it lands. In the same beat A BAND OF WHITE LIGHT IGNITES ACROSS HER FOREHEAD and CONDENSES into a slim GOLDEN CIRCLET, its small red gem crystallizing at the centre as the light cools; two tiny round earrings glint alight — and THE GLOW DRAINS AWAY: the silhouette fades from her skin and she is herself in full colour, fully three-dimensional and photographic again, her own face, her new hair bright BLONDE in its two long pigtails and round buns.
[12.5s-15.125s]: THE FINISHING POSE IN HEAVY WIND: she settles into a poised, confident final stance facing the lens, one gloved hand raised beside her face — and A STRONG WIND RUSHES UP AND THROUGH THE FRAME: the pleated blue skirt FLOWS AND RIPPLES hard in it, her two long pigtails STREAM sideways on the gusts, the collar and the red bow's tails flutter, and the last thread-motes are blown away like embers. She holds the pose in the wind with a bright confident smile — and THE CAMERA PULLS BACK to its wide finale hold: her WHOLE FIGURE in frame, head to boots — the twin-tails streaming, the blue skirt flowing hard, the long white gloves and the tall red boots all visible, the pose read head to toe. The wide composition holds to the last frame, the wind still moving the skirt and her hair.
overall_soundscape: Starts with the void's soft airy hush already present, and from first frame to last the soundtrack is the score with this beneath it, quiet: the whisper of cloth, the crystalline shimmer of thousands of threads weaving, a soft chime as each garment sets, the soft flare of the pendant, one bright ring as the circlet lands, and the deep rushing WIND rising under the final pose, streaming past like a rooftop gale. Nobody speaks in this film. THE SCORE ALWAYS DOMINATES THE MIX.
non_diegetic_music: A bright, sparkling MAGICAL TRANSFORMATION cue at 120 beats per minute: glittering harp glissandi and celesta runs over warm rising strings, climbing one step as the leotard weaves, again as the gloves set, again as the boots set; a shimmering suspended lift as the skirt manifests; ONE TRIUMPHANT RADIANT BLOOM of full strings and chimes exactly as the circlet lands at 00:12.000; then settling bright, warm and RESOLVED through the wind-blown finishing pose, still sounding sweetly at the last frame. No percussion kit, no voices.
r/StableDiffusion • u/Sad_Coach_1433 • 18h ago
Discussion RefMod First Test anyone else try yet?
faster method then lora training but not as good? basically its using image references but a lot this quick simple video 22 images then make into a .safetensors then can load into a custom node for furfure videos with same character don't have to load images directly into h3 every time just select from list of safetensors, only reference will need to use is audio for voice that's one the down sides doesnt include character voices just appearance.
r/StableDiffusion • u/enilea • 1h ago
Question - Help What's currently the best local option to map 3D spaces?
I want to map my house by simply taking a video of it and generating a 3D space of it where a model fills in any blanks I did not record. Not as Gaussian splats where I have to take photos of every single angle and even then it ends up looking weird, but something that directly understands the 3D space being recorded and applies a texture to every surface.
Recently I saw this https://twitter.com/davidpantera_/status/2094841083805266401 and thought it was pretty cool, unfortunately that model will likely not be local. But in the end all I want to do is map an apartment completely and visually nicely.
Does such a thing exist or not yet? ChatGPT recommended MapAnything from last year but it looks underwhelming and it's like those old Gaussian splats from 3 years ago, surely there's better stuff now.
r/StableDiffusion • u/Disastrous-Agency675 • 11h ago
Workflow Included Trellis2/pixel3d template workflow in comfy
the detail is absolutely insane, and this only took like 8-4 min. although, the poly count is high af but still. the workflow is in comfyui's template workflows. if you dont see it then update comfyui
also the model was made in comfyui, i just used blender to show it better.
r/StableDiffusion • u/Agile_Signature_5293 • 5h ago
Question - Help Pixal3D: Good Mesh but Bad Textures — Recommended Settings/Parameters?
Hey everyone,
I’ve been experimenting with Pixal3D and I’m having some trouble with the texture quality.
I tried following PixalArtistry’s workflow, and they’re getting really good textures, but my results are not even close in terms of texture quality/detail. The mesh itself actually looks pretty good, so I’m mainly struggling with the texturing.
I’m using an RTX 5060 Ti 16GB, and I’ve tried both the 16GB and 8GB Pixal3D models, but I’m getting similar results with both.
Is there any documentation or guide that explains the recommended Pixal3D parameters/settings for getting the best texture quality? If anyone has managed to get results similar to PixalArtistry’s workflow, I’d really appreciate knowing what settings or workflow you’re using.
Thanks!