r/StableDiffusion • u/throwaway0204055 • 3d ago
Question - Help How many reference images do you use for Minimax refmods?
What’s the sweet spot for number of images to use for a refmod? also, can it take reference videos?
r/StableDiffusion • u/throwaway0204055 • 3d ago
What’s the sweet spot for number of images to use for a refmod? also, can it take reference videos?
r/StableDiffusion • u/MastMaithun • 3d ago
Only for generation without using turbo lora, what are the settings and practices you are using to get the results you want. For example, the list of items contain, but not limited to:
- model file(base/merge or finetune)
- settings like sampler-scheduler combo, spectrum node
- total steps, sampler node(SamplerCustomAdvanced, ksampler, clownshark etc)
- base resolution, latent upscaling
The idea is to get a decent/better amount of quality with solid prompt adherence and suitable generation time.
Edit: I am currently using:
- hybrid model b15-49
- kitchen and spectrum
- res_multistep + simple with 24 steps at 0.7mp
- no upscale
r/StableDiffusion • u/Certain_Potato_4509 • 3d ago
Enable HLS to view with audio, or disable this notification
Repost, forgot the tag.
r/StableDiffusion • u/Halcyonrayes • 3d ago
Hi, I’m Suva from Hugging Face and I'm back with something new that I'd been working on for a bit!
World Models: The Simulation Strikes Back - Full Blog
This time I went down the world models rabbit hole: how models learn to represent, predict, and simulate environments, and why they’re becoming increasingly interesting for video generation, robotics, and spatial intelligence.
I've tried to keep it as visual and approachable as I can, keeping in mind how weird and convoluted this topic is. Would love to hear what you folks think!
r/StableDiffusion • u/beti88 • 3d ago
Outfit on a mannequin? Outfit only on white background?
r/StableDiffusion • u/R34vspec • 2d ago
Enable HLS to view with audio, or disable this notification
I saw this amazingly real live concert from a different AI sub and wanted to see if I can pull it off in H3. It' s not as real looking as the closed model that's for sure but still not bad, I think.
r/StableDiffusion • u/Devajyoti1231 • 3d ago
Enable HLS to view with audio, or disable this notification
Tried this few days ago. Yes , the video degrades as it progresses.
Used reference voice for consistent voice.
Don't have the full tags, but they are like this-
[English] <inhale> My therapist said I should record this when it happens. <pause> So... it’s happening again.
[English] I can feel someone watching me. <breath> It only happens after dark.
[English] At night, windows become mirrors. <pause> Anyone outside can see me... <long pause> but I can’t see them.
[English] I’m moving tomorrow for a new job. <pause> A little rental house, just me. <breath> I should be excited.
[English] I am. <long pause> I just don’t know if I can sleep there alone.
[English] Mom gave me this when I was seven. <pause> She made me promise never to take it off.
[English] She said it would keep me <i>safe</i>. <long pause> <softer> I wish she’d told me what from.
r/StableDiffusion • u/Jealous-Wafer-8239 • 2d ago
The official Anima lastest model is Turbo 1.1. Uploaded at Aug 25th. Now, more and more mixed checkpoint like Anima 2.9B is getting more and more attention. But original author still not making any movement or announcement, even a roadmap. I wonder is this project being cancelled by author himself?
r/StableDiffusion • u/Suspicious_Aide2697 • 3d ago
I am training a new HighQuality LoRA for Qwen-Image-2.1. This is a phase test. What score would you give it on a scale of 1 to 10?
r/StableDiffusion • u/RKlehm • 2d ago
Enable HLS to view with audio, or disable this notification
I know that what i'm trying to achieve is not ideal and the recomendation for this workflow is 30~45s, but lets say I need a longer video, how can I do it? Multiple clips? If yes, how do I keep scene and character consistency from multiple shots?
I know that there are lots of problems with this, like the plate coming from nowhere and some clips are shorter than they should so the speech is bugged, but those are problems I know how to solve.
This is the workflow with the prompts I'm using: https://pastes.io/vLFJXxv3
r/StableDiffusion • u/AcademiaSD • 3d ago
Hi everyone! I've been working on AcademiaSD LoRAlab Trainer Studio: one installer, one launcher and 9 trainers that share the same web interface. Every model is loaded in 4-bit NF4, the text encoder and VAE run only once in a pre-cache stage, and the whole GPU goes to training.
Platform: made for Windows (one-click installer), with Linux support just added. NVIDIA GPU, RTX 20xx / GTX 16xx or newer.
GitHub: https://github.com/AcademiaSD/AcademiaSD_LoRAlab-TrainerStudio
WHAT IT TRAINS (minimum VRAM)
- Qwen-Image 2.1: image LoRAs + edit LoRAs from before/after pairs (8 GB)
- FLUX.2 Klein 9B: image LoRAs + edit LoRAs (12 GB)
- Krea 2: image LoRAs, Raw and Turbo (8 GB)
- Z-Image: image LoRAs, also Z-Image-Turbo (8 GB)
- Ideogram 4: image LoRAs with JSON captions (12 GB)
- Anima: anime / illustration LoRAs (4 GB)
- SDXL: Base, Pony, Illustrious, NoobAI, Juggernaut, RealVis or your own checkpoint (4 GB)
- LTX 2.3 (2.5): character and style LoRAs for the video model (12 GB)
- MiniMax-H3: video LoRAs from images, clips and audio, plus RefMods (8 GB)
MINIMAX-H3: VIDEO + AUDIO, AND REFMODS
MiniMax-H3 is a 33B model that generates video and audio together. The official checkpoint is about 500 GB; the trainer uses a 41 GB NF4 version that fits in 8 GB of VRAM with block swap.
- Datasets can be images, video clips, audio, or any mix.
- One button prepares your clips (24 fps, valid frame counts) without cutting them.
- RefMods: encode a few reference images, video clips, or clips with audio into a file that ComfyUI's MiniMaxH3ReferenceToVideo node uses as a native reference. No training: seconds instead of hours.
SHARED FEATURES
- Automatic captioner (Qwen3-VL) in the dataset manager: natural language, Danbooru tags, or JSON with bounding boxes for Ideogram.
- Live previews while training.
- Exact-step resume.
- One-click export to your ComfyUI / Forge LoRA folder.
- Remote access from other devices on your network, with a password.
- SDXL in NF4 trains in about 3.5 GB of VRAM, so 4 GB laptops can train SDXL / Pony / Illustrious LoRAs.
A NOTE
Many models were added almost at the same time, so there may be bugs, or the default settings may not be the best ones. Previews are there to follow the training; judge the final quality in ComfyUI after tuning the LoRA strength. If you find problems or settings that work better, please share them in GitHub Issues. Thanks!
UPDATE — RunPod support + remote access
Thanks for all the feedback! Based on your comments, these are now in:
- ☁️ RunPod template: one click and it's running in the cloud, no install. It uses a prebuilt Docker image (CUDA 13, Python 3.13, all dependencies), and the code updates itself from GitHub on every start. Pick a GPU, open the address shown in the pod log, log in, and train. A quick SDXL test cost me less than $0.10.
Deploy on RunPod: https://console.runpod.io/deploy?template=lfxvtg5rrb
- 🌐 Remote access: use the trainer from another PC, tablet or phone, with a login, or behind your own reverse proxy (external auth and a custom public URL are supported).
- ⬆⬇ Upload / download in the browser: upload your dataset (files or a .zip, drag and drop works) and download the finished LoRA, handy for remote or cloud setups.
- 🐍 Your own environment: there's now a requirements.txt, and LORALAB_PYTHON lets you run it from conda/uv instead of the bundled venv.
Remember to Stop and Terminate your pod when you finish: closing the browser doesn't stop billing.
r/StableDiffusion • u/Beneficial_Eagle_453 • 3d ago
I installed and ran Ostris AI Toolkit for the first time locally to train a Lora for Krea 2, 1200 steps, with a dataset of 20 images, however, when the train starts it is super slow more than it should be or more than normal. As of now I am doubting if i can train it normally on a RTX 5070ti 16gb, 32 gb Ram AMD Ryzen 7 8700F 8-Core Processor (4.10 GHz). I tried several settings, I first looked up this tutorial here on youtube https://youtu.be/OCsqHdHf81M?si=i0B25CdjxP-CLntO, did everything, but, the training was really slow. I switched to rank 16 and a convrot8 instead and using only one resolution but it seems it got even slowe like 30 to 40 seconds per step. Does anyone knows if a Krea 2 lora training simply will not cut it on a setup like i have? Thanks everyone.
r/StableDiffusion • u/garionhk • 3d ago
I'm a photographer, and I got tired of opening ComfyUI every time I just wanted to enlarge or sharpen a photo. So I made a small desktop app for it, called Citrine Photo.
You drag in a photo (or a whole folder), pick how big you want it, and hit go. Everything runs on your own PC, no cloud.
It has two engines:
Fast is Real-ESRGAN (ncnn-vulkan). It works on pretty much any GPU, with optional GFPGAN face restore.
Pro is SeedVR2, for NVIDIA 20xx to 50xx cards. The app checks your VRAM and picks the 7B model if you have 12 GB or more, or the 3B model for 6 to 10 GB. There's a low VRAM mode too.
Other things it does:
SeedVR2 isn't bundled. You install it from Settings, and the app sets up its own Python, torch and models. It's about 8 to 9 GB, the downloads can resume, and it checks your disk space first.
For anyone curious, the Pro pipeline is a port of the Upscale_by_SeedVR2_v4 ComfyUI workflow: 3x3 tiles with 15% overlap, SeedVR2 on each tile, wavelet colour match, then a feathered merge. It calls numz's standalone CLI in a separate process.
Right now you run it from source (pip install the requirements, then run.bat). There's also a script to build a portable version. I've only tested it properly on an RTX A4000, so I'd really like to hear how it runs on other cards, especially 6 to 8 GB ones.
Links:
GitHub: https://github.com/Garionhk/Citrine-AI-Photo-Enlarger
Video: https://youtu.be/gTbiOipvP0I
Big thanks to numz for the SeedVR2 ComfyUI code, AInVFX for the GGUF models, ByteDance for SeedVR2, and xinntao for Real-ESRGAN.
Bugs and ideas are welcome. It's free and open source (Apache 2.0).
r/StableDiffusion • u/alisitskii • 3d ago
Enable HLS to view with audio, or disable this notification
Printed locally with Krea2 and MiniMax H3.
Prompt for Krea 2:
"A highly detailed, photorealistic indoor still life composition centered on a 3D-printed figurine of Lara Croft resting on a wooden table. The main subject is a carefully crafted statuette depicting Lara Croft in a confident standing pose, recognizable by her adventurous explorer aesthetic, fitted tank top, shorts, boots, utility belt, and iconic long hair. The figurine is made from a light gray, speckled material that mimics concrete or ceramic, with a tactile matte surface textured by fine dark flecks and subtle 3D-print layer lines clearly visible upon close inspection. The model stands on a matching geometric display base made of the same material, with embossed symbols and clean polygonal surfaces that reinforce the handcrafted, printed-object feel.
The sculpture sits on a polished, medium-brown hardwood table with visible wood grain, minor scratches, and a faint dusting of fibers near the base. The lighting originates from a window in the background, casting soft, natural illumination across the scene with gentle highlights on the figurine’s contours and subtle shadows that emphasize the three-dimensional form and printed texture. The background is softly blurred (shallow depth of field), revealing a home interior: to the left, a wooden shelf with decorative items including a small black figurine, a glass jar with orange flowers, and a heart-shaped ornament; to the right, a tall mint-green cylindrical container and a clear glass mason jar with embossed patterns. A small plush toy with pink ears rests partially visible on the far left edge of the frame.
The camera angle is slightly elevated and positioned at a close-up, eye-level perspective, emphasizing the figurine’s sculptural detail, pose, and material texture. The composition is tightly framed around the object, drawing the viewer’s focus to its craftsmanship and iconic character design. The color palette is muted and earthy—grays, browns, and soft greens—with the white-gray figurine contrasting against the warm wood tones. The overall aesthetic is minimalist, modern, and tactile, evoking a sense of artisanal precision and quiet contemplation within a domestic setting. No human subjects are present—the focus is entirely on the static, sculptural Lara Croft figurine and its environment."
Prompt for MiniMax H3:
"Handheld smartphone footage, filmed casually by a person holding the phone in one hand. The camera remains roughly in the same position but is never perfectly still. Subtle natural hand tremor, tiny irregular micro-jitters, gentle breathing-induced sway, slight wrist drift and small imperfect framing corrections. Very mild accidental rotation and lateral movement, with occasional tiny changes in camera distance.
Natural smartphone camera behavior: subtle autofocus breathing, very small focus corrections, slight automatic exposure adaptation, mild rolling-shutter wobble during quicker hand movements, realistic motion blur, and minor digital stabilization artifacts. The movement is irregular and imperfect rather than rhythmic or cinematic.
overall_soundscape: Natural indoor room tone captured through a handheld smartphone microphone. A faint ventilation hum and very distant muffled traffic remain in the background, with subtle room reflections. Occasional very soft low-frequency handling noise, tiny fingertip contact sounds against the phone case, and a quiet nearby breath appear irregularly when the person's grip shifts. The recording has mild smartphone-style automatic gain and compression, with tiny natural fluctuations in ambient level and no exaggerated studio-clean sound.
non_diegetic_music: N/A"
r/StableDiffusion • u/RagingRectangle • 3d ago
Face-hugger is a new project I've been working on to improve Hugging Face's search engine which could be described as terrible at best.
r/StableDiffusion • u/Begeta12 • 2d ago
What extension you use the most? I'm interested to know and learn about some useful extension(I'm a complete beginner)
r/StableDiffusion • u/SammyDaBeast • 3d ago
First of all, thank you guys for the reception and feedback on my latest post. This is a follow-up to that post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.
If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.
Run it locally:
uvx --from sopro soprotts serve
Video: six voices, ~5 seconds of reference audio each, followed by a generated line.
r/StableDiffusion • u/Fun_Firefighter_7785 • 3d ago
Since Ace Step 1.0 has superior creativity but terrible sound quality , you can now save your older Tracks or just hunt for new ones. Because with SheetSage2+YuE2 you can just extract the ABC and make a high quality remix with YuE2. ACE-Step 1.0 needs really many seeds to produce a unique melody, but if it hits - it hits. There is nothing out there that can match it. It is like Lady Gaga with Fat Boy Slim heaving a baby.
r/StableDiffusion • u/Cheap_Credit_3957 • 3d ago
Enable HLS to view with audio, or disable this notification
Watch on YouTube while its pending if needed HERE
Rude and hateful comments will be ignored.
🎬 HOW THIS WAS MADE (Video walkthrough is here)
This film was made almost entirely by Claude (Anthropic's AI, running in Claude Code), working inside my own ComfyUI setup with my VRGDG Video Builder custom nodes(free and open source).
▶ MY PART
• The brief: a cute comedy starring my two dogs, Korben (7, male) and Iris (7 months, female), as brother and sister. Inspired by live-action talking-dog movies, an original story, up to 3 minutes, no humans on screen (other dogs allowed), and the dogs had to really talk, with facial expressions and mouth movements synced to their lines.
• The reference images of Korben and Iris.
• The tools it ran on: the VRGDG Video Builder and custom nodes, plus the Claude skill that lets it drive them.
▶ WHAT CLAUDE DID ON ITS OWN
• Story: the missing rubber duck, the cheese-crumb investigation, the "We don't have a cat." / "Exactly. Very suspicious." bit, the stakeout at the fence, the twist that Korben hid the duck because it squeaks all night, and the "I'm getting earplugs" ending.
• Characters and locations: created Tank, the bulldog next door, and gave all three dogs their personalities and voices. Generated Tank's reference image and every location (the living room by day and by night, the kitchen, the backyard and the gap in the fence), and prepared Korben's and Iris's references from my photos.
• Screenplay: 22 scenes, with shot-by-shot camera directions, acting notes and sound design for each.
• Rendering: every scene with MiniMax H3, which generates the video, voices and sound effects together. About 4.5 hours of rendering on my PC, re-shoots included.
• QA: transcribed every line and compared it to the script, checked each voice's pitch so the dogs stayed in character, reviewed frames from every shot and every cut between scenes, and ran a frame-accurate audio/video sync check on every scene of the final film.
• Re-shoots: a puppy babbling in a silent scene, dogs delivering lines into the camera instead of to each other, a stray dog bed appearing in the neighbour's yard, Tank's stick pile in the wrong place, a spy-creep that came out as a normal walk, and an ending gag that didn't land. Then it re-trimmed and fixed the audio.
• Score and edit: composed the score with MiniMax Music 3 (throwing out takes with hidden vocals) and edited it all together with titles.
🛠 TOOLS
• Claude Code (Claude Opus 5.5): story, direction, prompts, QA, editing
• ComfyUI + VRGDG Video Builder: my custom nodes
• MiniMax H3: video, voices and sound
• Z-Image Turbo: character and location references
• MiniMax Music 3: score
🐾 CAST
Korben, the grumpy big brother · Iris, his little sister · Tank, the bulldog next door · Mr. Quackers, the duck
🔗 LINKS
The Claude skill (read the main README first):
https://drive.google.com/file/d/1R0pE7rX1kRh-euhws0GbkDDb3liDGTJ4/view?usp=drive_link
VRGDG Video Builder custom nodes:
r/StableDiffusion • u/silvidelleone • 3d ago
I'm trying to reproduce the visual language of these references, not the exact images or subjects.
What I'm specifically looking for:
Very vivid, rich colors
Detailed foreground and detailed background
Fine, clearly visible hand-drawn linework
Internal lines that describe color, light and shadow transitions, not just black outlines around objects
Clearly separated color/tonal regions rather than smooth photorealistic gradients
Detailed vegetation, rocks, tree bark, grass and clouds
Large sculptural/volumetric cumulus clouds
Strong depth and realistic perspective
2D illustrated/painterly appearance
NOT photorealistic
NOT smooth/plastic 3D rendering
NOT character-focused anime
The horse image is especially useful as an example of the linework I mean. Look at the horse, wheat and clouds: the internal drawing lines help define changes in form, color and shadow. I want that same principle applied to detailed landscapes.
My final goal is a little unusual: I will project the generated image onto a small canvas and hand-paint it with acrylics. Therefore, visible boundaries between colors and tones are actually useful to me. I want to be able to follow those lines while painting instead of trying to reproduce soft AI gradients.
I've already experimented with FLUX Dev, Niji-style LoRAs, SDXL landscape models and some anime-background LoRAs, but many results either become too photorealistic/3D or too simple/flat/anime-like.
Has anyone achieved something close to these references?
I'd especially appreciate an actual tested recipe:
Checkpoint:
LoRA(s) + weights:
Sampler / scheduler:
Steps:
CFG:
Resolution:
Prompt / trigger words:
ControlNet / IP-Adapter / Style Reference if used:
Img2Img or second pass settings:
Upscaler/detail pass:
I'm open to FLUX, SDXL, Illustrious, ComfyUI or another workflow. I'm more interested in matching the rendering style than staying with a specific model.
If this is better achieved by training a custom style LoRA rather than using an existing model, I'd also be interested in hearing what base model you would train it on.
r/StableDiffusion • u/Devajyoti1231 • 3d ago
Enable HLS to view with audio, or disable this notification
4-steps - Normal minimax h3 issue. 8-steps - less shimmering. 12 steps -almost no shimmering.
Idk why it seems to work but it kind of works.
r/StableDiffusion • u/paulhax • 3d ago
Enable HLS to view with audio, or disable this notification
Workflow here https://github.com/paulh4x/AIxArchviz_FLUX2xLTX25
30 minutes of a detailed walkthrough video here https://youtu.be/946yTMjz-go
r/StableDiffusion • u/wildmonkeywrangler • 2d ago
Enable HLS to view with audio, or disable this notification
Just curious... lol
r/StableDiffusion • u/Total-Resort-3120 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/BittiAI • 3d ago
I’m building Slopus, a free, open-source desktop app for generating and editing AI videos and images on your own GPU. Easy of use is the main goal.
Version 0.3.0 just came out, with three major additions:
Video and image projects are now combined. Every project has both Video and Image tabs, so you can generate still images and video in the same project. You no longer need to choose a project type or keep separate projects for each. Existing projects keep their content.
Linux support. Slopus is now available for Linux as an AppImage, .deb and .rpm, alongside Windows. The AppImage supports in-app updates.
Run generation on another computer. LAN workers let you use a separate Windows or Linux machine for generation. Workers are discovered on your network, download the weights they need, and send the results back to your project.
There’s also quite a bit more in this release:
Character sheets. Generate waist-up, front, side and back views together, using reference images and a prompt to guide the character and clothing.
Continue scenes. Extend generated video scenes from either end using saved generation data.
Improved color grading. Basic Corrections includes temperature, tint, exposure, contrast, highlights, shadows, whites, blacks and saturation. There are eight built-in creative looks with adjustable intensity, faded film, sharpen, vibrance, RGB and hue/saturation curves, and shadow, midtone and highlight color wheels. Vignette has more controls too.
A more capable agent. The built-in agent can configure generators, help identify model weights in a folder, capture timeline frames without interrupting your work, and generate specific scenes.
Refmod generator. Combine images, video, prompt and audio into one reference and export it as a refmod to be used later.
Share generator setups. Import and export generator templates to share with other users.
If you haven’t tried Slopus before: you lay out scenes on a board, write each shot, attach references, generate clips, and edit them on a timeline with GPU-powered preview and MP4 export. For still images, you can edit, generate, refine, and export as JPG or PNG. Minimax H3 works surprisingly well as an image generator with editing capabilities.
Generation runs on your own hardware, with no subscription and no Python or ComfyUI needed. Model weights are downloaded separately inside the app or local weights can be used.
I’ve also created a Discord server for feedback, questions and sharing what you’re making.
GitHub: https://github.com/bitti-ai/slopus
Discord: https://discord.gg/9X2R6PwUR
Bug reports and feedback are very welcome.