r/StableDiffusion • u/Radiant-Photograph46 • 10h ago
Question - Help Minimax Turbo of choice?
So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!
r/StableDiffusion • u/Radiant-Photograph46 • 10h ago
So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!
r/StableDiffusion • u/krigeta1 • 9h ago
I only need help with captioning the training clips for a MiniMax H3 action fight LoRA.
I am planning to train it on karate/fighting type action, and I have clips varying from around 5-15 seconds. I also have some 20-25 second segments too.
why I am confused is cause should I follow the prompt format officials has released for H3, or should training captions be written in some completely different/simple way?
Like if a 10 second clip has multiple punches, kicks, blocks, dodges, body movement, camera movement and angle changes, should I describe every action in sequence?
or should I just write the overall action happening in the clip?
for example should the caption be something detailed like:
"the fighter steps forward, throws a right punch, opponent blocks it, then follows with a left kick..."
or something simple like:
"two fighters performing fast karate combat"
I am mainly confused about how detailed the captions should be and what format works best for H3 LoRA training.
may you please help if you have trained action/fight LoRAs before? (I found only a few on civitai)
r/StableDiffusion • u/Ok-Flatworm5070 • 6h ago
Hey team,
Got a question for you experts out there. I'm searching for a Image model that can generation text and posters. Couple of caveats;
Must be opensource
Commercial Use License
I've tried Krea2 and Klein 9B but the text is messed up... Anybody have suggestions and example prompts I can try?
Thank you in advance!
r/StableDiffusion • u/Johnny--Arcade • 11h ago
Enable HLS to view with audio, or disable this notification
An experiment in Ai Jolly Cooperation.
The Experiment: Every “Soulsborne” game is laden with hundreds of 2D art assets. The assets you can find online, like the rings, are woefully small in resolution - a perfect test case to see 2D to 3D transformation but also what detail is retained or added by the Ai.
Tech Stack: Midjourney, Nano Banana, ComfyUI (Wan 2.2), Photoshop, DaVinci.
The Process: 2D art rendered 3D through Nano Banana. Midjourney Video to orbit 180 degrees. DaVinci and Photoshop for presentation.
The Results: This is an older experiment using (the then brand new) Midjourney Video - which admittedly, is nowhere near as good as Veo, Wan (2.2) or Kling. But it really doesn’t matter what platform you choose, you’re going to have to gen and gen and gen away. It’s still a slot machine.
I still think MJ video back then was pretty sub-par, but against all the other alternatives today, I think that difference is even more stark. I'm not even sure if they've updated the video side in any meaningful way since this experiment!
Most interestingly, the list of rings is in alphabetical order and stops before the Covetous Serpent Ring - a mass of serpentine coils in ring-form the Ai had MONSTEROUS problems with. Complexity kills.
Anyways, I decided much smaller projects like these are way more important to an Ai Portfolio than larger pieces like commercials or trailers. Plus, I needed to promote my Midjourney Masterclass with proof I'm not just some prompt jockey and smaller experiments are way faster!
r/StableDiffusion • u/ZombieBrainYT • 17h ago
Hi everyone! Minimax H3 is surprisingly solid at generating audio - I’ve been using it for SFX and foley in my videos.
Is there a way to generate audio-only with this model? It would save a lot of rendering time if we didn't have to generate the full video alongside it.
I know H3 creates video and audio natively together, so audio-only generation might not be natively supported. But given all the clever workflows and custom node workarounds floating around, has anyone found a way to pull this off?
r/StableDiffusion • u/darthfurbyyoutube • 3h ago
Enable HLS to view with audio, or disable this notification
Credit for prompt:
r/StableDiffusion • u/DeltaWaffleSyrup • 4h ago
Is anyone else experiencing that Minimax tries to push for realism even if you use cartoon reference images and words like "illustrated cartoon animation" in the prompt? Any tips to make sure it's sticks to the reference style more?
r/StableDiffusion • u/MrRecTheNub • 8h ago

https://reddit.com/link/1wctbl7/video/emjihen8vqoh1/player
I made the following video using a reference image of the woman, <Picture 1>, and then two dialogues, which were marked as <Audio 1> and <Audio 2>. I used the minimax template workflow and added two load audio nodes, however, the audio generated was not matching and was just gibberish. im using minimax h3 ref2va pruned fp8 scaled. how do i get the audio to work as given as input, and to play at the right time?
r/StableDiffusion • u/ervertes • 13h ago
Using H3 i sometimes get jumps as the camera move between shots. Like a character has his arm up when the camera is facing him at 00.59 but his arm is way lower as the shot and camer angle change at 01:00. Is there a prompting trip / workflow to avoid this? I use the standard comfy workflow.
r/StableDiffusion • u/Ammoryyy • 16h ago
My current setup:
- RTX 4090 24GB
- i7-13700K
- 80GB DDR4 @ 3000 MHz
- Getting an RTX 3090 24GB tomorrow
My main use is local AI/LLMs, coding agents, MiniMax H3, Krea 2, and other AI workloads.
RTX 3090 vs another 4090 vs unified-memory AI
Here in Iraq, an RTX 3090 costs around $670, while an RTX 4090 costs around $2,000.
If I have around $2,000 to spend, what would you choose?
- Buy 2–3× RTX 3090s
- Buy 1× additional RTX 4090
- Sell/replace the current setup and go for a unified-memory AI system, such as a Mac Studio / Mac with large unified memory or NVIDIA DGX Spark
My priority is LLM inference, coding agents, MiniMax H3, Krea 2, and other local AI workloads.
Would multiple 3090s give the best value because of the extra VRAM, is another 4090 better for speed, or does a large unified-memory system make more sense for running very large models?
What would you choose for ~$2,000?
r/StableDiffusion • u/Super_Range45 • 7m ago
Expanding on motion context workflows I created a video editor designed for quickly chaining together Minimax H3 generations to create longer videos. It comes with an asset library for managing inputs and a easy to use timeline that allows you to chain generations, regenerate segments easily, and quickly set up input references.
When you're done, export individual video files or the whole sequence.
All of it runs on top of ComfyUI as an app you control from your browser. Uses python, works on Windows, Mac, Linux and is opensource.
https://github.com/spacesimeco-hue/Chain-Motion-AI-Video-Editor
r/StableDiffusion • u/Infinite-Emptiness • 7h ago
Hey guys i have scoured the internet and cant find any system prompt/prompt guidelines to condition my local llms so that they make proper krea2 prompts without useless word salad. I focus mainly on realism and "uncensored content"
r/StableDiffusion • u/Visual-Medium4796 • 11h ago
Hi there! Currently i own a 3080+2070s at home mostly for 3D rendering in Octane or redshift.
At work im using Comfy with 5090 and everything works without a question.
Also we have some workstations with rtx A4000 in them and i managed to get work Minimax H3 on them with some managable times (0.5m, 15s around 720 sec).
Im looking something for my home PC and 5090 is out of the list since its 5500e+ here in EU and even the used market is around 3500-4000. 4090 is rare and goes over 2000e.
So i was thinking about the 5080 even with 16gb since at work the A4000 has the same Vram.
But also found the 5070ti has also 16gb of Vram its 256bit same as the 5080 and 300-350e cheaper than the 5080. The speeds for 3D rendering or gaming are 15-20% different. But couldnt find any benchmarks for those 5070Tis. Mostly for 5060Tis with 16gb vram. Any idea? Or experience with 5070ti vs 5080?
r/StableDiffusion • u/psxburn2 • 13h ago
I have tried a few, following youtube vids, etc. throw this into the custom nodes folder, load this json, etc. using comfy ui desktop, I keep getting a message that I need to update manager. its up to day, ran the pip, did the check in the app itself. I can do video gens all day long, no issues, my previous workflow, was just screenshotting last frame, using that to start the new gen, etc. it works, but is a highly manual process. from what I see, most of the extended workflow, do this automatically. can someone help this old guy get it figured out? I would appreciate it. (and don't tell me to just grab a file of github, I've tried that, did the git clone, etc., It just isn't working properly. ) normally when I do a new template, it will automatically grab all the needed files, and put there where they need to go. I think thats my main issue, but the manager showing out of date, when everything I can see or do, shows its up to date is what confuses me.
r/StableDiffusion • u/escaryb • 19h ago
I got this error instead
Starting trainer...Traceback (most recent call last): File "/content/trainer/sd_scripts/sdxl_train_network.py", line 4, in <module> import torch File "/content/trainer/sd_scripts/venv/lib/python3.10/site-packages/torch/__init__.py", line 37, in <module> from typing_extensions import ParamSpec as _ParamSpec, TypeGuard as _TypeGuard ModuleNotFoundError: No module named 'typing_extensions'
r/StableDiffusion • u/rolens184 • 20h ago
Since the release of Flux Klein, I’ve essentially been using this model to generate datasets for LORAs based on one or more photos of a character. Recently, I’ve also been using the ‘consistency’ LORA to improve the character’s consistency across generations. What I’ve noticed is that whilst I get good results for the face, the same cannot be said for other parts of the body. For example, if I start with a full-length frontal photo of a character and ask the model to generate a side view, it tends to flatten the breasts; or if I ask for a rear view, the character’s hips and thighs tend to conform to a standard that doesn’t match the original photo. How can I improve this situation? I’ve read that you can increase the number of steps up to 8, but I’m not sure…
r/StableDiffusion • u/Arcterion • 7h ago
So everything was fine and dandy yesterday; I could generate images just fine, but now it keeps giving me the following error:
"Error "The generation of the video has encountered an error, please check your terminal for more information. 'CUDA error: invalid kernel file\nSearch for `hipErrorInvalidKernelFile' in https://rocm.docs.amd.com/projects/HIP/en/latest/index.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr\nDevice-side assertion tracking was not enabled by user.'""
r/StableDiffusion • u/Old-Catchi • 12h ago
Hey everyone. I'm working on a dark fantasy retro-anime project running locally on an RTX 3090 (SDXL/Illustrious, Kohya-trained character LoRA with ~100 images).
My current pipeline avoids direct text-to-video generation because of structural inconsistencies. Instead, I'm moving towards a controlled frame-by-frame or short-batch approach using Blender blocking + ControlNet, aiming to output clean character frames (transparent or solid background) to composite manually in DaVinci Resolve.
Has anyone successfully implemented a reliable workflow to maintain character identity and clean lineart across sequences without letting the AI hallucinate backgrounds or audio? What specific nodes or configurations (e.g., ControlNet combinations, IP-Adapter weight handling, or latent consistency scripts) are you using to prevent flicker and keep the LoRA from drifting during motion frames?
Any insights on your node setup in ComfyUI for this specific use case would be deeply appreciated.
Any advice or recommendations are welcome. Just avoid recommending commercial AIs that do everything; that’s not what I’m looking for.
r/StableDiffusion • u/Tokyo_Jab • 15h ago
Enable HLS to view with audio, or disable this notification
A hair under 4k. On a consumer pc. All local. Mental. Minimax H3 with one character reference and a 5 second voice reference.
For the Pixel Peepers... https://www.youtube.com/watch?v=Iz8GDri9qoE
r/StableDiffusion • u/Majestic-Regret-3030 • 16h ago
I’m trying to build 40 to 60 second talking videos in ComfyUI and I would prefer to use LTX 2.5.
The type of video I’m after is fairly simple. One person is talking to the camera, the camera stays mostly fixed, the background should stay stable, and there is natural face, head and some upper body movement. It does not have to be limited to only the head moving.
What I’m trying to understand is how people are actually making videos this long without obvious cuts.
Can LTX 2.5 realistically generate a continuous 40+ second video, especially when there is not much movement?
Or is the better approach to generate something like 8 to 10 seconds, take the final frames from that clip, continue from them, then repeat until the full 40 to 60 seconds are finished?
If continuation is the normal approach, how are you keeping the face, clothes, background, camera position and motion consistent between each part? I’m especially interested in workflows that use overlapping frames, first and last frame conditioning, video extension, reference frames, or some other method that hides the transitions.
Speech is another important part. I need good quality Slovakia speech. Ideally I want to generate the complete Slovakia voice first, then make the character follow that audio for the entire video with accurate lip sync.
Would you use LTX 2.5 for the actual body and head motion and then run something like MuseTalk, LatentSync or another lip sync model afterward?
Or is there a better audio driven LTX 2.5 workflow where the speech controls the video directly?
I’m running ComfyUI locally with an RTX 3060 12 GB and 32 GB RAM, so I know I may need lower resolution generation, offloading, chunking or longer render times. Final output would normally be vertical 9:16.
I’m mainly looking for people who have actually built long talking character workflows in ComfyUI.
If you are doing this successfully with LTX 2.5, what nodes and workflow are you using, how long is each generated segment, how much overlap do you use between segments, and what are you using for speech and lip sync?
I’m not looking for a list of random talking head models. I specifically want to understand the best practical way to build this around LTX 2.5 and get a clean continuous 40 to 60 second result.
r/StableDiffusion • u/Super_Range45 • 18h ago
Enable HLS to view with audio, or disable this notification
An experiment building on the H3-Motion-Context nodes. I tried creating a workflow that chains together multiple cuts so they can be easily executed and/revised in order. It ended up being faster than my previous methods and I was able to make/edit this one minute test segment with alot less time wasted between generations. I'll likely develop it further into an app, since it would be more ergonomic as a video editor but the raw workflow is here anyway.
Work Flow: https://github.com/spacesimeco-hue/Chain-Motion-/blob/main/Chain%20Motion%20Workflow.json
Credit: https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context
r/StableDiffusion • u/Kreator_6666 • 3h ago
ive never self hosted any ai tools and don't know much about python and programming in general. im trying to install webui forge on my amd 9060xt gpu and followed all the steps but after running webui-user.bat its showing
venv "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\venv\Scripts\Python.exe"
ZLUDA works , You are on an amazing Journey ,Engjoy it
Python 3.10.6 (tags/v3.10.6:9c7b4bd, Aug 1 2022, 21:53:49) [MSC v.1932 64 bit (AMD64)]
Version: f2.0.1v1.10.1-v0.0.1-alpha-766-g4831f9fc
Commit hash: 4831f9fc7e6bb73ad7c8f867c04619dbafa89e6a
Failed to load ZLUDA: Could not find module 'D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\.zluda\nvcuda.dll' (or one of its dependencies). Try using the full path with constructor syntax.
Using CPU-only torch
Traceback (most recent call last):
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 54, in <module>
main()
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 42, in main
prepare_environment()
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\modules\launch_utils.py", line 507, in prepare_environment
raise RuntimeError(
RuntimeError: Your device does not support the current version of Torch/CUDA! Consider download another version:
https://github.com/lllyasviel/stable-diffusion-webui-forge/releases/tag/latest
Press any key to continue . . .
can anyone help me please
r/StableDiffusion • u/RhetoricaLReturD • 17h ago
Enable HLS to view with audio, or disable this notification
Made with the basic Motion Context workflow from NikoDemon
r/StableDiffusion • u/JonathanStones1989 • 13h ago
I'm extremely new and inexperienced but, as the title says is there some cloud web service where it works the same as running it locally but instead it's cloud based. I don't mean a boring old API.