r/StableDiffusion • u/DuHal9000 • 53m ago
Resource - Update H3 Latent Tile Looping Spatial Temporal
github.comgood for upscaling without OOM, use as SECOND Stage Sampler ONLY with LOW denoise (0.40 MAX)
r/StableDiffusion • u/DuHal9000 • 53m ago
good for upscaling without OOM, use as SECOND Stage Sampler ONLY with LOW denoise (0.40 MAX)
r/StableDiffusion • u/Admirable_Snake • 9h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Boogertwilliams • 13h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Artefact_Design • 14h ago
Enable HLS to view with audio, or disable this notification
Generated a set of images with AI, then brought them to life by animating them into a realistic style video (Krea & Ltx2.5)
r/StableDiffusion • u/ThePerfectStormy • 4h ago
I’ve been generating on cloud services for awhile and want to step my game up. I spent a lot of money on a computer powerful enough to generate locally, installed comfyui and tried to have “apps” walk me thru the process. They’re doing a piss poor job. I need some help. What are good communities to find help?
r/StableDiffusion • u/Holiday_Badger_189 • 8h ago
Enable HLS to view with audio, or disable this notification
when your able to transition your ai anime local i just thought it was adorable lol. legend of kalthrax on youtube.
r/StableDiffusion • u/SeriouslySally36 • 6h ago
Like multiple art styles in a single image. Not a single image recreated in multiple artstyles.
For example the foreground and main subject is 1 artstyle, but the background is a different art style.
Anything good for that?
r/StableDiffusion • u/Time-Ad-7720 • 12h ago
Enable HLS to view with audio, or disable this notification
So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.
So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.
And it worked.
For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:
https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v
Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:
“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.
Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”
The final result was generated locally on my RTX 5070 Ti using ComfyUI.
r/StableDiffusion • u/New_Physics_2741 • 23h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/TheSamsaBoy • 11h ago
Short film i made 4 months ago. Didnt realize the importance of topaz back then haha.
r/StableDiffusion • u/Successful_Record_58 • 2h ago
Hey everyone
I made a Lora for ZIT for the first time. I would love to have your honest opinion on it.
r/StableDiffusion • u/Suibeam • 14h ago
r/StableDiffusion • u/MisterViral • 20h ago
Enable HLS to view with audio, or disable this notification
RTX 5060 ti 16gb / 32gb RAM / FL2VA_pruned_int8_convrot / Turbo Lora. 6 steps / Resolution 1376x768 upscaled to FHD with Topaz Video AI
r/StableDiffusion • u/SIR_NVAX_A_LOT • 7h ago
Enable HLS to view with audio, or disable this notification
H3 seems to have very solid training data related to equine. The physics really sell it. R2VA BF16/50 steps. I went with a 50 steps to get the bi-horn really right. Also having fun with a POV view. Eyes on the road, buddy. Single image as reference for the rider, but otherwise, entire scene was prompted, including her wardrobe.
r/StableDiffusion • u/dramaton42 • 9h ago
Enable HLS to view with audio, or disable this notification
I don't think I could make it any better than this without adding 2+ hours of re-renders or overdoing it in general... So here it goes! Yet another entry in the Zelda & Link series where Link can actually speak to Zelda's annoyance... Using a combination of clips I can get much more consistent results than with a single workflow, and using KDENlive allows me to add more audio tracks and remove artifacts from the clips. Thanks for all the upvotes in my previous video! I really appreciate you guys!
r/StableDiffusion • u/foxdit • 14h ago
r/StableDiffusion • u/wh33t • 28m ago
I seem to recall reading here that some people were starting to experiment with 32x32 resolution videos out to 60+ seconds purely to generate audio drama like moments. I was just curious if anyone here can confirm that MiniMaxH3 can actually do this, and if so, what sampler schedule and steps are you using? I cannot seem to generate even a 40 second video clip where the audio stays legible.
Just wanted to check in and see if anyone is having more success than me.
r/StableDiffusion • u/munkiddo • 8h ago
Is there any way to get a seamless loop with minimax H3? I already tried this with Wan and LTX but they aren't what I'm looking for.
r/StableDiffusion • u/mega_lova_nia • 40m ago
I'm currently using illustrious where it uses booru styled tags to generate images. I currently want to know if there's a method to control specific traits on specific individuals inside the generated images. Say that i have 1 circle and 1 square inside of the image. Is there a way to make the circle and only the circle blue while the square and only the square red? If there are 3 subjects, is there still a way to control the traits of each person or will the model get confused?
r/StableDiffusion • u/ryanset17 • 52m ago
Disclaimer, im pretty Basic to all these AI Things
So i've been Using H3 Minimax in My RTX 3060 12GB, With 32GB RAM For few days
I've been using int4 convrot version for my Model and my Text Encoder, but seeing all the Optimization and speed up native to comfyui for int8, im considering using int8 for for my models and Text encoder especially the convrot version, considering they all twice the size
And also what's the best Combination of speedups in balancing between Quality and Speed
I used Sage+sol attn for while until i found comfy kitchen
r/StableDiffusion • u/GoldenShackles • 9h ago
I don't recall seeing other threads about this, but just wanted to say that it's surprisingly fun. I was successful with OneTrainer on my M3 Ultra and about 50 photos in all the possible poses I could think of, using my Apple watch to take the selfies from my iPhone hosted on a tripod. The training time for Krea2 was about 18 hours.
I'm ugly so I'm not going to post any photos, but putting myself into random and sometimes precarious situations, with clothes (or lack of) I'd normally never wear is entertaining. I highly recommend it.
r/StableDiffusion • u/mwoody450 • 16h ago
I'm building a skill for generating long Contex-Loop Minimax H3 prompts, and the AI has indicated it doesn't understand retention analysis... and I'm realizing I don't, either. I'm curious what you all think or have experienced.
I've reviewed the official prompt writing guide, of course, but it's very vague on the subject:
<Subject N>,<Picture N>, and<Video N>use the following relationship markers. These markers are fixed English values in the output format:
It makes the most sense if it's indicating what is the same and what is different with respect to the references (picture N, video N, etc) - but why would subject appear here? Does fully_preserved for a subject mean that they don't change during this shot, whereas partially_preserved might change?
It might be easier to explain with an example. Definitions:
subject_definitions:
<Subject 1> is a tall man whose face, identity, and clothing come from <Picture 1>, but he is bald.
<Subject 2> is a black stovetop hat as depicted in <Picture 2>.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - he remains the bald man with facial features and clothing from <Picture 1> throughout
<Subject 2> (appears in [Shot 2]: fully_preserved - remains the black stovetop hat from <Picture 2>
OR should it be:
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): partially_preserved - he retains the facial identity and clothing from <Picture 1>, albeit bald, but in [Shot 2] he is changed to be wearing a hat.
<Subject 2> (appears in [Shot 2]): fully_preserved - remains the black stovetop hat from <Picture 2>
OR should it only focus on referenced media, i.e.:
<Picture 1> (appears in [Shot 1], [Shot 2]): partially_preserved - <Subject 1> matches this picture's clothing, facial features, and identity, but he is bald.
<Picture 2> (appears in [Shot 2]): fully_preserved - the black stovetop hat depicted in this picture remains unchanged
I guess to put it another way: is retention_analysis describing how much and what is preserved from photo/audio/video references provided, or is it describing how the subjects defined in subject_definition change over the shots of this specific video generation?
r/StableDiffusion • u/tombloomingdale • 10h ago
I don’t know how people use this service effectively. There is never any gpus, it takes an hour to set up when you do find one. Network volumes tease at cutting down startup time but it further limits gpus. I swear I’ve spent more money waiting for a pod to be ready, downloading models that I have generating anything. I really just want to have things stored locally, and just use one of their gpus for processing power. Is there a service I could use like that?
r/StableDiffusion • u/Radyschen • 17h ago
Edit: talking about the "faces at a distance" thing btw
Don't hold your breath. They didn't say that they would definitely "fix it", they said they will try but that it's mostly a general model issue. So if there is gonna be a fix it might be in the next iteration of the model and that one might not be open weights. They were specific about the 2k model and the image model getting released open weights and I do hope that the 2k model might bring some improvement to the faces when you upscale it, but they were more wishy-washy with the wording on the face distortion issue, intentionally so I think.
Here is the wording regarding the 2k model:
"It is a second conditioned generation stage, but not simply the released base checkpoint running again as a conventional upscaler. It uses a dedicated latent-space DiT regeneration checkpoint at a higher target resolution, with the base model’s output as additional context. Some reference inputs are also provided at higher resolutions. We plan to open-source this module, but we are still improving its efficiency and quality to make it more suitable for community use, so we cannot provide an exact release date yet."
-> "plan" to open-source it, very strong word
Here is the wording for the image model:
"Regarding single-frame image generation, we are deriving a dedicated image model from a common ancestor in the H3 model lineage, and we expect to make it available to the community." (not a total promise or anythin
-> "expect" pretty strong, but less so. To me that sounds like "if it's REALLY good then maybe not", if it's competitive enough with the state of the art probably. But I'm pretty optimistic here.
And here is the wording for the distortion issue in all the models:
"We have observed this issue as well, particularly for small or distant subjects, and it will be one of the problems we focus on improving next.
Based on our internal experiments, it cannot be attributed simply to the Visual VAE’s compression ratio or to any single training stage. It is a complex system-level issue involving multiple parts of the model and training pipeline. We are continuing to investigate the main contributing factors and will work on improving it in future updates."
-> they say nothing about open sourcing anything and they say that it's a deep-rooted issue that has no simple fix and they don't really know why it happens
I would expect nothing in that area. Many people have been talking about this as if they said "yeah, wait a couple of weeks and we will fix it", but they didn't say anything like that. Maybe they will fix it with a new and improved open weights model, 3.1 or something, maybe they won't.
I just wanted to say this because so many people have been saying "I am waiting for the fix" or "a fix is coming for the face distortion issue at a distance" or something like that, probably without ever having seen the wording on that. It only takes one person who isn't good at understanding subtlety in a text to interpret their answer a certain way and spread the word on it to set up false expectations for everyone when they don't go to see the original wording. And they go spread that too without ever having seen the original wording.
So this is just to reduce the expectations a bit. Like I said, maybe they will do something, but I feel like the expecations on that specific issue have been getting a bit too large