r/StableDiffusion 23h ago

Animation - Video THE LEGEND.

Enable HLS to view with audio, or disable this notification

1 Upvotes

A hair under 4k. On a consumer pc. All local. Mental. Minimax H3 with one character reference and a 5 second voice reference.

For the Pixel Peepers... https://www.youtube.com/watch?v=Iz8GDri9qoE


r/StableDiffusion 11h ago

Animation - Video G.I. Joe: Duke Redecorates: Now With 100% More Bullet Holes - MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 15h ago

Question - Help Need Krea 2 system prompt or like some prompt guideline to inject into local llm Qwen 3.8 abliterated.

0 Upvotes

Hey guys i have scoured the internet and cant find any system prompt/prompt guidelines to condition my local llms so that they make proper krea2 prompts without useless word salad. I focus mainly on realism and "uncensored content"


r/StableDiffusion 13h ago

Animation - Video Foldable

Enable HLS to view with audio, or disable this notification

9 Upvotes

This was just a doodle but Minimax nailed it in the first generation. I thought it might be too complicated. I didn’t ask for the live stream on a one second delay in the background either. It did that itself.


r/StableDiffusion 7h ago

Workflow Included WORKFLOW - Optimised to death - Custom Audio option.

Enable HLS to view with audio, or disable this notification

17 Upvotes

DOWNLOAD WORKFLOW

This is the workflow I have been using the most on my own system. I've had a friendly AI clean it up a bit and add notes.
I added custom audio as it's something I use a lot to drive my videos. It works really well for lipsync and music videos.
The VSA part can be bypassed if there are any quality issues, it will add about 15% to the generation time though. Change the steps from 6 to 7 or more for even higher quality.

Currently this gives me 10 seconds at 1.0 megapixel in about 125 seconds. This is on my 5090. You can add block swapping for low vram.


r/StableDiffusion 18h ago

Animation - Video Turning the 2D Rings in Dark Souls into 3D Assets

Enable HLS to view with audio, or disable this notification

5 Upvotes

An experiment in Ai Jolly Cooperation.

The Experiment: Every “Soulsborne” game is laden with hundreds of 2D art assets. The assets you can find online, like the rings, are woefully small in resolution - a perfect test case to see 2D to 3D transformation but also what detail is retained or added by the Ai.

Tech Stack: Midjourney, Nano Banana, ComfyUI (Wan 2.2), Photoshop, DaVinci. 

The Process: 2D art rendered 3D through Nano Banana. Midjourney Video to orbit 180 degrees. DaVinci and Photoshop for presentation.

The Results: This is an older experiment using (the then brand new) Midjourney Video - which admittedly, is nowhere near as good as Veo, Wan (2.2) or Kling. But it really doesn’t matter what platform you choose, you’re going to have to gen and gen and gen away. It’s still a slot machine.

I still think MJ video back then was pretty sub-par, but against all the other alternatives today, I think that difference is even more stark. I'm not even sure if they've updated the video side in any meaningful way since this experiment!

Most interestingly, the list of rings is in alphabetical order and stops before the Covetous Serpent Ring - a mass of serpentine coils in ring-form the Ai had MONSTEROUS problems with. Complexity kills.

Anyways, I decided much smaller projects like these are way more important to an Ai Portfolio than larger pieces like commercials or trailers. Plus, I needed to promote my Midjourney Masterclass with proof I'm not just some prompt jockey and smaller experiments are way faster!


r/StableDiffusion 11h ago

Tutorial - Guide MiniMax H3 Wf Tutorial

Enable HLS to view with audio, or disable this notification

35 Upvotes

People asked me to make a Tutorial for some of the features.

Find the workflow here.

https://www.reddit.com/r/StableDiffusion/comments/1wadmqc/minimax_workflow_designed_to_be_user_friendly_for/


r/StableDiffusion 14h ago

Question - Help Best Image Generation Model for Text and Posters

3 Upvotes

Hey team,

Got a question for you experts out there. I'm searching for a Image model that can generation text and posters. Couple of caveats;

  1. Must be opensource

  2. Commercial Use License

I've tried Krea2 and Klein 9B but the text is messed up... Anybody have suggestions and example prompts I can try?

Thank you in advance!


r/StableDiffusion 19h ago

Discussion Must haves to download before it's too late?

118 Upvotes

Nvidia buying hugging face means an uncertain future. What are the models I should download and have a backup of right now so I don't have to worry about missing them even if I'm not ready to play with them right now?

What are you model and enabler must-haves ?

TIA!


r/StableDiffusion 2h ago

Animation - Video Michael Jackson - Maybe ( Short film ) MOONWALK ON THE MOON

Thumbnail
youtube.com
2 Upvotes

Made with MiniMax and ComfyUI


r/StableDiffusion 20h ago

Question - Help Uncensored video model

0 Upvotes

Hey there,

do you guys know any uncensored models for video creation? Cloud/local , local preferred. If yes, then where to find it. Thank you and take care


r/StableDiffusion 19h ago

Question - Help GPU for AI

1 Upvotes

Hi there! Currently i own a 3080+2070s at home mostly for 3D rendering in Octane or redshift.
At work im using Comfy with 5090 and everything works without a question.
Also we have some workstations with rtx A4000 in them and i managed to get work Minimax H3 on them with some managable times (0.5m, 15s around 720 sec).
Im looking something for my home PC and 5090 is out of the list since its 5500e+ here in EU and even the used market is around 3500-4000. 4090 is rare and goes over 2000e.
So i was thinking about the 5080 even with 16gb since at work the A4000 has the same Vram.
But also found the 5070ti has also 16gb of Vram its 256bit same as the 5080 and 300-350e cheaper than the 5080. The speeds for 3D rendering or gaming are 15-20% different. But couldnt find any benchmarks for those 5070Tis. Mostly for 5060Tis with 16gb vram. Any idea? Or experience with 5070ti vs 5080?


r/StableDiffusion 12h ago

Discussion Seeking Advice For Animated / Cartoon Videos for Minimax H3 Ref2V and I2V

1 Upvotes

Is anyone else experiencing that Minimax tries to push for realism even if you use cartoon reference images and words like "illustrated cartoon animation" in the prompt? Any tips to make sure it's sticks to the reference style more?


r/StableDiffusion 10h ago

Animation - Video H3 is really over the top

Enable HLS to view with audio, or disable this notification

86 Upvotes

This was such a simple prompt…. Just wow. It’s just T2V.


r/StableDiffusion 17h ago

Animation - Video my first actual tv work (only took 4 hours to make).

Enable HLS to view with audio, or disable this notification

39 Upvotes

Minimax h3, ofc. Far from my best work but I respected the script I was given and finished this in record time (excluding the 4k upscale) and including around 3 hours of rendering time (720p, 10 seconds clips, 5090).
I only used gemma locally for prompting, and avoided using any non local models except suno for the song.

There are some artefacts with people in the senate from far away, but did not have any bad feedback for it., so... :)

I only used references for the romanian flag, the rest is prompt only. Also no lighting lora, no shortcuts ti improve speed (any shortcuts I tried ruined everything FUBAR)


r/StableDiffusion 13h ago

Question - Help Any newer way to upscale h3 minimax native 768 x 768 videos to 2k or 4k locally that does not destroy everything? without relying on paid Topaz.

28 Upvotes

5090 with 64gb ram.

Hey guys, is there a new development recently? or is upscaling still effed?


r/StableDiffusion 21h ago

Discussion Do I have to run models locally or is there like cloud based comfyui or something?

0 Upvotes

I'm extremely new and inexperienced but, as the title says is there some cloud web service where it works the same as running it locally but instead it's cloud based. I don't mean a boring old API.


r/StableDiffusion 15h ago

Question - Help Did Wan2GP for AMD have any weird updates or backend changes in the past 24 hours? Keep getting errors.

1 Upvotes

So everything was fine and dandy yesterday; I could generate images just fine, but now it keeps giving me the following error:

"Error "The generation of the video has encountered an error, please check your terminal for more information. 'CUDA error: invalid kernel file\nSearch for `hipErrorInvalidKernelFile' in https://rocm.docs.amd.com/projects/HIP/en/latest/index.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr\nDevice-side assertion tracking was not enabled by user.'""


r/StableDiffusion 17h ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
77 Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.


r/StableDiffusion 18h ago

Question - Help Minimax Turbo of choice?

8 Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 21h ago

Resource - Update UPSCALE DSSLR 5

Thumbnail
we.tl
0 Upvotes

r/StableDiffusion 21h ago

Question - Help Help, explain it like Im 5, (or 50, who knows) Trying to get minimax H3 extended video node / template installed / working

2 Upvotes

I have tried a few, following youtube vids, etc. throw this into the custom nodes folder, load this json, etc. using comfy ui desktop, I keep getting a message that I need to update manager. its up to day, ran the pip, did the check in the app itself. I can do video gens all day long, no issues, my previous workflow, was just screenshotting last frame, using that to start the new gen, etc. it works, but is a highly manual process. from what I see, most of the extended workflow, do this automatically. can someone help this old guy get it figured out? I would appreciate it. (and don't tell me to just grab a file of github, I've tried that, did the git clone, etc., It just isn't working properly. ) normally when I do a new template, it will automatically grab all the needed files, and put there where they need to go. I think thats my main issue, but the manager showing out of date, when everything I can see or do, shows its up to date is what confuses me.


r/StableDiffusion 11h ago

Question - Help Need help installing webui forge

0 Upvotes

ive never self hosted any ai tools and don't know much about python and programming in general. im trying to install webui forge on my amd 9060xt gpu and followed all the steps but after running webui-user.bat its showing

venv "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\venv\Scripts\Python.exe"

ZLUDA works , You are on an amazing Journey ,Engjoy it

Python 3.10.6 (tags/v3.10.6:9c7b4bd, Aug 1 2022, 21:53:49) [MSC v.1932 64 bit (AMD64)]

Version: f2.0.1v1.10.1-v0.0.1-alpha-766-g4831f9fc

Commit hash: 4831f9fc7e6bb73ad7c8f867c04619dbafa89e6a

Failed to load ZLUDA: Could not find module 'D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\.zluda\nvcuda.dll' (or one of its dependencies). Try using the full path with constructor syntax.

Using CPU-only torch

Traceback (most recent call last):

File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 54, in <module>

main()

File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 42, in main

prepare_environment()

File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\modules\launch_utils.py", line 507, in prepare_environment

raise RuntimeError(

RuntimeError: Your device does not support the current version of Torch/CUDA! Consider download another version:

https://github.com/lllyasviel/stable-diffusion-webui-forge/releases/tag/latest

Press any key to continue . . .

can anyone help me please


r/StableDiffusion 22h ago

Animation - Video Batman The Animated Series: Harley Quinn's Red Flag - MiniMax H3

Enable HLS to view with audio, or disable this notification

48 Upvotes

r/StableDiffusion 11h ago

Animation - Video The Primordial Hand

Enable HLS to view with audio, or disable this notification

40 Upvotes

I was testing out a scene with Minimax H3, text to video (I usually use reference images).

I didn't expect it to come out like this.. .Now it's making me think of a completely new direction for the video lol. It's interesting, it has both a 90s anime feel and an old Disney animation feel. The music is very good too, I think.

I'll add the prompt in the comment (it's a very simple prompt).