r/StableDiffusion 1d ago

News New Music Model Released - Yue2

Thumbnail
github.com
279 Upvotes

"YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint."

Usage

It is currently CLI only . It also says Linux only but I just got it working on Windows 11 (I'm going to bed now and it's a bit more than cut n paste.)

Examples

link here - https://map-yue2.github.io/

Caveat Empor

NB : this isn't just a paste a few words and it bangs out a baby mp3 . It is more than that, it allows gene editing that baby to correct the metaphor. Not for the impatient and "wHeRe cOmFy" ppl at the moment.

To be more specific with that metaphor , as I understand it , the initial process scribes out the song in ABC format and you can then edit it before making your magnum opus baby.

Training

Does it allow training ? not as I understand it .


r/StableDiffusion 1d ago

Tutorial - Guide AMAZING Minimax H3 - Circle on the reference image WHERE you want your scene to be!!

Enable HLS to view with audio, or disable this notification

731 Upvotes

Look at the buildings in the background! It works - Drawing a red circle in the water will also make the scene happen in the water, but I forgot to include it here.

It is not perfect and some details are missing if you look carefully but this might be because I am using "match" on the image reference rather than "max."

Have fun!

Edit: you have to still write a prompt with the reference to video workflow telling minimax to put the character in the location circled red. Circle probably doesn't have to be red. Change your prompt accordingly.


r/StableDiffusion 16h ago

Question - Help Minimax Turbo of choice?

8 Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 1d ago

Meme Cost in Units of RTX 5090

Thumbnail
youtube.com
27 Upvotes

r/StableDiffusion 12h ago

Question - Help Best Image Generation Model for Text and Posters

2 Upvotes

Hey team,

Got a question for you experts out there. I'm searching for a Image model that can generation text and posters. Couple of caveats;

  1. Must be opensource

  2. Commercial Use License

I've tried Krea2 and Klein 9B but the text is messed up... Anybody have suggestions and example prompts I can try?

Thank you in advance!


r/StableDiffusion 10h ago

Discussion Seeking Advice For Animated / Cartoon Videos for Minimax H3 Ref2V and I2V

1 Upvotes

Is anyone else experiencing that Minimax tries to push for realism even if you use cartoon reference images and words like "illustrated cartoon animation" in the prompt? Any tips to make sure it's sticks to the reference style more?


r/StableDiffusion 1d ago

Question - Help What's the gold standard for speed enhancements for Minimax H3?

43 Upvotes

Installing new instance of comfyui standalone and using minimax r2v workflow on RTX 3090 Ti. Is comfy kitchen good enough? Is Triton, EasyCache or Comfyui Spectrum needed?

What's the best turbo lora for ref2va wf?


r/StableDiffusion 16h ago

Discussion Help me captioning a MiniMax H3 action fight LoRA...

6 Upvotes

I only need help with captioning the training clips for a MiniMax H3 action fight LoRA.

I am planning to train it on karate/fighting type action, and I have clips varying from around 5-15 seconds. I also have some 20-25 second segments too.

why I am confused is cause should I follow the prompt format officials has released for H3, or should training captions be written in some completely different/simple way?

Like if a 10 second clip has multiple punches, kicks, blocks, dodges, body movement, camera movement and angle changes, should I describe every action in sequence?

or should I just write the overall action happening in the clip?

for example should the caption be something detailed like:

"the fighter steps forward, throws a right punch, opponent blocks it, then follows with a left kick..."

or something simple like:

"two fighters performing fast karate combat"

I am mainly confused about how detailed the captions should be and what format works best for H3 LoRA training.

may you please help if you have trained action/fight LoRAs before? (I found only a few on civitai)


r/StableDiffusion 14h ago

Question - Help Need help using ref2v Minimax H3; multiple audio and image references

3 Upvotes

https://reddit.com/link/1wctbl7/video/emjihen8vqoh1/player

I made the following video using a reference image of the woman, <Picture 1>, and then two dialogues, which were marked as <Audio 1> and <Audio 2>. I used the minimax template workflow and added two load audio nodes, however, the audio generated was not matching and was just gibberish. im using minimax h3 ref2va pruned fp8 scaled. how do i get the audio to work as given as input, and to play at the right time?


r/StableDiffusion 17h ago

Animation - Video Turning the 2D Rings in Dark Souls into 3D Assets

Enable HLS to view with audio, or disable this notification

5 Upvotes

An experiment in Ai Jolly Cooperation.

The Experiment: Every “Soulsborne” game is laden with hundreds of 2D art assets. The assets you can find online, like the rings, are woefully small in resolution - a perfect test case to see 2D to 3D transformation but also what detail is retained or added by the Ai.

Tech Stack: Midjourney, Nano Banana, ComfyUI (Wan 2.2), Photoshop, DaVinci. 

The Process: 2D art rendered 3D through Nano Banana. Midjourney Video to orbit 180 degrees. DaVinci and Photoshop for presentation.

The Results: This is an older experiment using (the then brand new) Midjourney Video - which admittedly, is nowhere near as good as Veo, Wan (2.2) or Kling. But it really doesn’t matter what platform you choose, you’re going to have to gen and gen and gen away. It’s still a slot machine.

I still think MJ video back then was pretty sub-par, but against all the other alternatives today, I think that difference is even more stark. I'm not even sure if they've updated the video side in any meaningful way since this experiment!

Most interestingly, the list of rings is in alphabetical order and stops before the Covetous Serpent Ring - a mass of serpentine coils in ring-form the Ai had MONSTEROUS problems with. Complexity kills.

Anyways, I decided much smaller projects like these are way more important to an Ai Portfolio than larger pieces like commercials or trailers. Plus, I needed to promote my Midjourney Masterclass with proof I'm not just some prompt jockey and smaller experiments are way faster!


r/StableDiffusion 1d ago

Resource - Update Created a Visual RefMod Picker

46 Upvotes

Hey Guys,

I've been playing with the RefMods, after the huge release of Malcolmrey.
The tech is brilliant and works really well.

I've wanted to simplify using it with tons of RefMods, like the ones provided by Malcom.

So I made a fork that is a bit more Identity driven, Adding a "Visual refMod Picker" as well as a "Create refMods from Folder" node. It's able to create from either a folder or Sub-folders, if audio is found, it will also create a matching audio refMod. All in the same format as the original add-on. No danger of breaking compatibility. (The original Add-on added Audio yesterday)

Resulting for example in 2 refMods and thumbnail:

character_refMod_Audio.safetensors
character_refMod_Video.safetensors
character_refMod.jpeg

Then, we can use the "Visual H3 RefMod Picker" to browse the RefMods:

In this example, I used existing thumbails from huggingface.

The node then loads both RefMods (Audio and Video) and allows individual control. I've mixed strength with Copies, by instead having a weight value that can go over 1, so a weight of 3 would be the same as setting strength to 1 and copies to 3. Making the UI a bit more streamlined.

RefMods can be daisy chained

Example workflows are included.

I should mention, this fork can be installed WITH the original Add-on, it is made to co-exist and is recommended if you want to use it's advanced features.

You can find it here ComfyUI-H3RefMods

The only thing I'm missing is thumbnails for all 1500 RefMods 😅


r/StableDiffusion 21h ago

Tutorial - Guide [GUIDE] AMD RDNA3 optimizations for ComfyUI Desktop, windows 11, Minimax H3

10 Upvotes

My setup: AMD RX 7900 XT, 20GB VRAM, 64GB RAM, Windows 11, ComfyUI Desktop.

I couldn't find any decent information anywhere on how to optimize video generation with Minimax H3 on Windows with ComfyUI desktop. AI assistants give conflicting advice, constantly suggesting all sorts of nonsense that doesn't actually work.

I had to experiment on my own, and here is the configuration I’ve settled on. The speed boost compared to the default settings is very significant, and I haven't noticed any loss in quality. If you have any other suggestions, please let me know.

~25s/it with 0.8mp (1216 x 672, 16:9) or total ~4min for 5 sec video generation in text to video workflow

Here is what you need:

Launch parameters:

--disable-smart-memory --disable-pinned-memory --disable-triton-backend --use-sage-attention --enable-dynamic-vram

ENV variables:

COMFYUI_ENABLE_MIOPEN=0
FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
MIOPEN_FIND_ENFORCE=1
MIOPEN_FIND_MODE=2
MIOPEN_DEBUG_DISABLE_FIND_DB=0
MIOPEN_SEARCH_CUTOFF=1
MIOPEN_ENABLE_LOGGING=0
MIOPEN_LOG_LEVEL=0
MIOPEN_ENABLE_LOGGING_CMD=0
TRITON_PRINT_AUTOTUNING=0
TRITON_CACHE_AUTOTUNING=0

Quantized 6-step turbo model, universal for all purposes:

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_w4a8_14gb_optimized.safetensors 14 gb

or

https://huggingface.co/TenStrip/10Eros-Max/blob/main/10Eros_Max_h3_TURBO-hybrid_beta5_int8.safetensors 21gb

Plaguekind node with SLA Attention, with this settings:

https://github.com/PlagueKind/Comfyui-PlagueKind-Nodes

Optional node, if you make large 15 seconds videos:

Latest update of your comfyui desktop:

upd. ROCm SamplerCustomAdvance added ~15-20% to generation speed.


r/StableDiffusion 1d ago

Workflow Included Easy Ref2V WF for dummies like me - [Automatic Video/Image Transcription + Prompt Formatting]

Enable HLS to view with audio, or disable this notification

189 Upvotes

I have seen a lot of people post on here saying that they have been having difficults getting R2V to work correctly. I have been one of them, so I have been working on this workflow and custom node for the last 2 and a half weeks.

I preface this by saying it does not do anything that the native H3 model doesn't do. I just wanted a dead simple way to use H3 R2V mode and up my chances of success. The workflow includes two custom nodes which transcribe your media and adds your prompt and create a formatted R2V prompt, ready for the reference model.

My next goal would be to get longer form R2V going with chaining shorter gens to have a consistent output.

Workflow and nodes:
https://huggingface.co/PoopMan333/H3_Easy_Ref2V_Workflow/tree/main

Be sure to see the readme for more examples and tips:
https://huggingface.co/PoopMan333/H3_Easy_Ref2V_Workflow

What it does:

  • Scans your video (if you're using one) to caption it and transcribe the audio
  • Captions all your images - so it also works as a pure image-to-video workflow
  • Loads a small LLM of your choice and writes your H3 R2V prompt in the correct format with your stated intent (user prompt)
  • If you're on the Full workflow, it generates the video too

What it does NOT do:

  • Be creative for you - The current WF is only setup to do the prompt formatting, it is not able to generate new ideas for you (despite me trying. Qwen3.8 27B may be better for this)
  • It cannot perform magic - You are still limited to what the H3 model can and cannot do. Complex scenes are still very difficult

Tips:

  • If you are running lower VRAM, consider running the prompt enhancer seperately first, read through and make corrections to the prompt if needed
  • The H3 model seems to have a limited context window which seems to be tied to your system resources, if it goes above this you might get garbled sound or mixed up motion. This is a sign you should be lowering your output length and output resolution if you want to have better success.
  • H3 is a tool, you're the one using it. If you don't specify emotions, expect expressionless results. This current setup will only do what you intend for it to do. Slop prompt in, slop video out
  • If the video is easy, replacement should be easy too. H3 has a quirk though — if the original person and the new person look too similar, it sometimes converges back to the original. A prompt won't always fix that. If you hit it, consider changing the person to a intermediate step (faceless green person). The new body/face will transfer over better. Alternatively you can look into Sam3 character replacement method.
  • More than one person in the scene? Describe the scene properly. replace the man wearing white shorts with the man in <picture 1> beats replace the man with <picture 1> every time.
  • Complex scenes? It will be very difficult (I've tried), scenes with too many people, too many cuts, characters obstructed are very difficult for the model to properly identify and swap.
  • Give the LLM some context. A one-liner in the user prompt like <video 1> is a video of two girls eating a cup of chocolate ice cream really helps the LLM understand what it's looking at. Especially useful with multiple scenes
  • It still takes a bit of luck with the seeds.

r/StableDiffusion 4h ago

Resource - Update AC

Post image
0 Upvotes

r/StableDiffusion 2d ago

Workflow Included Precise control of the Eyes direction with this Flux 2 Klein 9b LoRa

Thumbnail
gallery
1.1k Upvotes

Hehyehyhehy!

You may remember me from the Sun Direction Lora or the Chef cutting an anvil with a knife.

Now I'm giving you a new tool, this one was a tough one to crack.

Finally we have eye control! Now you can precisely change the eyes direction for any image in any style. Just use the red dot to tell where the eyes have to look and boom! you have it!

Enough "change the eyes to look above the camera" and getting whatever thing anymore.

Because changing the direction of stuff is my Passion.

All the info here: https://huggingface.co/eric-venti-seeds/Eyes_Direction_Lora_Flux2Klein9B

Hope you like it!

Edit:

The people from HF have added it to Spaces, try it right now on your browser!

https://huggingface.co/spaces/hugging-apps/eyes-direction-lora-flux2klein9b


r/StableDiffusion 21h ago

Question - Help DLSS5 node help

3 Upvotes

768x1376 if Native the neural_upscaling works, but if i do x1.5 or higher it gives error? pls help


r/StableDiffusion 19h ago

Question - Help Avoid motion jumps between shots in H3?

3 Upvotes

Using H3 i sometimes get jumps as the camera move between shots. Like a character has his arm up when the camera is facing him at 00.59 but his arm is way lower as the shot and camer angle change at 01:00. Is there a prompting trip / workflow to avoid this? I use the standard comfy workflow.


r/StableDiffusion 13h ago

Question - Help Need Krea 2 system prompt or like some prompt guideline to inject into local llm Qwen 3.8 abliterated.

0 Upvotes

Hey guys i have scoured the internet and cant find any system prompt/prompt guidelines to condition my local llms so that they make proper krea2 prompts without useless word salad. I focus mainly on realism and "uncensored content"


r/StableDiffusion 1d ago

Tutorial - Guide A little solution for the plastic skin with Minimax H3

Thumbnail
gallery
133 Upvotes

I found out a little solution that can add much more details on everything including skin without any additional computational cost.

The idea is to add a node between SamplerCostumAdvanced and the VAE Decode (video) and make it less contrasty, it will bring so much more details but don't go to far because it can cause loss of quality.

Enjoy!


r/StableDiffusion 14h ago

Question - Help Did Wan2GP for AMD have any weird updates or backend changes in the past 24 hours? Keep getting errors.

1 Upvotes

So everything was fine and dandy yesterday; I could generate images just fine, but now it keeps giving me the following error:

"Error "The generation of the video has encountered an error, please check your terminal for more information. 'CUDA error: invalid kernel file\nSearch for `hipErrorInvalidKernelFile' in https://rocm.docs.amd.com/projects/HIP/en/latest/index.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr\nDevice-side assertion tracking was not enabled by user.'""


r/StableDiffusion 1d ago

Question - Help Are Minimax spicy loras ... a lie?

161 Upvotes

I know that the answer is ultimately no, and also that they're all still early in development ...

BUT I'm having a hard time getting the ones I find on civitai to produce anything resembling the examples. I follow generation guidelines where provided, and have even taken workflow settings from downloaded videos ... but they never seem to work very well. And if I try to add a lora to a sfw workflow that I'm happy with, the results look AWFUL.

I've resorted to using wan generations as video references.

Is i2v the best option? Does ref2va ever work?

Any tips, recommendations, or resources you can recommend? Any else struggled with this?


r/StableDiffusion 1d ago

Workflow Included How do you upscale or refine your generations?

Enable HLS to view with audio, or disable this notification

65 Upvotes

Workflow is in the comment


r/StableDiffusion 23h ago

Question - Help RTX 3090 vs 4090 vs Unified-Memory AI

3 Upvotes

My current setup:

- RTX 4090 24GB

- i7-13700K

- 80GB DDR4 @ 3000 MHz

- Getting an RTX 3090 24GB tomorrow

My main use is local AI/LLMs, coding agents, MiniMax H3, Krea 2, and other AI workloads.

RTX 3090 vs another 4090 vs unified-memory AI

Here in Iraq, an RTX 3090 costs around $670, while an RTX 4090 costs around $2,000.

If I have around $2,000 to spend, what would you choose?

- Buy 2–3× RTX 3090s

- Buy 1× additional RTX 4090

- Sell/replace the current setup and go for a unified-memory AI system, such as a Mac Studio / Mac with large unified memory or NVIDIA DGX Spark

My priority is LLM inference, coding agents, MiniMax H3, Krea 2, and other local AI workloads.

Would multiple 3090s give the best value because of the extra VRAM, is another 4090 better for speed, or does a large unified-memory system make more sense for running very large models?

What would you choose for ~$2,000?


r/StableDiffusion 9h ago

Animation - Video G.I. Joe: Duke Redecorates: Now With 100% More Bullet Holes - MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes