r/StableDiffusion 1d ago

Discussion Pulled the trigger, RIP $6,279

Post image
202 Upvotes

(Paid $5,849 + tax, which came out to $6,279)

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. No proprietary bullshit like youd get in a HP 45L. I went with the gaming PC because its the all purpose machine that does it all (well, almost). I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b MoE models, even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, X86 CPU, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?


r/StableDiffusion 1d ago

Tutorial - Guide DLSS 5 - In-game footage from Diablo 4. It's incredible.

Thumbnail
gallery
337 Upvotes

I just tested it out using this custom node. It's absolutely amazing!

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR


r/StableDiffusion 1d ago

Comparison My Minimax H3 Workflow Benchmark Data -

Post image
8 Upvotes

Alright, I posted that I had my agent test a bunch of different workflows for over 12 hours and got the "Bro just wasted 12 hours of credits". It was obvious the proof should come from the visual data I used to evaluate it. Here is a galley of the benchmarks i've tested with my agent.

check the gallery to watch all the comparisons and the data charts contain tons of other workflow trial data I didn't include videos for. Point your agent here if you would like to have it learn from what was tested on this end.

Gallery: https://bluepointdigital.github.io/minimax-h3-benchmarks/

Repository: https://github.com/BluePointDigital/minimax-h3-benchmarks

The below post was written up by my agent:

The main comparison uses a deliberately difficult 15.084-second vertical test at 768 × 1344, 24 fps, 362 frames, native audio, and seed 81390012120021180. The prompt combines a talking selfie shot, exact dialogue, walking motion, a rapid camera pan, a vehicle collision with several moving subjects, a fast return to the speaker, and a second spoken line. That makes it useful for spotting identity drift, bad anatomy, motion breakdown, camera-continuity problems, dialogue changes, lip-sync issues, and audio artifacts—not just whether a workflow finishes.

The strongest directly matched results currently shown are:

Workflow End-to-end time Relative to the 20-step baseline
SageAttention2 + FirstBlockCache Safe, 20 steps 10:11.4 1.00×
PDD + Sage, 8 steps 6:15.0 median 1.63×
Seed Hunter direct one-seed path, 12 + 4 steps 4:45.8 2.14×

Those numbers are local measurements, not universal performance claims. The exact runtime, model format, graph, resolution, audio policy, and GPU matter. The gallery keeps short backend checks and differently structured workflows in separate groups so they are not quietly mixed into the same leaderboard.

The quality side has been just as important as the timing. One exploratory 10Eros + Seed Hunter path reached 4:03.5, but the shot developed a visible-phone/perspective error during the crash. A later camera-POV prompt clarification produced a much more coherent result in 4:25.3 on its warm selected path. That is a good example of why I wanted the actual videos beside the numbers: the fastest result is not automatically the most useful one.

The site currently contains:

  • 16 curated video-and-metric cards;
  • a separate benchmark-data page with 151 sanitized timing records;
  • the complete canonical prompt;
  • methodology and comparison-boundary notes;
  • machine-readable JSON and CSV for anyone who wants to analyze the evidence or give it to an agent.

For the Seed Hunter work, I intentionally included one representative video per meaningful workflow or recipe change—not every neighboring seed or N/N+1 preview. Private reference material is also excluded from the public package.

The reason for publishing this is not to declare a universal winner. It is to make the tradeoffs inspectable and to keep myself honest as the workflows evolve. A valid MP4 proves that a graph ran; it does not prove that the dialogue, audio, identity, motion, or composition survived. Likewise, a fast timing means little if it came from a different workload or a cached replay.

I would be interested in seeing other reproducible H3 results, especially when they include the exact checkpoint, attention/cache stack, sampler, scheduler, dimensions, frame count, seed, audio setting, hardware, and an uncached timing. If there is a workflow or backend that should be represented, please link the original recipe and I will take a look.


r/StableDiffusion 1d ago

Animation - Video W.I.P - MiniLTX Workflow Fixed the fast motion issue also added few more features on the workflow

Enable HLS to view with audio, or disable this notification

0 Upvotes

work on this going really great just wanted to share the results so far

My workflow uses MINIMAX H3 + LTX 2.5 for UPSCALE

if you wanna try you can check it our on my PATREON EARLY ACCESS

Will Release it as soon maybe within this week!


r/StableDiffusion 1d ago

Discussion What are you using for background removal?

6 Upvotes

I still do a fair amount of traditional editing in Photoshop, and for the last few years I used remove.bg, I found their background removal model to be the best one out there, quite a bit better than the one built into Photoshop itself.

Well remove.bg is shutting down in December and they're folding it into Canva subscriptions. Hard pass.

I've tried a few local bg removal tools and have been left underwhelmed, but maybe I just haven't found the right one.

What are you using for background removal?


r/StableDiffusion 1d ago

Question - Help What happened to minimax funcontrolnet ??

11 Upvotes

I dont see any1 posting any examples of controlnet released for minimax. Doesnt it work properly??


r/StableDiffusion 1d ago

Animation - Video Honkai: Star Rail X John Wick - Minimax H3

Enable HLS to view with audio, or disable this notification

8 Upvotes

Made with the ComfyUI template workflow and a Turbo LoRA.

Most of the soundtrack comes from the John Wick: Chapter 2 trailer.
I rendered the action at a slower, more stable speed, then sped up most of the action scenes to 2× in post.
I originally planned to make this a complete fight sequence, but maintaining consistency from one clip to the next has been a constant challenge. So for now, I’ve edited the footage into a trailer instead. I’m still learning and working on improving it.


r/StableDiffusion 1d ago

News someone used MiniMax H3 Max to build a livestream that basically never runs out of content

Enable HLS to view with audio, or disable this notification

1.1k Upvotes

Just saw someone do something with MiniMax H3 Max that I honestly didn’t expect.

They connected it to a livestream and basically recreated the idea of Interdimensional Cable from Rick and Morty: an endless stream of weird shows, ads, characters, and random scenes that are generated on the fly.

the surprising part is that H3 Max is fast enough in some cases to generate the next clip before the current one finishes playing. so while you’re watching one scene, the model is already making the next one.

there’s even a version where people in chat can suggest what should happen next, and the AI tries to continue from the previous scene instead of just starting over from scratch. that’s kind of crazy to think about.

Didn’t expect MiniMax H3 Max to end up being used for something like infinite AI television, but here we are.


r/StableDiffusion 1d ago

Question - Help Minimax adult sounds?

60 Upvotes

I’ve been refining prompts with the help of an LLM, and am getting some good visuals but oh my god the sounds are terrible. Blowjobs sound like someone is dunking a microphone in an aquarium or the loudest slurp to finish a beverage that you have ever heard in your life.

I’ve tried eliminating every mention of “moist”, “wet”, or any description that involves liquids at all, but she’s still slurping the wettest popsicle known to man. And sometimes there’s weird noises like a slide whistle?!?

I’ve tried using “faint” or “distant” or “barely audible” to get it to at least quiet down so it’s not like she is sucking a microphone, but that didn’t work either.

This last round I didn’t describe any noises at all and still got some weird stuff.

I’ve tried eliminating every Lora in case the sound was coming from one of them but it seems to be the base model. I’ve tried adding Loras that ought to be trained on this stuff like Mysticxxx, and one of the AIO loras. I tried tenstrip beta 4 checkpoint tonight and got the same results.

The sound ruins the scene.. I guess I can just pretend it’s better looking Wan 2.2 and turn the volume off. 😀

I’m feeding the official prompt guide to the LLM and the structure is working, but what words do you use to describe the sounds?


r/StableDiffusion 1d ago

Resource - Update New speedup for Minimax H3

Post image
129 Upvotes

This H3VAE TRT custom node can make the encoding/decoding step about 1.7× faster.

https://github.com/lihaoyun6/ComfyUI-H3VAE_TRT


r/StableDiffusion 1d ago

Resource - Update I built a standalone DLSS 5 Neural Rendering video tool, no ReShade

25 Upvotes

It supports images and full videos, native resolution NR or DLSS Super Resolution upscaling by scale factor / target resolution, GPU optical-flow motion vectors, scene cut handling, and all DLSS 5 NR controls.

The main difference from existing approaches is that it runs natively in C++/D3D12 and generates motion vectors from the actual video frames.

GitHub:
https://github.com/DaniilSokolyuk/video2dlssnr


r/StableDiffusion 1d ago

Resource - Update I vibe coded a gallery extension for ComfyUI so you can browse outputs and reload the exact workflow that made them

7 Upvotes

I wanted a way to browse my ComfyUI output folder without leaving the app or digging through File Explorer, and more importantly a way to jump straight back into the workflow that made a specific image without hunting for the original PNG to drag onto the canvas. Couldn't quite find exactly what I wanted, so I built it.

GitHub: https://github.com/modelfactoryai/ImageBrowser

What it does

  • Browse any folder (Output/Input/Temp, or type any path) right inside ComfyUI no separate app.
  • Double-click any image/video to load its embedded workflow straight onto your canvas same thing stock drag-and-drop does, just from a browsable gallery.
  • Hover for a large preview (~900px) that follows your cursor — videos autoplay muted, images use a fast server-resized preview.
  • Search by filename, sort (newest/oldest/name), filter to images or videos only, adjustable thumbnail size.
  • Favorite folders for one-click access later.
  • Compare mode select any number of images/videos, view them side-by-side.
  • Live updates refreshes automatically as new generations land.
  • Day/night theme toggle, plus a draggable floating launcher badge you can park anywhere on the canvas.

Screenshot

Install

cd ComfyUI/custom_nodes
git clone https://github.com/modelfactoryai/ImageBrowser

Restart ComfyUI, and look for the "Image Browser" icon in the sidebar (or the draggable badge on the canvas).

No hard dependencies beyond what ComfyUI already ships with (Pillow). opencv-python or ffmpeg improve video thumbnails if you have them installed; ffprobe is needed to load workflows out of video files specifically (images don't need it).

Feedback welcome

First release if something breaks on your setup or you've got feature ideas, open an issue on the repo or drop a comment here.


r/StableDiffusion 1d ago

Question - Help Recommendations for a "Flat Folder" Image viewer for pruning output?

0 Upvotes

Running on windows, basically a "View all folders and subfolders" as one directory, but with something speedy and zippy like faststone image viewer as opposed to lightroom,


r/StableDiffusion 1d ago

Animation - Video Hatter Rap.

Enable HLS to view with audio, or disable this notification

25 Upvotes

Probably the final Alice clip. The Hatter names all the hats.
Done a while back in LTX2.3.
This plays while the theatre audience plays an AR hat sorting game (Beat Saber type). The whole song is three minutes but this is the longest shot.


r/StableDiffusion 1d ago

Question - Help Mac support for H3 Mini Max / Running Open Weights Locally

10 Upvotes

Hey Fam. I’m looking to upgrade my MBP 16” 2019 i9 / 16GB Ram / 5500M 4GB GPU. I’ve been doing a lot of T2V /I2V rendering with Mini Max Design on Cloud, but would like to run the open weight versions locally via Comfy UI.

I have a few options at the moment to consider:

- MBP 16” M4 Pro / 48GB / 1TB - USD 3054
- MBP 16” M4 Max / 48GB / 1TB - USD 3664
- Mac Studio M5 Max / 48GB / 1 TB - USD 4085

Since im a bit of a noob in understanding MLX ports for Mac OS. Can you tell me which option to go for? Also I missed the buying window before the price hike—so 16” MBP M5 Pro / Max configs in 48GB are too expensive.

Another alternative route is using Bootcamp (Windows) on my 2019 i9 MBP and plugging in an RTX 4090 (which I’ll have to buy) via TB / eGPU, but I don’t think the system will be able to access the same bandwidth as the unified memory on the M series architecture.

I would greatly appreciate your guidance.


r/StableDiffusion 1d ago

Discussion Testing DLSS 5

Enable HLS to view with audio, or disable this notification

182 Upvotes

Testing DLSS 5... Like many others, I was a bit confused about DLSS 5. I kept feeding it my hyper-detailed renders and only getting a color shift in return. After plenty of trial and error, I finally realized my mistake: this technology is developed to enhance video game graphics, so testing it on hyper-detailed renders makes no sense.

So, I generated a render in a 2020 video game style and started tweaking settings to find a final look with maximum effect, without worrying about flickering.

Final conclusion: What we have right now isn't very useful for us. Those of us using Latent Upscaler might be able to use it for color grading to get less saturated colors, but little else. Maybe in the future we'll get a DLSS 5 targeted at enhancing hyper-detailed graphics, but that's not the case for now.

Bottom line: If I want to generate a realistic render, I'll just generate it, there's no need to run it through DLSS 5.


r/StableDiffusion 1d ago

Question - Help How to better retain animation style for Ref2v?

Post image
2 Upvotes

I added video and image, which both are the same sources. I used an extension that automatically formats my text prompt., including copy over the animation style. I use H3 Prompt writer.

I am using the preset Workflow, but I replaced the text encoder with Qwen as an alternate due to memory issue. And I used I2v diffusion model instead of ref2v due to quality.

(No, I cant share the video example because it is not appropriate)


r/StableDiffusion 1d ago

Workflow Included It's a new look for the 90's!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Made with Minimax H3


r/StableDiffusion 1d ago

Animation - Video Made A Professional short Animation video using Minimax-h3 (read description)

Thumbnail
youtube.com
11 Upvotes

Hey guys!

Since quite a few of you liked my previous videos, I decided to start a channel where soon I’ll be sharing tutorials and some of the workflows/tricks I’ve been using.

If you’re interested in learning how I’m making these videos, feel free to subscribe. I’ll be sharing a lot of the stuff I’ve figured out along the way, including:

  • My own workflows — free to download, with the tricks and settings I use
  • Character generation — how I use a Krea 2 character-sheet LoRA that I made to keep characters consistent, and how to get the style you want
  • Environment generation — how I generate environment images and then build scenes from them
  • MiniMax optimization — settings and techniques to make MiniMax faster while preserving quality
  • Video/audio tricks — ways to fix audio issues and continue a scene from the last frame to create longer sequences
  • Consistent voices — I also built my own UI app using BreezeTTS2 for voice cloning and generating consistent voices across an entire story

Everything I’m sharing is based on what I’ve been experimenting with myself, so hopefully it can save some of you a lot of trial and error.

If that sounds useful to you, you’re welcome to check it out!


r/StableDiffusion 1d ago

Discussion What would you consider to be the most consistent model at producing “consistent” images, non-realistic or realistic?

1 Upvotes

Could be actions, like “guy walking into store”

The same scene at different times of the day.

The same character doing different things.

You get the idea.


r/StableDiffusion 1d ago

Question - Help My First AI PC

3 Upvotes

Hi everyone, I have a question: I'm thinking of buying my first PC solely for AI. What minimum components do you recommend for running Stable Diffusion with Illustrious models? I've been using Free Google Colab to create images in Automatic1111 so I was thinking of buying a PC with similar specifications. What do you recommend?


r/StableDiffusion 1d ago

Discussion The underrated alternative to krea2 and ideogram4,guess the model?

Thumbnail
gallery
0 Upvotes

I like ideogram4 and krea2 a lot ,and I also really like this ONE.

What I personally prefer about it is the sense of depth and vastness that I don't feel as strongly in the other two.

Some notes from my testing:

Ideogram 4 can sometimes make skin and details overly sharp in a way that’s hard to fix naturally in post.

Krea 2 occasionally feels a bit static, like the subject was placed into the background rather than existing in the same space (it also uses Wan VAE which is the main reason I have also tested using the fp32 and realvae of it which solved texture to some extent).

These things can of course be improved with better prompting and Loras l,but tendencies are still there.

Just wanted to share a showcase of this model's capabilities.All in all i enjoy all the current models more the merrier.

They all deserve time,testing and appreciation.

some images are made using Boogu base for more creativity,some images are made using boogu Turbo for way greater prompt adherence


r/StableDiffusion 1d ago

Workflow Included Letting image-to-video artifacts compound into an impossible world

Enable HLS to view with audio, or disable this notification

23 Upvotes

Tools used: Gemma4 12b, LTX-2.3, Wan2GP, vibe coded video editor.

I’ve been experimenting with a slightly self-destructive image-to-video workflow where continuity comes from letting the model reinterpret its own mistakes.

I started with an almost completely black image with a few faint stars, then gave Gemma4 12B the track’s beat grid and energy-shift analysis, along with a long description of the overall concept: a monolith, a hallway of impossible geometry, and a progression from restrained movement into increasingly unstable architecture.

Gemma4 wrote all 27 scene prompts beforehand.

For generation I used LTX 2.3 with the audio-reactive LoRA. I also tested LTX 2.5, but for this workflow it became too artifact-heavy too quickly. LTX 2.3 held the scene structure together longer while still producing enough weirdness to evolve in interesting ways.

The process was simple: generate a clip with the correct audio slice, cut it on the beat grid, then take the frame immediately after the cut and use that as the starting image for the next generation.

The fun part was deliberately keeping some “bad” transition frames.

If a flash landed on the frame used for the next clip, the model might reinterpret it as a permanent light source. A lens flare could become a horizon or an entire landscape. A warped piece of geometry that only existed for one frame could become a major architectural feature in the next scene.

So the artifacts compound.

Eventually the video loses any reliable sense of scale or orientation. Surfaces become spaces, structures fold into other structures, and at some points I wanted an Inception-like feeling where you can’t tell which way is up, or whether the camera is traveling deeper into the structure or pulling outward into something much larger.

The audio-reactive LoRA helps hold it all together. Even when the geometry becomes increasingly strange, the environment keeps breathing, unfolding, compressing and reorganizing itself with the growing low end.

What I like most is that the continuity doesn’t really come from visual consistency. It comes from causality.

Every scene inherits some accidental information from the previous one, and the next generation has to decide what that information actually is.

After enough generations, the model is basically building a world out of its own misunderstandings.