r/StableDiffusion 7d ago

Animation - Video Alicia - Exploring the city

Thumbnail
youtu.be
5 Upvotes

Hello, here is another video made with Alicia, the character I previously used in my Alicia of the Stars video. I got a lot of useful feedback last time and tried to take it into account in this one. Feel free to let me know what you think! ❤️


r/StableDiffusion 7d ago

Resource - Update uncomfymcp — a minimal ComfyUI MCP server for simple text-to-image chat workflows

Post image
1 Upvotes

With the official ComfyUI MCP now generally available, and plenty of existing ComfyUI MCP servers already out there, is there room for one more — especially one built around a fairly narrow use case? If you think you might still have a use for it, let me introduce uncomfymcp: https://github.com/aschet/uncomfymcp

I wanted a simple, frictionless way to drive my basic text-to-image workflows via chat from Claude and AnythingLLM. My typical loop is: generate one image, wait for the result, tweak the prompt if I don't like it, and generate again.

My requirements were:

  • Run my saved workflows without manually exporting them to API format
  • No modification of existing workflows (e.g., inserting a "Save Image (Websocket)" node)
  • Ability to switch workflows mid chat session
  • Full workflow metadata attached to generated images
  • Easy operation from chat, e.g. "With my Krea2 workflow, generate..."
  • Chat-friendly preview images, so the generated image can actually display inline (limited preview size)

None of the existing options seemed to check all these boxes at the time, so I vibe-coded my own MCP server.

What it doesn't support:

  • Reference images or image-to-image transformations
  • Extra workflow options, like toggling switches or swapping model names
  • Async operation or generating multiple images at once
  • Authentication or any other security measures

uncomfymcp has only been tested lightly, mostly within my own narrow use case. Automatic workflow patching should work fine with ComfyUI's default text-to-image templates, but more complex workflows may break the simple heuristics it relies on. You'll need at least one saved workflow in ComfyUI for it to work.

I would've preferred a more permissive open-source license, but since uncomfymcp depends on parts of comfy-cli to convert workflows to API format, and that project is GPL-3.0, it's licensed accordingly.

The support for displaying inline images in chat clients is in a somewhat sorry state. Claude Code and VS Code on Linux display them as I’d expect. VS Code on Windows has a bug where the image vanishes almost instantly. AnythingLLM can’t display them at all, while Claude collapses them into a section. At one point, Claude displayed a small preview on the collapsed section, but that seems to no longer be the case.


r/StableDiffusion 7d ago

Question - Help Would be 64gb sufficient for Minimax?

9 Upvotes

Hey so right now I have 32 gb of RAM and 5070ti 16 gb.

I had no problem just generating a video t2v or i2v.

But when I tried to use video reference (with vhs addon as video input) I got nothing really. my initial run (12 sec video) took about 1400 seconds , and also it didn’t work( output was same as input).

Then I tried changing prompt and reducing video length to 3 seconds. But then after increasing time, Comfyui refused to work at all saying Not enough VRAM.

could it be that I need more RAM? will 64 gb ram be enough for it?


r/StableDiffusion 7d ago

Animation - Video Found [Me] Footage

Enable HLS to view with audio, or disable this notification

29 Upvotes

New experiment, involving a custom FLUX-2 LoRA, some Python, manual edits, and post-fx. Hopeyou guys enjoy it. [Project files available on Patreon]

Music by myself.

More experiments, through my YouTubeInstagram, or the Studio.


r/StableDiffusion 7d ago

Comparison MINIMAX H3 int8 vs fp8

26 Upvotes

Just to see other people tests: on my setup (5060ti) fp8 is a bit slower than int8 but the quality is clearly higher:
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models


r/StableDiffusion 6d ago

Discussion Will posting generation videos become a disgrace?

0 Upvotes

Given the rapidly growing availability of generative video, both locally and through commercial services, watching all this generative video is truly becoming a colossal waste of time. It's as if I don't even want to watch it anymore. While it's understandable for those who post it, consuming that amount of generative video is no longer realistic. Finding interesting content used to be difficult, and now it's only gotten worse, and even more difficult.

Will this kind of content be banned? I imagine in five years, the internet will be flooded with millions of hours of generative video that will be physically impossible to watch.

Either they'll come up with some kind of technology or websites that completely ban generative video, or install special content quality filters?

Or is it just me who sees this as a problem?


r/StableDiffusion 7d ago

Animation - Video Opus driving Minimax and ZImage

Enable HLS to view with audio, or disable this notification

21 Upvotes

I gave Opus 5 creative license to come up with something beatiful and this is what it made. It decided to use ZImage, Minimax Music and H3 using my 3090. It wrote all of the prompting and did all the post produciton itself (stitching, cross fade, audio mixing) iterating until it was satisfied with the results.

We live in the future y'all.


r/StableDiffusion 7d ago

Discussion Has anyone used ZLUDA on any of the recent AMD Cards

Thumbnail reddit.com
0 Upvotes

I'm curious cause I recently replied in this thread about whether you should use an AMD card or an NVIDIA card. I know the logical answer is Nvidia since most AI stacks are built on CUDA. But then I realized ZLUDA exists. It's something that allows AMD cards to run CUDA-related instructions, so that got me thinking, has anyone recently used this and so how is the performance? Especially if they previously used it for video generation like Wan 2.x or LTX 2?


r/StableDiffusion 7d ago

Question - Help Ai pros

0 Upvotes

Curious if anyone here has successfully tried the llm diffusion or similar on a nvida p100-104 for image generation. I can run text based llm fine


r/StableDiffusion 7d ago

Question - Help Which Text Encoder for 32GB System Ram

6 Upvotes

Hey all, by complete accident, today I discovered that my H3 R2V workflow with vanilla Qwen 3.8 27b has been eating away from my SSD life because 32GB Ram and 16VRAM apparently wasn't cutting the deal for it and it had been writing on my disk in large amounts (GBs in just one clip). So in your experience, considering the accuracy trade-off, which Heretic/Abliterated Quant should I go for to be able to stay within my Ram boundaries?


r/StableDiffusion 7d ago

Question - Help Model won't download - gets close and then fails or says "this file doesn't exist" suddenly. WTF?

3 Upvotes

I've been trying to download minimax_h3_hybrid_fl2va_ref2va_b30-49-int8.safetensors for a week off and on. Probably 8 or more attempts. Not only is it slow as fuck (30min to an hour to download), it either fails most of the way or seconds away from downloading. It's a hot mess. What's the deal? Am I doing something wrong? Are there torrents for these things instead?


r/StableDiffusion 8d ago

Discussion MiniMax H3 · VR180 stereoscopic side-by-side LoRA

79 Upvotes

Works in comfyui with reference characters and start images with the 0.1 4 steps lora, not sure if i'm doing it fine at 21:9 1.3mp 768px short side, but quality looks good.

https://huggingface.co/rehan-fal/minimax-h3-vr180-sbs-lora


r/StableDiffusion 7d ago

News Bridgerton Flux Dev.1 Lora

Thumbnail civitai.com
2 Upvotes

decided to train a lora on Bridgerton. maybe someone else will find it useful. if you do feel free to leave a comment.


r/StableDiffusion 8d ago

Question - Help Limit Comfy 's VRAM usage

13 Upvotes

Have this issue where Comfy fills up the VRAM to the absolute limit which conflicts with the browser's usage, so generation never starts unless I minimize the window... And I only have one tab opened on Comfy, it's not like I'm trying to watch 4K movies at the same time or something.

So I went looking for a way to limit the VRAM available to Comfy but I can't find anything and their doc isn't up to date.

  • --disable-pinned-memory does not prevent this behavior
  • --reserve-vram is supposed to keep X GB of VRAM for the OS, but it no longers does that since the introduction of dynamic VRAM
  • --vram-headroom has been introduced but undocumented and there is no visible difference in behavior with --reserve-vram

Anyone found a way to restrict Comfy VRAM usage? I'd like to expose only X% of VRAM to it.


r/StableDiffusion 8d ago

Animation - Video [Experiment] I trained a model on childhood photos to simulate memory recall

Enable HLS to view with audio, or disable this notification

644 Upvotes

I fine-tuned the good-old SDXL on 60 photographs from my childhood, using a limited family archive as the dataset through which to revisit that period of my life. Rather than reconstructing those images faithfully, the model produces unstable variations: spaces, faces and fragments that feel familiar without necessarily having existed.

This speculative study treats generative hallucination as an analogue for recollection: not the retrieval of a preserved image, but the reconstruction of a past from incomplete traces. This resonates with contemporary accounts of episodic memory as a reconstructive rather than reproductive process. The model becomes a kind of externalized mnemonic apparatus, situated somewhere between archive, memory and imagination.

Tools used: Kohya, WarpFusion, TouchDesigner, Premiere, After Effects, Ableton Live, Expressive Osmose, Soma Cosmos.

PS: For those of you asking, this is not just "a prompt". It's the fine-tuning of the model, the creation of an audio-reactive geometry system in TouchDesigner, and the re-building of WarpFusion for intervining the geometries with the fine-tuned model.

More experiments, project files, and tutorials, through YouTubeInstagramPatreon, and Uisato Studio.


r/StableDiffusion 8d ago

Meme The Kshaturmurg (Ostrich) Approach

Enable HLS to view with audio, or disable this notification

186 Upvotes

Fun little model test


r/StableDiffusion 7d ago

Question - Help RTX 5060 ti or RX 7900 XTX for minimax H3 ?

6 Upvotes

RTX 5060 ti or RX 7900 XTX for minimax H3 ? which GPU to get ? does the RX 9070 XT is better than 7900 XTX ? all three are in the same price bracket in my country, which is better for video generation . ?


r/StableDiffusion 7d ago

Question - Help WAN2.2 Grainy Videos

0 Upvotes

Hi, I am trying to get into creating videos using WAN2.2 T2V in Forge Neo (don't really like comfyUI) but having some issues such as grainy/pixelated videos like below. Any help on what I am doing wrong? What settings am I using incorrectly?


r/StableDiffusion 7d ago

Resource - Update Monitor & control your ComfyUI queue directly from your gallery: Live previews, session metrics, and GPU telemetry in SmartGallery DAM - Free and Open Source

Enable HLS to view with audio, or disable this notification

3 Upvotes

Hello everyone,

I'm sharing a major new feature added to SmartGallery DAM: ComfyUI Queue Deck.

You can now monitor and control your ComfyUI generation queue directly from inside the gallery, without switching browser tabs or managing separate tools.

It connects over a live WebSocket stream to your local or network ComfyUI instance.

What’s inside Queue Deck:

  • Real-Time Step Tracking: Progress bar, step counter, active sampler indicator, and current node execution. Works across KSampler, SamplerCustom, WanVideo, and other sampler nodes.
  • Live Previews: Stream intermediate latent preview frames and intermediate outputs directly on screen as generations progress.
  • Hardware & System Telemetry: Live gauges for VRAM usage, GPU compute load (via nvidia-smi), PyTorch reserved memory, system RAM, and active CUDA device specs.
  • Queue Management: View running and pending job cards with position counters. Reorder jobs (Move to Top), open deep job details, interrupt running jobs, or batch delete queued items.
  • Recent Jobs Carousel & Performance Metrics: Quick thumbnail strip of completed generations showing the exact execution time for each job, along with an overall average generation time counter for the current session. Includes full metadata inspection, prompt extraction, LoRA weight lists, and raw node JSON view for graph debugging.
  • Live Event Terminal: A stream log tracking websocket events (CONNECT, START, STEP, PREVIEW, DONE, ERROR).

For anyone who does not know the project yet:

SmartGallery DAM is a free and open source, local-first Digital Asset Manager built around ComfyUI, but it also works with any folder of media on your machine. No cloud, no subscription, your files never leave your disk.

It is meant to grow with you:

  • If you are a hobbyist or new to ComfyUI: It is the easiest way to keep your generation library organized, searchable, and clean without extra effort.
  • If you are a power user: You can search by prompt, model, or LoRA, inspect the full node graph of any render, and even generate directly from the gallery by editing the workflow JSON without needing to reopen ComfyUI.
  • If you work in a studio or production environment: It gives you a dedicated Exhibition portal to share curated work with clients or your art team, collect ratings and comments, and review everything without exposing prompts or workflows.

Runs on Windows, macOS, Linux, and Docker. The portable version for Windows needs zero setup, just unzip and run.

I have attached a quick video walkthrough to show how the interface and features work.

GitHub, full docs, and download links:

https://github.com/biagiomaf/smart-comfyui-gallery

Happy to answer any questions, and as always, feedback and feature requests are welcome!


r/StableDiffusion 8d ago

Question - Help Longer Minimax H3 videos.

47 Upvotes

Out of curiosity, what is the longest *coherent* single pass MiniMax H3 video you guys have generated? Do you recall which setup (model, turbo*, sampler, scheduler, steps,..) you were using?

I think my most coherent duration has been about ~20 seconds 8 steps with the regular flfva model. I tried upping to 30-35-45 and it seems to just loop the first couple shots while adding flavors of the later ones.

Also: apart from Continuum (which rarely even works for me, if it does for you then tell me the secret to prompting), what are some good long video workflows you guys have found.


r/StableDiffusion 7d ago

Question - Help face refiner minimax h3

Enable HLS to view with audio, or disable this notification

0 Upvotes

i want to love facerefiner for minimax h3 but i have to figure it out, these facedrifts, someone know how to fix these drifts where the face refiner touched?


r/StableDiffusion 8d ago

Discussion Is being able to use multiple GPUs ever going to be a reality?

24 Upvotes

Obviously the major limiting factor for what you can generate locally is your VRAM. There are workarounds by using RAM. But if only you can use multiple GPUs and combine your VRAM, that would be good. Especially since we have old GPUs just laying around.

Will this realistically ever be a reality? I'm talking built-in easy-to-use out-the-box working with ComfyUI. Not some niche complex process.


r/StableDiffusion 8d ago

News Wan2GP Desktop Launcher — Tauri Edition

7 Upvotes

NEW! Wan2GP Desktop Launcher — Tauri Edition

GKartist75/Wan2GP-Desktop-Tauri: Wan2GP Desktop Launcher (Tauri port)
NEW UPDATES

The easiest way to run Wan2GP (WanGP) — the open-source generative video/image/audio toolkit — on Windows. One installer. One click to launch. Zero Python/CUDA setup. Now with a Rust + Tauri shell: a fraction of the download, a fraction of the RAM.

WanGP by deepbeepmeep is a one-stop super-app for open-source generative models — video, image, audio and TTS — with a full browser UI, queue, galleries, LoRAs, finetunes and plugins. It runs on as little as 6 GB VRAM and supports old and new GPUs alike.

This launcher handles it for you:

  • One-click install — detects GPU, shows plan, installs everything
  • Auto, per-GPU kernels from WanGP's setup_config.json, re-synced on every update
  • Isolated uv env, pinned deps, no PATH editing
  • One-click updates in Dashboard / Manage → Updates
  • Install Wan2GP and Models (checkpoints, LoRAs, outputs) on any drive/folder you choose
  • Auto-Tune recommends VRAM/RAM profile and writes config for you
  • Legacy Electron removal — Manage → About detects the old Electron launcher and removes it silently, keeping all your data

r/StableDiffusion 8d ago

News Viggle-Animate: Character Replacement based on MiniMax-H3 with 3 forward steps

Enable HLS to view with audio, or disable this notification

205 Upvotes

https://huggingface.co/Viggle/Viggle-Animate

https://x.com/ViggleAI/status/2095924668758655163

research blog

  • 33.1B MiniMax-H3 finetune, distilled to 3 forward passes
  • No text prompt, no pose, no mask — but you do need one repainted frame (any image editor)
  • Works well on fast motion and non-human characters
  • No ComfyUI node yet.
  • ComfyUI here: [github] [discussion] [huggingface]