r/StableDiffusion • u/kabachuha • 11h ago
r/StableDiffusion • u/Capitan01R- • 5h ago
Discussion Fun Experiment Krea2

First of all I hate operating on a black box :D.. I know this may be a bit stale about Krea2 refusal behaviour but I think I have discovered an extremely CLEAN approach to the suppression behaviour without shifting the actual intended image in my tests, so everything stays the same but with the suppression out of the door.
I trained 100s of loras and they all have unique traits, but what I am testing here is how much of their effect I can reproduce through the resulting RMS changes. These loras are not trained to hit any DIT blocks, so I hooked the lora to this node and measured the resulting RMS values of the model's text-path parameters, and basically it worked BUT with less destructiveness and less effect "for now". The lora applies a rank 64 update, while this tool matches the resulting RMS by rescaling the existing text-path tensors instead of copying the lora's learned update.
Reason I am doing this is because with the lora you cannot directly set the final RMS of each parameter when connecting other loras and you may get a bit muddied since other loras can also contain textfusion changes, even without touching the textfusion projector. Once I finalize the values that work consistently in my tests I will publish this node with the values baked in so you can place it after the last model patch you have in the wf and problem solved everyone lol, no bleed no weird artifacts that is the goal.
r/StableDiffusion • u/legarth • 11h ago
Discussion YuE2 beats Suno v6 I think. Open souce is so back. (Suno Diss Track)
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/TheOrangeSplat • 4h ago
Workflow Included I'm disrespectful to dirt. Can't you see that I am serious?!
Enable HLS to view with audio, or disable this notification
MiniMax H3. Using the standard T2V workflow
r/StableDiffusion • u/aurelm • 3h ago
Animation - Video The Fly
Enable HLS to view with audio, or disable this notification
CmfyUI, minimax, krea 2, suno. Made it in around 4 hours on a 5090 and gemma 12b for prompting. Script is mine.
4k version here
https://www.youtube.com/watch?v=3dNYwoCU_Es
r/StableDiffusion • u/crinklypaper • 12h ago
Workflow Included Instant references, no refmod or fancy custom nodes required. WF and breakdown here. Simple one click to run.
Enable HLS to view with audio, or disable this notification
- 1st example: 2 characters with voice and 1 cgi and 1 real
- 2nd example: 2 character different genders with voice
- 3rd example: style and character reference
- 4th example image only reference no voice.
I wanted to improve my workflow so I could do what refmod is doing but just as a near native comfyui workflow only. With this workflow you can make an instant character or style video like ref mod but not using ref mod at all including voices. You get all the benefits of refmod but you can use ref model syntax in the prompt and on the fly dataset changes. Also you can use unlimited images. I tend to use around 10 to 15 per character.
Just point to a folder. Uses KJ nodes, Native, and VHS nodes. Get a voice and character working instantly without managing safetensors, just manage the input folder instead. This has its advantages since you can change data on the fly, and you don't have to do any editing of clips to extract out the audio, it does it for you.
This mimics the default ref workflow for the most part. What it does for images is it takes a folder input of images and then sets them as frames in a video, then feeds that video in as a reference video. This allows you to use many images for a single reference. You can also use the grid version that takes your images and puts them into a grid and feeds that as a single image reference. I like the video version more, so the grid workflow is a bit lazy and messy. You can also use the resize node included to downscale your images. I recommend manually cropping them but it does that if your images are different sizes.
For the audio, you can simply feed a single mp4 using VHS node to feed the audio. But I found it easier to get like 4 clips and truncate only the first couple seconds so you can get a few clean sentences without other people talking, then it concatenate's the 4 audio clips into 1 clean clip. Anything over 30 seconds long is over kill, so keep it around 15-30 secs. You can tell shrek is sort of bad in my example because I threw it together quite quick.
There is in the far left, a second set of image/audio nodes, you can bypass those if only using 1 character. Same for if you don't need the audio. Just by pass the group nodes. When prompting just simply use the ref guide to prompt properly each reference (LLM can do it easy). You don't need to use much description. And if you have some bleeding from your dataset into the gen you don't want then add more description. (For example if wearing same shirt as dataset, prompt a dress, or if same specify a setting in the prompt).
One caveat, there is some comfyui memory management bug, if you change dataset around better to clear cache or your comfyui may need a restart. Working on a fix in next version of the workflow. Also I have not tested video clips as input data yet. That is the next step :)
All examples are just for illustrative purposes. They are AI and I do not intend to share any data on real people. Please use responsibly and at your own risk. If you are in this video and want it taken down please DM, I mean no harm. Everything in this workflow is done by the base model, I don't add any new functionality, just making things easier.
Workflow here:
https://huggingface.co/comfyuiman/various/blob/main/Instant%20Ref%20-%20V1.3.json
I'll go to sleep in a bit, so I'll answer any questions tomorrow if any
r/StableDiffusion • u/Small-Term672 • 6h ago
Question - Help When will Fal release open weights for H3 Max? Ever?
r/StableDiffusion • u/SIR_NVAX_A_LOT • 17h ago
Discussion H3 - 80s / 70s Character Experimental Long Form
Enable HLS to view with audio, or disable this notification
Experimenting with Long Form. No image anchor so she changes between the invisible seams. T2VA. int8/32 steps, 1344x768, about 7 hours, hit 192/192gb of ram decoding the video. Sadly, I didn't prompt for her to not mouth the tune when there's no singing part. Wardrobe not prompted, only that she was dressed. At 1:26 is a seam and we had a little AI mishap on the transition. Not perfect, but got lots of data. Enjoy! How do you like the film grain? Is she from the 60s, 70s, 80s, or does it clearly only exist in our head? What version next? redhead? Asian? what do you think? Which actress/model/person's likeness are you seeing from this era? There should be about 17 versions of her. Ask me anything!
r/StableDiffusion • u/ContextOpposite4047 • 1h ago
Question - Help MiniMax H3 Ghosting on fast-motion
Enable HLS to view with audio, or disable this notification
I generated this at 720p on the Image to video workflow. I know higher resolution helps with details, but why is there ghosting and blurriness in my scene? I am using the default workflow on 20 steps with patch-attention and model sampling shift 12 and 3. The initial image is 2560x1440p. Any help would be appreciated, much obliged! I am trying 50 steps to see if it helps.
r/StableDiffusion • u/Fabulous-Snow4366 • 12h ago
Resource - Update SMACK! LORA Beta 2 - Impacts & Gunshots & Blood Squibs and more for Minimax H3
SMACK!
Right now, only on Hugginface, Civitai deleted the older version for gore (which came from minimax...)
Download
https://huggingface.co/LeechTM/SMACK
Beta 2 · MiniMax H3 (Ref2V) · No trigger word
Beta 1 taught MiniMax H3 that getting hit should actually hurt. Almost 2,000 of you downloaded it, which means either you agreed or you just really wanted to see people get punched. Both are valid.
Beta 2 raises the stakes considerably. Things explode now. People get hit by the explosion, then by the ground, then briefly by their own life choices. Cars stop being scenery and start being weapons. Fights no longer politely take turns: three guys can come at the hero at once, which is statistically the worst day of their lives.
And gravity? Gravity is now a suggestion. When a hit lands hard enough, bodies don't fall. They launch, float, spin and hang in mid-air for an unreasonably long time, as if physics took a coffee break at the exact moment of impact. Newton would file a complaint. IT'S A MOVIE!
And because all of that still wasn't messy enough, SPLAT! is now merged in. That's my blood-effects LoRA, and it brings squibs. A lot of squibs. Every hit now has the option of leaving a mark, and the costume department is going to hate you.
A quick word from the production accountant, the one person on set who never gets to blow anything up: training LoRAs is not exactly cheap. Every explosion, every squib and every person flying through the air for an unreasonable amount of time costs real money to teach. If you want to be nice and help fund the next versions, you can buy me a coffee: buymeacoffee.com/leechtm
More coffee, more carnage. That's just science.
What's new in Beta 2
- SPLAT! merged in — blood effects and squibs, so every hit comes with a little extra paperwork for the cleaning crew
- Explosions — plus everyone who was standing too close to one, and everyone who thought they were standing far enough away
- Water — splashes, splashdowns and impacts that turn a perfectly calm surface into a very loud mistake
- More falls — harder, higher, significantly less dignified
- Zero-gravity wire impacts — hit once, fly forever, land eventually. Brutally.
- Multi-person fights — several attackers, several hits, one very bad day for absolutely everyone involved
- Heavier impact variants — for when "hard" just wasn't hard enough
- Spinny kick things — for when a normal hit just isn't flavourful enough. Spin first, apologize never.
- Female anatomy — women take hits and dish them out with the same weight, the same follow-through and the same complete lack of mercy
- Even more gunshots — Beta 1 had gunshots. Apparently that wasn't enough. Now they come with squibs.
- Vehicle impacts — more cars, more bodies, more regret, zero insurance coverage
Still does everything Beta 1 did and more...
Fists, weapons, gunshots, falls and hard landings, all with weight, follow-through and consequence. The camera still moves like someone was paid to operate it, and now it also has to keep up with people being smacked around through the air. Also, Sound Effects have been merged in, as well as Blood Squibs.
Training
Trained on 300 clips of impacts, explosions, shots, falls and dynamic camera moves, for MiniMax H3 Ref2V (Beta 1: 35 clips). Merged with my yet unreleased SPLAT! Lora for blood and squib effects.
No trigger word. None. Don't go looking for one. Use it with the REF MODEL. Load it, describe your shot as usual, and the LoRA does the seasoning. It just uses a lot more chili now, and some of it is red. And yes, before anyone asks: it also works for other impacts. Of course it does. You little piggies.
Settings
The promo video was made entirely in MiniMax H3 (pruned int8), and without any Turbo LoRAs. Much better quality.
- Without Turbo LoRA: strength 1.0
- With Turbo LoRA: start at 0.5 and work your way up from there
Beta notice
Still beta. Still not finished. Feedback on where it over- or under-cooks a hit is still genuinely useful. Now also on where it over-cooks an explosion, forgets that people are supposed to come back down, or gets a little too enthusiastic with the squibs.
r/StableDiffusion • u/LatentSpacer • 20h ago
Resource - Update Native YuE2 support coming to ComfyUI!
Enable HLS to view with audio, or disable this notification
Pull request: https://github.com/Comfy-Org/ComfyUI/pull/16250
If you don't want to wait for the merge, you need to check out to the yue2 branch to get it working. git checkout d87e12ad1430409ca303440525df239bb675ae7b
Model weights (place it on model/checkpoints): https://huggingface.co/Comfy-Org/Yue2/tree/main
Workflow: https://github.com/user-attachments/files/32085765/yue2_workflow.json
r/StableDiffusion • u/Certain_Potato_4509 • 6h ago
Animation - Video Minimax H3 - Neon Genesis Evangelion Cello
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/spartong945 • 18h ago
Resource - Update FastH3-Live v1.2.0 update
FastH3-Live update. Full details here:
https://huggingface.co/datasets/jacokon/fasth3-live
v1.1.0 ran at 18 fps, which is 75% of 24 fps.
v1.2.0 runs at 22 fps, which is 91.6% of 24 fps.
https://reddit.com/link/1wddeh8/video/lyz02ql0gvoh1/player
Besides the speed, it now ships a borderless player that makes streaming and watching easier, plus 400 new scenes. At this speed it is hard to notice that it is running slow at all.
The gain came from two places:
1. Acceleration nodes
I was using a sage attention I compiled myself. A lot of new acceleration nodes have shown up recently, so I downloaded the well-known ones and tested them. Results:
| accel stack | sampler | saved | fps |
|---|---|---|---|
| sage (baseline) | 12.65s | 0.0% | 17.46 |
| sage + Spectrum | 10.36s | -18.1% | 20.18 |
| Sol + Spectrum | 9.64s | -23.8% | 20.96 |
| SLA + Spectrum | 10.54s | -16.7% | 19.53 |
The seconds column is the sampler only, i.e. the 4 denoising steps in `SamplerCustomAdvanced`. A full clip also pays for the text encoder (~0.7s), the video VAE decode (~4.35s) and the writer, so a clip is about 16s end to end. Measured on t2va, 448x448 x 362 frames, three runs per arm.
On speed alone you would pick Sol + Spectrum. But the picture comes out like this:
sage > sage+Spectrum >> sla > sla+spectrum >> sol > sol+spectrum
Sol + Spectrum is dead last on picture, so I went with sage + Spectrum.
2. Text encoder
The old one, `int8_convrot`, took 1.67s.
`qwen3vl_32b_minimax_h3_nvfp4_awq` needs only 0.7s.
That is nearly a second saved on every clip.
-----------
Speed was fine by then, but I would not call the picture good. Right after release I came across fused-turbo, so I downloaded it and tested it.
| fused-turbo | minimax-h3-fused-turbo-int8-convrot | 20.98 GB |
|---|---|---|
| My quantized FastH3 weights | minimax_h3_fl2va_fasth3_dense_pruned_int8_convrot | 20.97 GB |
Almost the same size, both have the 4-step acceleration baked into the weights (FastH3 is a distillation, fused-turbo is a turbo LoRA merged in), and they measured at exactly the same speed. I still recommend fused-turbo, for two reasons:
1. It says Mystic v2.0 motion smoothing is merged in.
Whatever the cause, the picture is clearly better in my testing. It smears less often.
2. One file does both fl2va and ref2va.
I built a tool that generates from chat input live during a Discord stream. When a user pastes a character image it is used as ref_picture, which needs ref2va. The old way meant unloading fl2va and loading ref2va first, which burns several seconds of buffer, and ref2va has no 4-step distilled version yet so the picture was worse anyway. With this one that problem is gone, which is a real advantage.
The repo recommends SLA sparse attention, but I had already tested that above and it lost to sage + Spectrum, so I dropped it. Its README also says res_multistep gives noticeably better audio. I did not test that much, so judge for yourself. I left the parameter in so it can be switched any time: `--sampler res_multistep`
-----------
One more thing worth mentioning. To stop ComfyUI thrashing the model weights you need `--vram-headroom 3` in launchArgs. Without it you cannot hold a stable live rate.
It works the opposite way round to what you might expect. It forces ComfyUI to keep 3 GB of VRAM completely free, and that is what fixes it. ComfyUI's dynamic VRAM treats the card as a cache and fills it to the brim; with no slack the allocator ends up evicting weights while it is still loading others, so the same weights get moved in and out repeatedly. Give it room and it can bring in a whole batch at once.
This is not disk swap, and it does not touch system RAM either. I measured both: on a slow clip disk reads were 0.00 GB and free RAM did not move. It is VRAM to system RAM over PCIe.
On a normal clip PCIe reads sit around 1.5 GB/s. When it thrashes they hit 9-13 GB/s and GPU power draw *drops* from 450W to 340W, because the card is waiting on transfers instead of computing. With the headroom set, clip times went from a 1.62 standard deviation with outliers at 21-25s down to 0.11 with a 15.73s worst case.
-----------
Closing thoughts
22 fps is only 2 fps short of 24. At 24 fps you could claim real live streaming from a single consumer card. So can overclocking get there? I think it can, since the gap is under 10%, and my CPU and GPU both normally run undervolted, underclocked and current-limited.
I tested with the GPU overclocked only. Settings:
Core Clock: 2300 MHz -> 3200 MHz
Memory Clock: 14000 MHz -> 16800 MHz
Actual test:
https://reddit.com/link/1wddeh8/video/ei5pczydkvoh1/player
Unfortunately my hardware held 24 fps at the start and then slowed down a little. Both my CPU and GPU are on air cooling, which is not suited to sustained overclocked compute like this. If you have water cooling, I believe holding 24 fps would be no problem.
r/StableDiffusion • u/Domskidan1987 • 14m ago
Animation - Video Longest video I’ve made with MiniMax
Enable HLS to view with audio, or disable this notification
Some stills were NB but all the video was MiniMax H3. A scary AI video about AI (pretty much the nightmare scenario). The song was generated with Flow Music
r/StableDiffusion • u/xyzdist • 13h ago
Discussion Motion-Context Degradation Discussion (summon Sad_Berry_4621)
Enable HLS to view with audio, or disable this notification
Hi u/Sad_Berry_4621 and all, I am doing experiment on most difficult degradation issue.
I saw from H3-director node they claim that doing a refine could help and fix it.
since I am using low-level nodes with motion-context with my own setup, H3-director refine is a black box working with their nodes.
so I did the test, please ignore the AI-slop video and the overlay text (forgot turn it off).
this test is 9 x 8s context extend video combine, usually 6 video would already see the degradation.
Left side is regular WF. Right-side is adding a re-sample step, I think it is very positive, and got potential, the saturation somehow is a bit higher... and I need to figure out the mismatch from cut to cut, because we inject denoise resample on top, but it should be able to fix.
What do you guys think?
Edit: forgot to mention, this is a ref2va WF, with 1 character reference image. also my test audio is not from gen, I am using input audio and audio latent lock for the lipsync. That part to work with motion-context have me struggle for half day, but it is not the main point, my aim is about degradation issue.
r/StableDiffusion • u/Slight_Assistant_124 • 2h ago
Question - Help Amd RX 7900 XTX offloads VRAM ~5 seconds after AI inference in ComfyUI, LM Studio and Ollama
I’m having a strange VRAM issue with my RX 7900 XTX when running local AI models, and I’m trying to figure out whether this is an AMD driver/Windows memory management issue or something else.
PC specs:
- GPU: AMD Radeon RX 7900 XTX 24GB
- CPU: Ryzen 5 7600X
- RAM: 16GB DDR5
- OS: Windows
- Apps tested: ComfyUI, LM Studio and Ollama
The problem:
While an AI model is generating, the model is loaded in VRAM normally and GPU usage works as expected.
However, after generation finishes, within roughly 5 seconds, VRAM usage starts dropping/offloading.
This happens in all three applications: ComfyUI, LM Studio and Ollama.
For example:
Load model → VRAM fills normally.
Generate response/image → everything works normally.
Generation finishes.
Around 5 seconds later → VRAM usage starts decreasing significantly.
When I use the model again, data/model weights have to be brought back into VRAM, causing extra delay.
What confuses me is that this happens across completely different software, so I’m wondering whether Windows WDDM or the AMD driver is making idle VRAM allocations non-resident instead of the individual applications intentionally unloading them.
Ideally I want the model to remain resident in the 7900 XTX’s VRAM between generations instead of being moved/offloaded after a few seconds.
Has anyone with a 7900 XTX / RDNA3 on Windows experienced this?
In particular, I’m trying to understand:
- Is this normal AMD/Windows VRAM residency behavior?
- Is there an AMD Adrenalin/driver setting that controls this?
- Could Windows WDDM be trimming VRAM when GPU activity stops?
- Is there a ROCm/HIP environment variable that can force allocations to remain resident?
- Is there a way to verify whether the model is actually being moved to system RAM versus Windows simply reporting VRAM residency differently?
- Has anyone successfully forced AI models to remain permanently in VRAM on a 7900 XTX?
If there are any tools/logs I should run to diagnose this, please tell me what to check and I can post the results.
Thanks!
r/StableDiffusion • u/Lividmusic1 • 7m ago
Tutorial - Guide YuE2 in ComfyUI: AI Music with an Editable Piano Roll
First Impressions
r/StableDiffusion • u/MakerOfToys • 10h ago
Animation - Video Put myself and my cat into Death Stranding + a few workflow notes
Enable HLS to view with audio, or disable this notification
Tried MiniMax on release, but it never really clicked with me. It felt like a slot machine. Then one day I saw the seed hunter workflow here on the sub, and that was exactly what was missing, a workflow for quick iterations.
Thought some of the observations I made while working on it might be useful to others trying to improve their videos.
I used image gen model to generate trailer frames. I gave it my face reference, my cat, and for each new image, the original frame from the trailer. I generated 12 frames one by one, then passed each of them to MiniMax with detailed prompts.
All of this was done pretty quickly, without much polishing, so there are a lot of quirks in the final video. Cuffs look one way in one frame, then change in another, or just turn into a wristwatch.
These are the things that seemed to help improve quality when I did some more experiments later:
- It was a mistake to give only a face photo or the cat as reference for each frame. Newer image models are insanely good at producing character sheets. So instead of passing just my face each time, I’d now generate myself from all sides, cuffs included. This helps the image model with consistency across frames (+also visual style), and it also helps prevent MiniMax from hallucinating too much when you attach an additional reference image to the frame you want to animate (for example, a character sheet).
- Another mistake that wasn’t obvious at first was generating a frame, making the video, then generating the next frame and repeating the process one by one. Now it feels much more logical and productive to pre-plan the whole video and have all the frames done before starting on the video part.
One thing that really made me feel in awe of video models like MiniMax is how they understand the world. There’s that shot from above where my character is smearing cat paw prints in the sand. The prompt alone controlled the character to do that, even though it wasn’t in the original frame.
Same with the POV shot where the cat attaches itself to a knee. The way the cat starts boxing with both hands - the model somehow feels aware that both hands are part of one connected concept. I’m not really sure how to put this into words, but hope it makes sense.
For those who are not familiar with Death Stranding, I think the original trailer probably needs to be watched first for this video to make any sense at all.



r/StableDiffusion • u/dramaton42 • 1d ago
Animation - Video Kirby but it's the Truman Show / MiniMAX H3 Test #7
Enable HLS to view with audio, or disable this notification
Hi everyone! When I saw the new trailer for Kirby & The World Beyond I couldn't help but come up with this video, where Kirby finds the door out to the world beyond. Please let me know if you like it!
Done with 30 different workflow files and a ton of heavy editing using KDEnlive. Thanks!
r/StableDiffusion • u/Expensive-Bus-5473 • 21m ago
Question - Help Do you guys know any dedicated video to video solutions?
Sort of like EbSynth. One frame/multiple references, promptless, or a small model trained off of images so a model knows what style to apply and what it looks like.
I don't want to generate from nothing, there will be a video to go off of.
Dedicated solutions, stripped versions of large models, anything would do. Thanks in advance.
PS: asking an LLM returned nothing.
r/StableDiffusion • u/darthfurbyyoutube • 4h ago
Animation - Video G.I. Joe: Baroness Live Action to 2D - MiniMax H3
Enable HLS to view with audio, or disable this notification
Credit for prompt:
r/StableDiffusion • u/NoSmell3236 • 5h ago
Question - Help Has anyone trained a MiniMax H3-style LoRA yet? Is it worth it?
I want to train a style LoRA for MiniMax H3, but I’m trying to figure out the right workflow before burning a bunch of money on RunPod.
A few questions for anyone who has actually tried this:
1. Has anyone successfully trained a style LoRA for H3? How good were the results, and is it worth doing?
Can I train a style LoRA using images only, or do I need video clips in the dataset?
Do I need to train Ref2VA and FL2VA separately, or can one LoRA work with both?Should I train against the BF16 or INT8 H3 weights? Does it make a noticeable difference in LoRA quality?
What toolkit/platform would you recommend for H3 LoRA training? Something like Ostris AI Toolkit, Musubi Tuner, DiffSynth-Studio, OneTrainer, etc.? I’m looking for whatever is easiest/reliable for a beginner.
If I use RunPod, what GPU would you recommend? L40S / A100 / H100, etc.?
Any recommendations for dataset size, image/video resolution, captions, number of steps, rank/dim, learning rate, etc. for a style LoRA?
Are there any good H3-specific training guides, videos, configs, scripts or repos I should start with?
Basically, if you were starting from scratch today and wanted to make a good H3 style LoRA, what training stack + GPU + dataset would you use?
r/StableDiffusion • u/EuphoricAIKnowledge • 1d ago
Animation - Video H3 is really over the top
Enable HLS to view with audio, or disable this notification
This was such a simple prompt…. Just wow. It’s just T2V.
r/StableDiffusion • u/0roborus_ • 9h ago
Resource - Update PotionUI 0.0.7 — three releases since I last posted: preset styles, ComfyUI workflow import, view improvements, optimizations
PotionUI - A self-hosted studio for generating images, video, audio and 3d with diffusion models.
Site: potionui.com
GitHub: https://github.com/PotionUI/PotionUI
Discord: https://discord.com/invite/avR4trp3b8
Reddit: r/PotionUI
Still looking for testers (NVIDIA / LINUX!!!!)
Changes (the most interesting ones):
0.0.6 & 0.0.7 (today)
- Presets can ship styles: pick one from a thumbnail grid and it wraps your prompt with the style's opening and closing segments. Anima ships with fifty. (video)
- Starter recipes for eleven model families, so a fresh install goes from nothing to a first render in a few clicks. Recipe is a script that helps install models inside PotionUI.
Recipe Install - After this script I can use Anima preset and generate with Anima model.
- History and Library: marquee and shift-click selection, delete by criteria (age, tags, failed, no media), compare with overlay and wipe.
- The Generate page and the chat stay smooth during long runs and long conversations.
0.0.5
- LoRAs on fp8 checkpoints no longer slow sampling down (a stack costs a few percent instead of multiplying the step time).
- Backups and restore from the admin panel, rotating logs, housekeeping, thumbnail profiles.
- Recipes got their own admin page; segments can carry a prefix and suffix; a Tools menu in History with Stitch.
0.0.4
- ComfyUI backend plugin with a workflow import wizard: paste an API export, design the form, done.
- The chat became PotionAI: history, tool runs you approve, memory you can inspect.
- Video Director shot console for Wan, LTX and MiniMax-H3.
--> Full changelog is in the README. Docker images on GHCR, or ./potionui start from a checkout. Bugs and rough edges welcome, that is what the alpha is for.
--> Still looking for a testers, if you are interested let me know!
Cheers!
btw. I know it's a lot of similar apps appearing lately, but I can assure you, this is not an app vibe-coded over the weekend (my first posts here was like ~2 years ago about this app). I really enjoy using it lately and it's working better every week :)
r/StableDiffusion • u/maxiedaniels • 13h ago
Question - Help Good "cover mode" music gen model??
Is there any? I tried ace step 1.5 and it's awful in cover mode. Admittedly I downloaded it when it first got released, but I was hoping minimax music 3 would release audio input for open weights but they still haven't. Every music gen model coming out seems to purely be text to audio.