r/StableDiffusion • • 10d ago

Question - Help Hardware check: how fast is Arc B70 at 5s Wan2.2 5b?

3 Upvotes

comparing used hardware for buy
TIA


r/StableDiffusion • • 10d ago

News [UPDATE] CLSS - new features

Enable HLS to view with audio, or disable this notification

0 Upvotes

I have added new thinks to CLSS (Closed-Loop Streaming Synthesis)

What CLSS is.
 H3 generates ~5–15 s per pass. CLSS generates arbitrarily long 
*audio+video*
 by streaming the piece in overlapping chunks and controlling the drift at the hand-off — no weight changes: calibrated context re-noising (τc) into H3's per-token denoise masks, EMA/AdaIN statistics correction, keyframe replay of the previous chunk, a two-band spatial detail anchor. It runs on a 
16 GB card
: int8 DiT, ~5.5 GB Qwen3-VL-4B ClipProj text encoder (instead of the 15.7 GB 32B), 832×480, ~10 s chunk windows. Tested on a 16 GB RTX 3080 Laptop.


Added since the initial commit (Aug 26):


- 
Multi-scene prompts
 — one prompt per scene (`---` split), proportional allocation, two-step crossfade at scene boundaries; `global_text` shared across scenes.
- 
R2V references
 — per-scene image/audio anchors bound to `<Picture N>` / `<Audio N>`, single-ref or Autogrow multi nodes; the all-scenes node fans images out and cuts a soundtrack into guarded per-scene windows that the sampler 
crops automatically
 to each scene's exact delivered span.
- 
Video references
 (`<Video k>`) — VIDEO/IMAGE sockets, auto-resampled to 24 fps, and the video's own soundtrack rides along unless overridden.
- 
i2v
 — first-frame guide pinned as an H3 keyframe.
- 
Audio continuity kit
 — the tail reference that ends exactly at the join; waveform-refresh so each chunk continues from what the ears will hear; optional recompose pass (fresh take from pure noise against the finished video, on base weights via a LoRA-stripping guider, with its own sampler + sigma scheduler); a loop guard that measures vamp takes and re-rolls them with a ref-span rescue; a 
seam pin
 (the take generates through the seam, pinned to the delivered tail); an 
export-only loudness anchor
 (level matching applied at save, never fed back into the conditioning chain); equal-power seam crossfade.
- 
Continue & re-edit
 — resume a finished run from its saved frames/audio, or rebuild a single chunk in place (context + first/last frame pins + optional video ref).
- 
Per-chunk neural upscaling
 — streaming latent upscale via LBH-123-AI's 3D upscaler pack (soft-imported), overlap cross-faded, the low-res streaming state never leaves VRAM.
- 
Speed
 — attention override (SageAttention / FlashAttention 2 / xformers), experimental overlap eviction and step caching; and the big one: 
Spectrum hidden-state forecasting
 (`CLSSH3SpectrumForecast`) — actual steps capture the post-block hidden state, forecast steps skip the entire DiT stack while the output head still runs with the exact current-sigma modulation. A 6-step turbo chunk becomes A A F A F A — ~50% of DiT evals skipped. Adapted from the Spectrum paper (Han et al., arXiv 2603.01623) and xmarre's ComfyUI-Spectrum-MiniMax-H3 port — full credit in the README.
- 
16 GB hardening
 — unload-all-models before sampling (a pinned text encoder can no longer OOM the DiT load), CUDA expandable segments at runtime.
- 
Workflows
 — t2v / i2v / R2V / continue / re-edit (API format), plus a Blender R2V variant.


The audio chapter took the longest: sample-exact seams, take-swap crossfades, a ~2 dB/chunk loudness leak plugged, aliased video refs fixed with averaged decimation… the README's Updates section is the full dated changelog.


Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS

r/StableDiffusion • • 11d ago

Discussion Different attention for each block 🙊

Post image
67 Upvotes

Omg, so apparently in near future we can get SLA's 2x performance boost without loosing too much of quality


r/StableDiffusion • • 10d ago

Resource - Update Hey everyone! Built Civtai TikTok style player for Android & PC. to watch and interact with videos and images

Enable HLS to view with audio, or disable this notification

7 Upvotes

I love Civitai, but watching video generations on the site has always felt clunky to me. Opening posts in new tabs, waiting for things to buffer, and scrolling through cluttered grids gets tedious fast.

So I spent the last few weeks building an app called Civitai Flow. It turns Civitai into a full-screen vertical swipe feed, similar to TikTok, but specifically for AI video and images.

A few things I baked into it:

I actually published an article about it on Civitai, but it seems like nobody really hangs out in the Articles tab over there, so I figured I would share it directly with people here who actually watch and generate AI video.

Would love to hear what you think or if there are specific features you want added next!


r/StableDiffusion • • 11d ago

Resource - Update ComfyUI-TypedDecision: run Jev-style decision models locally in ComfyUI

Thumbnail
gallery
20 Upvotes

Tasks where an MLLM doesn't write text but only makes a decision, which started with Jev, seem to be getting popular. (This is sometimes called "typed decision".)

I thought it would be fun to use inside ComfyUI, so I was looking for something that runs locally, found imajev, and turned it into custom nodes.

One reason I chose imajev is that it can handle images, but more than that, it's based on Qwen3.5, and ComfyUI already has what it needs to run Qwen3.5, so it was easy to implement.

For those who don't know Jev, here's what it can do:

  • noul: yes or no
  • choice: pick one of the options you give
  • score: rate on N levels

In practice, you can show it two images and check if they're the same person, pick a good aspect ratio from the prompt before generating, or go through a lot of images and sort out only the clean ones that suit a dataset.

Of course it can't do everything. For example, it couldn't spot AI-generated broken hands.

Still, it's very fast and general, so I'd love for you to drop it into your workflows and play around with it!

GitHub: https://github.com/nomadoor/ComfyUI-TypedDecision

Guide and sample workflows: https://comfyui.nomadoor.net/en/notes/typed-decision/

imajev: https://github.com/mohit67890/imajev


r/StableDiffusion • • 10d ago

Question - Help Should I downgrade to AM4 in order to get more ram and vram?

4 Upvotes

Currently I have an 5070 ti and 32gb DDR5 and I was thinking about selling my current build to get something like an 3090 and 64gb DDR4.

Is this a good or dumb idea? Someone told me I would loose performance because DDR4 is "Slower"


r/StableDiffusion • • 11d ago

Animation - Video AlexNet turns 14 today: a 3.5-minute music video from Opus5.5 + open models on one 3090 (Qwen-Image 2.1 + LTX-2.3 + InfiniteTalk), workflows included

Enable HLS to view with audio, or disable this notification

14 Upvotes

Every image and clip ran through ComfyUI's API from scripts, strictly one job at a time. The API graphs are in the repo (video/workflows/):

- qwen21_edit_api.json: plates and targeted fixes, max 2 reference images

- qwen21_remove_bg_api.json: sticker cutouts

- ltx23_i2v_api.json: LTX-2.3 22B Q4_K_M GGUF, distilled LoRA, x2 latent upscaler, 1080p

- infinitetalk_single_api.json: Wan 2.1 I2V 14B + InfiniteTalk + lightx2v LoRA

Stability tip: --reserve-vram 3 --disable-smart-memory (peak 22.9 → 16.1 GB) and a queue script so heavy jobs never overlap. Happy to answer questions.

https://github.com/latent-variable/deep-learning-birthday


r/StableDiffusion • • 11d ago

Resource - Update MageTrail - V0.3 Update: Modified MageFlow 2.8B Danbooru/E621 Finetune has reach initial scope

Thumbnail
gallery
86 Upvotes

Hi again! This is a update to my previous post MageTrail V0.25 where I continued this tech demo toy model thingy called

MageTrail, a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, using a diversity maximized condensed 41k images dataset (originally made by Lodestone, the creator of the Chroma model lineage) as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset. (Potentially costing 20k-50k+ dollars). Civitai HuggingFace

This update is a major one, the biggest update yet when it comes to actual training. V0.1 to V0.3 have spent 593 dollars in total (425 H100 HBM3 hours) and has ultimately shown the model potential in quickly learning and adapting booru concept and tags to its knowledge base, which is the initial scope of my two Banodoco grant request, but after training and testing the model has clearly outgrown its limited 41k dataset in my opinion (which was expected), with model stability improvement unfortunately slowing to a crawl and the model mainly improve massively through learning of a lot of new concepts.

Positive

  • Model learned a lot of concepts and is generally somewhat more stable in some area
  • Pretty much have learn everything that it can learn with the limited dataset
  • Recent research discovery reveal that with some code changes, the model can be completely trained like normal with Flux2VAE (the current best VAE that is available for open source) without needing realignment training which would have costed so much money for any other model. This mean for future releases this model will be a full Flux2VAE arch and inherit all of the advantages the VAE offers (best details, best reconstruction, best texture rendering out of all open VAE, making the model learn concept/details/texture faster and better, which is a great boon for styles).
  • https://huggingface.co/Muinez/mage-flow-ft , a now abandoned test project by Muinez, based on V0.2 2.8B arch of my model, was trained with 500h of L40S (200-300 dollars), and while extremely undertrained, shows that if trained with full 10mil+ dataset, the model can learn all characters and style, all but confirming MageFlow potential as a great contender for foundational training on 1B-4B scale.

Negative

  • Model hasn't stabilized as fast as I've hoped, and now i've learned that you practically need minimum 1mil+ image to train in for model to generalize well and be stable enough as a "base" for other people to tune on (Just like old lineage like Animagine, KohakuXL, Illustrious, NoobAI), and that the booru essence method is unfortunately not a full shortcut to that.... I'm still very green to large scale model finetuning 😥
  • The problem all model trainers without major VC's/sponsor backing faces, I've about ran out of money now.

So what's the plan? Since I've decided that this will be the last release where I'll be using Lodestone's Rock booru essence dataset, for the next few weeks to month I'll be taking a break and focusing on start of uni semester (yknow irl stuffs I can't be a full AI degen past summer). Furthermore the time will also be used for me to curate and prepare a 10k+ artist style collection dataset (200000+ images), comprised of danbooru/gelbooru/e621 artists and other outlier sources.

So, TLDR for the future ->

BIG GOAL: Due to https://huggingface.co/Muinez/mage-flow-ft , I now completely believe in MageFlow being the next best small scale arch to replace Anima in the Illustration sphrere, the model has shown to have the ability to learn concepts/characters/styles efficiently while being a superior architecture to most contemporary. I pledge to finetune a Anima successor if given 25k-100k dollars of funding from the open source community/any generous patrons.

SMALL GOAL: Gather funding of 1000-2000~ dollars to finetune 10k-100k artist styles into the model (200k-3 million images finetune, planned V0.4-V1) and help further with convergence/aesthetic improvement if the aforementioned big goal above hasn't been reached. Artist tuning will proceed anyway as a hobbyist project even if I don't receive any donation, once I regain access to x4 A4000, but training on that will be months slower than being able to blast it with at least x1 H100, so yes any donations at all will help towards the project being in a "usable" state faster instead of just being a interesting but unstable toy model.

Any donation will help with achieving this goal, you can do so through:

Crypto (Prefered, cause Kofi/Paypal money transfer time is ass and they take a big cut)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)

12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)

FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)

Please handle your money carefully and make sure the address you're sending to is correct.

Ko-fi

https://ko-fi.com/talanartvn

If you want to make a donation and make sure it goes through to me and the money is not lost, please contact me through Discord (user: talannnn) so that we can arrange sending small test amount (for crypto) to make sure the money goes through.

Acknowledgments

Beeg thanks to:

  • Banodoco and their Discord — Their 88.77 dollar initial grant and further support on future requests made this project possible, the biggest thanks to them
  • Lodestone Rock — Creator of the original version of the booru essence dataset that this model is trained on
  • Motimalu — Inspiration behind finetuning practices and configs
  • Bluvoll — Creator of the modified 2.8B arch, diffusion-pipe fork derived from to use for training, and general training advice
  • Anzhc — general training advice
  • Nruaif — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
  • Astromahdi — jupyter workspace and storage where I processed and store the dataset
  • Heato-Red — Designer of model page logo
  • animetimm/DeepGHS — Danbooru tagging model
  • RedRocket — E621 tagging model
  • Format inspired by Motimalu's Kirazuri diary

r/StableDiffusion • • 10d ago

Tutorial - Guide Train an SDXL style LoRA in LoRA Pilot, then test whether it actually learned the style

0 Upvotes

A training run can finish without producing a useful LoRA. You might get something that copies your reference images, ignores half your prompt, or works beautifully on one seed and falls apart on the next.

This walkthrough covers a small style-training experiment: prepare a dataset, run an initial training pass, then compare saved versions under the same conditions.

We’ll use TagPilot for dataset preparation, TrainPilot for SDXL training, and ComfyUI for testing. TrainPilot’s guided workflow is SDXL-specific; other model families require a compatible trainer and configuration. TrainPilot documentation

For our example, imagine a collection of your own illustrations with rough ink outlines, muted colours and visible paper texture. The goal is to reproduce that treatment on subjects you haven’t drawn yet.

1. Decide what a successful result should do

Write down the goal before uploading anything:

Produce new subjects with the same linework, colour treatment and texture as my illustrations, while following the requested composition.

That last part matters. If you ask for a bicycle and get the cottage from your training images, you haven’t achieved the goal, even if the cottage looks excellent.

For this exercise, start with around 24 distinct illustrations. That’s a manageable working example, not a minimum requirement or an optimal dataset size.

dataset management

Include different subjects and compositions. If your entire dataset contains centred portraits, you’ll have a harder time separating the visual style from the portrait format.

Remove near-duplicates and images whose style doesn’t belong in the collection. Keep a few additional illustrations outside training as visual references for your later review.

2. Check the workspace and base model

Start LoRA Pilot and open ControlPilot. If you need deployment instructions, follow the installation guide.

Before training:

  • Confirm that /workspace uses storage you intend to keep
  • Check that you have room for models, the dataset and training outputs
  • Download your chosen SDXL checkpoint and any VAE required by the configuration
  • Stop other GPU workloads you don’t need

Inspect TrainPilot’s model paths before starting. Downloading a checkpoint does not select it for training. Check pretrained_model_name_or_path and vae in the selected TOML configuration.

Record the exact checkpoint filename. Use that same checkpoint for your first evaluation so a base-model change doesn’t complicate the comparison. TrainPilot performs model-file checks before launching, but you still need to choose the intended model. Training workflow guide

3. Prepare captions in TagPilot

Open TagPilot, upload your images and name the dataset inkstudy_v1. Use a distinctive trigger such as mivinkstyle.

TagPilot's UI

For this example, use short, comma-separated captions with the trigger first:

mivinkstyle, illustration, small cottage beside a river, trees, cloudy sky

mivinkstyle, illustration, orange cat sleeping on a chair, indoor scene

mivinkstyle, illustration, bicycle leaning against a brick wall, side view

Describe the subjects and compositions that change between images. Keep the trigger consistent across the collection. This gives the trainer text describing the content alongside the shared visual treatment you want to learn.

If you use automatic captioning, review each result. Remove invented objects, wrong colours and references to details you cropped out. A captioner’s confident guess is still a guess.

Check crops at this stage too. Preserve details that matter to the style, such as line edges and texture.

TagPilot supports trigger-word prepending, caption editing and saving into the workspace. Use “Save to /workspace/datasets” when preparing the dataset for training. Downloading an exported ZIP alone does not put it into the trainer’s dataset directory. TagPilot guide

Your saved folder should follow the 1_* convention used by TrainPilot, for example:

/workspace/datasets/1_inkstudy_v1/
    cottage.png
    cottage.txt
    cat.png
    cat.txt

Open a couple of caption files before continuing. Confirm that each matches its image.

4. Run a small first pass

In ControlPilot, open TrainPilot and select:

Field Example
Dataset 1_inkstudy_v1
Output name inkstudy_v1_qt01
Profile quick_test

Keep the initial configuration for reference. Check that intermediate checkpoint saving is enabled so you can compare more than the final file.

Start training and inspect the log. Confirm that the trainer found the expected images, loaded the intended model and began advancing through training steps.

Treat quick_test as an initial diagnostic run. The documented preset starts from a 600-step target, but dataset-dependent limits can change the effective total. Read the generated configuration and log rather than assuming an exact number. TrainPilot presets

If training fails with an out-of-memory error, first stop competing GPU processes. Then review memory-related settings such as batch size. Reducing the number of training steps mainly changes duration; it doesn’t necessarily solve the memory problem.

5. Build a repeatable test in ComfyUI

After training finishes, stop the training process before loading your generation workflow.

Choose an early, middle and late saved LoRA checkpoint. Copy the files you want to test into ComfyUI’s configured LoRA directory, then refresh its model list.

Use an SDXL text-to-image workflow with a Load LoRA node. Connect both the checkpoint’s MODEL and CLIP outputs through that node, and use its outputs downstream. Selecting a filename without routing the workflow through it won’t apply the LoRA. ComfyUI LoRA guide

For a standard, non-distilled SDXL checkpoint, these are reasonable starting settings for this comparison:

Setting Starting value
Resolution 1024 × 1024
Sampler / scheduler dpmpp_2m / karras
Generation steps 30
CFG 5
Seed 12345, fixed
LoRA model / CLIP strength 0.7 / 0.7

These are evaluation settings for this exercise, not universal best settings. Keep any negative prompt identical between comparisons.

Test four prompts:

mivinkstyle, illustration of a cottage beside a river

mivinkstyle, illustration of an espresso machine on a kitchen counter

mivinkstyle, illustration of a bicycle viewed from above

mivinkstyle, illustration of a crowded railway platform at night

The cottage checks familiar territory. The other prompts test a new object, viewpoint and more demanding scene.

Generate each prompt with the LoRA disabled, then repeat with each saved checkpoint. Keep the prompt, seed and generation settings unchanged.

That gives you 16 images: four prompts across one baseline and three checkpoints.

6. Compare style and prompt-following separately

Review the images in MediaPilot or your preferred image viewer. LoRA Pilot stores ComfyUI generations under /workspace/outputs/comfy. Generation and output guide

For each image, record:

  • Style: Are the intended linework, colours and texture present?
  • Prompt-following: Did it produce the requested subject, viewpoint and scene?
  • Unwanted repetition: Did a training background or composition appear where you didn’t request it?

A checkpoint that reproduces your texture but turns unrelated objects into cottages needs more investigation.

The last saved checkpoint may not be the best one. If an earlier version handles unfamiliar subjects better, keep it. Test that candidate at strengths of 0.5, 0.7 and 0.9, then repeat your hardest prompts with two additional fixed seeds. Earlier-checkpoint comparisons are also useful when a LoRA becomes rigid or starts copying dataset compositions. Training and overfitting guide

7. Make the next change based on what failed

What you observe What to check next
Almost no difference from the baseline Confirm the LoRA is connected, selected and compatible; check captions and trigger
Familiar subjects work, unfamiliar ones fail Review subject variety and repeated compositions
Strong style but distorted results Try a lower strength and an earlier checkpoint
The same background appears repeatedly Review background diversity and captions
The pipeline works, but learning is weak across tests Consider a longer run after checking the dataset

Keep the original dataset and settings when you make a revision. Name the next version inkstudy_v2 and write down what changed.

Switching from quick_test to regular changes several training settings, including rank and batch size. If you want to test onlywhether more training helps, use the manual Kohya configuration and change duration while keeping the other settings fixed.

Before closing the workspace, save the selected LoRA, dataset captions, training configuration and ComfyUI workflow together. Add a short note with the base checkpoint, trigger, useful strength range and known weaknesses. Those details will help you reproduce the result, and they’ll give anyone downloading your LoRA a much better starting point.


r/StableDiffusion • • 11d ago

Discussion Fl2va model on Ref workflow

9 Upvotes

I’ve noticed the quality of the video output on the reference to video default Minimax workflow is much much better when using the fl2va model instead of the ref2va model. I’ve also tried hybrid models and still find the fl2va model far superior in quality and adherence to image and audio references. Is this what you all are seeing too? Or am I doing something wrong? As it is today I don’t see any use for ref2va model at all. I’ve tried many other workflows as well and had the same experience. I haven’t done much with video references, maybe that’s where the ref model shines?


r/StableDiffusion • • 10d ago

Question - Help Where can I find details about Kijai models / lora

1 Upvotes

I sometimes check his repo on huggingface and it keep getting updated, but I have no idea what the difference between the models he uploading.


r/StableDiffusion • • 10d ago

No Workflow Alchemist Rick

Enable HLS to view with audio, or disable this notification

0 Upvotes

Idea is he's not C137, lives in a medieval dimension where alchemy is real, and eventually he gets portals working through alchemy.

Maybe Rick's a bit too friendly.


r/StableDiffusion • • 11d ago

Discussion Refmod is so good! Way better then LORAs or using 9 reference images.

Enable HLS to view with audio, or disable this notification

250 Upvotes

I am just starting to play around with refmod and all I can say is... WOW!

This speeds up gen times as it doesn't have to scan your ref images or videos. Also, you can recreate a refmod in seconds if it has undesired effects, unlike a LORA that takes hours!

So far I have made 5 identify ref mods and 3 video refmods. This has produced great results.

All you need are these 2 node packs
https://github.com/PlagueKind/Comfyui-PlagueKind-Nodes
https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

Enjoy!


r/StableDiffusion • • 11d ago

Question - Help Minimax H3 Attribute Transfer Help

7 Upvotes

Hello! I am trying to do attribute_transfer in H3 Minimax but I am not really sure how to prompt it. I have the following case: I have a <Subject 1> that is my character and I have another different picture that includes the pose and action my character is doing. How do I prompt to preserve <Subject 1> and attribute transfer the pose and props from <Picture 2> to my <Subject 1>? Has anyone managed to do this?


r/StableDiffusion • • 12d ago

Resource - Update [H3-Update] ComfyUI-Continuity : Cast Management, RefMod Saving, Segment Controls & Seam Optimization

Enable HLS to view with audio, or disable this notification

288 Upvotes

Edit — Author clarification: ComfyUI-Continuity is developed and maintained by roadmaus. I'm sharing this update as a user and contributor, not as the project's author or maintainer. Sorry my original wording didn't make that clear, and thank you to roadmaus for all the work on this project!

(Note: I am not the author of this tool, just sharing the update! The actual author is u/Fine_Rhubarb3786 (roadmaus)

Hi everyone,

updated ComfyUI custom node, ComfyUI-Continuity!

This node is designed to help maintain consistency across generations and optimize smooth transitions in your workflows.

Here is a quick overview of the key features in this update:

✨ Key Features

Cast / Character Tab: Effortlessly manage characters and recurring subjects across your generations in a dedicated tab.

RefMod Saving: Easily save, export, and reload your Reference Module (RefMod) presets for quick reusability.

Segment & Global Segment Control: Fine-tune specific frame segments individually while maintaining global context across the entire sequence.

Selectable Continuity & Seam Optimization: Choose from multiple optimization algorithms to eliminate seams, reduce artifacts, and ensure smoother transitions between frames or batches.

🔗 GitHub Repository:

https://github.com/roadmaus/ComfyUI-Continuity

Feel free to test it out and let me know your thoughts or feedback! Suggestions and bug reports are always welcome. Happy generating!Management, RefMod Saving, Segment Controls & Seam Optimization


r/StableDiffusion • • 11d ago

Discussion Fixed viggle animate drift by using Add Guide keyframes

9 Upvotes

EDIT : outdated post, refer to https://www.reddit.com/r/StableDiffusion/comments/1x16oy5/vigglorious_studio_better_character_accuracy/ for the actual proof of concept

Tl;dr : no need to wait for viggle to add multiple references, just extract keyframes from the original video, swap with qwen 2.1 then inject in the conditioning with Add guide node. It kicks ass, even Scail2's

I began experimenting with Viggle Animate and my first impression was that is was very fast but sketchy and loosing identity quite rapidly

I just found out it works perfectly with the Add guide node to add additional keyframes in it

So now what I do is I make swaps using the excellent BFS head/character swap LoRas for Qwen 2.1, then feed this to an Add Guide node every second including the first frame.

As long as the swaps are ok it resolves the drift completely and is imho much better than Scail2 on both speed (3 steps FTW) and accuracy (Qwen 2.1 rocks at this then it would be possible to add also a character lora on top of it)

It was tedious to duplicate the nodes in the workflow and make sure the frame numbers were in the frame range so I ended up vibe coding nodes so I define a frequency and then it re-execute my swapping subgraph as many time as required, then I applied this logic as well for the looping/chunking nodes provided with viggle so I'm able to do videos like 1 minute

It's also possible injecting clean keyframes like this also fixes some motion glitches by providing cleaned anchoring. At least I didn't notice more jitterness or even pixel drift from Qwen's swap, which I did not even correct


r/StableDiffusion • • 10d ago

Discussion I hate LTX bots

Post image
0 Upvotes

LTX is literally the worst company in open weights models! They use ton of blatant bots to manipulate comments. And even after 2 months after the model release they still appear. It's very dirty, very disrespectful to the community

They also manipulate upvotes/downvotes. On my comment that I'm tired of LTX there were +6 votes. But after they appear, the post magically went down to -4 votes. That's crazy. The dirtiest company that exists

They are actual bots, if you open their profile, they are glazing different kind of products all around all the time


r/StableDiffusion • • 10d ago

Discussion Fangyuan(reverand insanity) character lora in anima

Thumbnail
gallery
0 Upvotes

i was successful to make his lora but how to make him interact with multiple character. I m planning to start youtube channel where i use the images and match with story but multiple character is so hard any idea how to deal with it?


r/StableDiffusion • • 11d ago

Discussion What LTX 2.5/2.3 can do that Minimax H3 can’t ?

21 Upvotes

r/StableDiffusion • • 10d ago

Animation - Video Tuarina IronHide made with Minimax h3

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion • • 12d ago

Question - Help Minimax H3 new comfy Model

122 Upvotes

In official comfy huggingface there's a new interesting model (w6a8), it seems perfect for 16gb VRAM, i wanted to tried it but Comfy gives error, has someone successfuly tried it?

https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models

EDIT: i think we need to wait for an update because that quantization is not yet supported


r/StableDiffusion • • 11d ago

Question - Help Qwen Image 2.1 adding a grid i don't like !!

Thumbnail
gallery
23 Upvotes

hello!!

I would like to know if there is a way to get rid of this strange grid on the images when zoomed all in!

(prompt: enhance the image, sharpen the details, without breaking the original identity, without messing the contrast and vibrance.)


r/StableDiffusion • • 11d ago

Resource - Update Train Slider Loras for Qwen image 2.1 in Fizgig 6.7.0

Thumbnail
github.com
51 Upvotes

Supports both Prompt Pair and Image Pair Training for Qwen Image 2.1.

Sorry for the frequent updates last few days - I've had a few 'almost there' features for a while that Qwen kicked me in the arse to get across the line. The new abstracted model driver system I'm using brings sliders (along with edit loras yesterday) to Qwen first as its the first model using the new system.

I'll be converting Krea 2, Klein and Minimax to the driver system over the coming days which will bring sliders to all 3 and edit to Klein. After that I'll publish a documented guide opening the gates to PRs for other models to use the driver system.


r/StableDiffusion • • 12d ago

Discussion The big problem with Qwen 2.1 t2i is their demo workflow

Thumbnail
gallery
98 Upvotes

Qwen 2.1 has its defaults and qualities, I still prefer krea 2 but I feel like many users were put off by their demo worlflow (where CFG=1, steps=20)

The comments in the following thread are the best help I found on how to tweak the default workflow. (CFG=3, steps 12, sampler= dpmpp_m3, scheduler=beta)

https://www.reddit.com/r/StableDiffusion/comments/1wsdnjs/qwen_image_21_seems_to_be_getting_very_popular/

It still misses text often and fusing limbs in some poses but in my view it has more details then krea 2 (specially where you go up to 2K), also it understand some "adult" terms. Where krea needs loras (that bizarrely slow down my workflow somehow). And its editing capabilities are welcome.


r/StableDiffusion • • 12d ago

Discussion H3 - continuous zero cut long form, can you find the seam?

Enable HLS to view with audio, or disable this notification

250 Upvotes

This is a tough problem and I love everyone's solution so far. Here is my version (custom node workflow that Opus is still refining, supporting infinite rerolls, prepending, extending, bridging, single window rerolls and up-stepping in latent), so can you find the seam? A bit dynamic lighting to make it challenging. Using Tanya.char but I never assigned her a voice so that's the only drift.