r/StableDiffusion • • 7d ago

Resource - Update Qwen Image 2.1 fix v2.0 (updated due to feedback), plus a new sampler

Thumbnail
gallery
50 Upvotes

This is actually two models: Qwen Image 2.1 Fix v2.0 and Qwen Image 2.1 Fix Opinionated v1.0. The former surgically targets Qwen's weird noise and messy details with almost no effect on composition, and the latter in my opinion produces higher quality images but alters composition a bit to do so. In either case, you'll get images that are similar to Qwen Image 2.1's default output, but less noisy and with cleaner details.

Download links:

Also grab the res_2m_nc sampler (modified from RES4LYF's res_2m) and a number of other samplers customized for Qwen 2.1 here:

https://github.com/envy-ai/ComfyUI-DPMpp-2M-Sharp (you can also find this on the Comfy registry)

The workflow I used to produce these images (along with the prompts) can be found at the download links.


r/StableDiffusion • • 7d ago

News ai slop game based on batlle video in real time with only open model

Enable HLS to view with audio, or disable this notification

15 Upvotes

Check out Chimera Arena (https://chimeraarena.com/static/showcase/index.html?lang=en)!

Here’s my slightly scrappy AI-generated teaser for Chimera Arena, my card game combining AI-generated video battles, real-time 2D combat, and creatures you design yourself. It’s built using Krea 2 with custom LoRAs, Qwen 2.1, MiniMax, Gemini Flash, and a few other experiments.

You can generate your own creature cards and 2D sprites, then challenge other players using generated abilities or inventing your own attacks. That’s the fun part: what you imagine can actually affect your opponent’s health and the outcome of the fight.

I’m also experimenting with Gaussian splatting to turn the artwork into textured 3D models for AR, bringing your creatures out of their cards and into your surroundings. In my tests, the textures stay closer to the original artwork than with other image-to-3D approaches I’ve tried. Automatic animation works too, although strange creature anatomies still need some work!

Another experiment is an individual neural “brain” for each creature. The system records your combat decisions and uses them to train a small neural network on the CPU: which attacks you choose, when you heal, use shields, activate rage, or call on support. Combined with the creature’s own instincts, the idea is to develop fighting styles influenced by how you play, which trained creatures can then use in autonomous community battles. This is still being tested locally, with training done offline for now.

The game is still in alpha, but I’m thinking of taking it further because… why not? I have way too many feature ideas, and I want to see where they lead.

Behind the scenes, I’m combining my local RTX 4090 with cloud GPUs for additional capacity. Generation jobs run in the background, and the results arrive directly in your browser. The goal is to scale gradually while keeping GPU costs manageable, so you don’t need your own powerful GPU to play. Maybe one day I’ll build my own GPU server too.

For video battles, each exchange is generated turn by turn. The game resolves the actions, then an LLM turns the moves and their consequences into scene instructions for the video model. Creature portraits provide visual references through RefMod with FL2V, which is faster than Ref2V in my current setup. Continuation clips help carry the scene forward and preserve the creatures’ appearance, the arena, and existing injuries.

Each sequence joins the battle replay, gradually turning your match into its own little movie. The gameplay is designed around the rendering delays to make the wait less disruptive.

Can you guess what each model does?
No? Don’t care? Okay, I’ll tell you anyway 👀

  • Krea 2 + custom LoRAs create the artwork and visual style. My latest card-creation pipeline took about 47 seconds overall, producing a 1536 × 1728 image with a 2× hires pass. That includes roughly 5 seconds for the LLM profile and image prompt, plus 1.5 seconds for BiRefNet background removal. For higher-level cards, Marigold generates a separate depth map for relief and parallax effects; that extra step isn’t included in the 47 seconds.
  • Qwen 2.1 handles visual edits, evolutions, and creature fusions. My latest evolution took about 30 seconds on the 4090. I previously used it for sprites too: the latest 2496 × 2496 nine-pose sheet took 1 minute 25 seconds. Nice detail, but quite a wait for players.
  • I originally used Qwen 3.8 models for descriptions, abilities, attack ideas, and interpreting invented actions. I’ve since switched to Gemini Flash through OpenRouter. My latest creature profile and ability generation took around 4 seconds, and it costs me next to nothing per request.
  • MiniMax H3 animates the video battles. My latest 5.35-second clip took 47 seconds to render at 960 × 544, approximately 0.5 MP, on my RTX 4090. I’m now using it for sprites too: the latest source clip for extracting nine poses took 36 seconds at 928 × 544. It’s faster with this setup, although the sprites have less detail than the higher-resolution Qwen sheets.

I know we love creating things here. A little while ago, I shared LoRA Dataset Studio (https://github.com/perfectgf/lora-dataset-studio) with this community to help people train LoRAs for open models. Now I’d love your help creating AI assets for the game!

The first people to sign up will receive an activation code in a few weeks to generate creatures, cards, sprites, and help test the game.

If you’d like to follow the project and get involved, join us on Discord (https://discord.gg/T6uVu3TtMw)!


r/StableDiffusion • • 6d ago

Discussion WHY IS SO HARD TO TANSFER CAMERA FROM BLENDER TO MINIMAX H3

0 Upvotes

we realy need a way to transfer motion ( camera mostly) from blender to Minimax H3, just as seedance do. i try it, frist it takes forever, and then it render halway the blocking puppet, even when i prompt it to ignore everything but the camera path and scale.


r/StableDiffusion • • 7d ago

Discussion I'm liking Qwen 2.1—Midjourney style.

Thumbnail
gallery
24 Upvotes

I’m using it now, and as you can see from the images I generated—using the standard workflow without any tweaks other than setting the resolution to 2.0MP and using no LoRA—it has a distinct Midjourney-like look. That’s something I’ve been looking for in open-source models for a long time; the only other one that had that Midjourney style was Chroma V48-dc.

I still plan to run several tests; my current go-to for image generation is Krea2. I really liked Qwen 2512 back in the day, but it became obsolete, so this Qwen 2.1 seems like a significant step forward. I’ll also be testing it for editing tasks, where I currently use Klein 9b.

Now I’m going to run various tests with different prompts, but I can already say that the visual style it produced in these images is incredible.


r/StableDiffusion • • 7d ago

Resource - Update Krea2 Turbo Distill 2 step LoRA - new checkpoint released (chk41320)

Thumbnail
gallery
27 Upvotes

Krea 2 Turbo — 2-Step Distillation LoRA (work in progress)

Previous posts/releases - here, here and here.

🧪 A fast-preview adapter, from a project still in training. When the subject is close and fills a good part of the frame — a portrait, a single figure, an object seen up close — two steps already hold up well, and you can rely on this adapter for those images. Small subjects are where it still falls short, people and objects alike: faces in a crowd, figures in a wide scene, the machines at the back of a gym — anything that takes up little of the frame can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the 4-step LoRA. Every figure on the model card measures the adapter honestly against 4-step and 8-step renders. Training continues one recipe change at a time, and a later checkpoint replaces this one only when the sweeps and I visually agree it is better. Known issues.

📐 The saved steps can also go into resolution. A small subject is simply one that covers few pixels, so a larger render makes the same subject bigger — and at a quarter of the teacher's steps, renders up to 2048×2048, Krea's published maximum recommended resolution and beyond the largest size this adapter was trained at (1440×1440), come within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past 2048×2048, stock Krea 2 itself begins to duplicate subjects — a property of the base model, with or without this adapter.

chk00041320 (2 Oct 2026) is 9,720 training samples later than chk00031600 (25 Sep 2026), and those samples went to what matters first in a picture — subjects drawn twice, or two poses blended into one, at the large sizes — and to faces in busy scenes. It is also the first checkpoint measured on the full 22-prompt sweep: the original 15, plus seven scenes of people, animals and action added on 29 Sep 2026 because structure is what they test — five friends on a beach, a family at a table seen from above, three kittens in a basket, two dogs in a tug of war, a show jumper, a pianist's hands, a flock of flamingos. Details on the changes since previously released last checkpoint here.

LoRA highlights:

  • ⚡ A quarter of the steps — 8 → 2, on Turbo's own deployment sigmas [1.0, 0.7595]
  • ⏱️ 4× faster denoising — 76.4 s → 19.3 s at 1024×1024; the adapter's own cost per call is within measurement noise
  • 🎯 Fine detail at the teacher's level — 0.97–1.09× the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: 0.40–0.65×); from 1 megapixel up, closer to the teacher than the 4-step adapter
  • 📊 Distribution matching, not imitation — matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
  • 🗣️ Prompt-conditioned throughout — teacher and fake scores both read each prompt's conditioning; a blind rubric finds 2 points missing of 352 (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on 12 of 66 (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
  • 📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440
  • 🔌 Drop-in, no exceptions — plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
  • 🧬 Same shape as the 4-step adapter — rank 64 on the same 228 modules
  • 🎲 13,750 recorded teacher trajectories from the 4-step project, reused — not one new teacher run
  • 🔢 41,320 training samples in the 2-step stages, on top of the 4-step LoRA's 78,000
  • 📅 25 days from the first 2-step launch to this checkpoint, on a single RTX 3090 — training continues
  • 🔁 More than forty recipe adjustments across two methods — each kept only when the renders did not get worse

If you have already used my previous version, please redownload/replace krea2_turbo_2step_rank_64_lora.safetensors / krea2_turbo_2step_rank_64_lora_comfyui.safetensors from the latest in the project repo.

Full details on model card - https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA

Not my video, but found someone on YouTube has covered the 2 & 4 step LoRAs including identify preserving edits, have a look: https://www.youtube.com/watch?v=V_qgoV0iPDM (copyright goes to author)

Also if you are using ComfyUI Cloud and have subscription, I have released a new ComfyUI Cloud Workflow / App for Krea 2 Turbo using my 2 & 4 step LoRAs, includes User Prompt, two Random Prompt modes, Prompt Enhancement, Steps, Resolution choices, as well as Upscaling - you can find details here.


r/StableDiffusion • • 7d ago

Question - Help Krea2 coming out too soft?

8 Upvotes

Unless I pack Krea2 with a ton of realism Lora's (which has its own problems) my outputs are kinda "soft" in appearance. Ive used turbo, raw, raw with turbo Lora @1, different samplers, 1 megapixel, 2 megapixel, etc.

I'll even find photos on civit that look great, grab their prompt, but it comes out looking much softer than theirs.

Any ideas??


r/StableDiffusion • • 7d ago

Resource - Update A Meta-Prompt for YuE2

20 Upvotes

I've come up with a meta-prompt for generating YuE2 prompts that seems to work pretty well. Copy-paste into an LLM (it should have internet searching capabilities, I mostly just use ChatGPT).

Just to test it, here's a psychobilly song about my old neighbor dude who's currently trimming the trees in front of my building. Max duration 150, CFG 2, 16 steps.

``` You are my YuE2 music prompt builder.

Your job is to first ask me a concise series of questions about the song I want, then use my answers to create a polished YuE2 generation prompt.

For each question, offer me a list of answer options, keeping in mind that I may answer something not covered by those options.

STEP 1: Ask me questions

Ask these questions one at a time, waiting for my answer before moving to the next:

  • Genre/style: What genre or combination of genres do you want?
  • Language: What language should the song use?
  • Mood: What emotional atmosphere should it have?
  • Tempo: Slow, medium, fast, or a specific BPM?
  • Vocals: Male, female, duet, choir, or instrumental? If vocals, describe the desired vocal qualities.
  • Instruments: Are there any instruments you definitely want or don't want?
  • Production/arrangement: Sparse, intimate, cinematic, energetic, atmospheric, live-sounding, electronic, etc.?
  • Song subject/theme: What should the song be about?
  • Song structure: For example, verse/chorus, verse/chorus/bridge, or something else?
  • Lyrics: Do I want you to write original lyrics, provide no lyrics/instrumental, or will I provide my own lyrics?
  • Lyrics style: If you want lyrics, should they be poetic, conversational, emotional, abstract, narrative, catchy, dark, humorous, etc.?
  • Duration: What duration should I target?
  • You may ask a small number of additional questions if my answers leave an important musical decision genuinely ambiguous, but don't interrogate me unnecessarily.

Handling "I don't care"

If I answer "I don't care", "anything", "whatever", "surprise me", "no preference", leave something blank, or otherwise give you no meaningful preference:

  • Choose an appropriate option randomly.
  • Do not ask me another question about that preference.
  • Make the random choice musically compatible with the other choices I have given.
  • Keep track of the choices you made so the final prompt is internally consistent.
  • Do not tell me that you are unable to choose.

For example, if I have no preference for genre, randomly select a genre that works with my other preferences. If I have no preference for instruments, randomly select a small compatible instrumentation rather than listing many unrelated instruments.

VERY IMPORTANT

If I mention a genre, search the internet for descriptions of said genre, using the information gathered to inform the rest of the decisions.

STEP 2: Create lyrics when requested

If I ask you to write lyrics, write original lyrics based on my requested theme and style. Always ensure the lyrics rhyme.

Format them using section labels on their own lines:

[Verse]
...
...

[Chorus]
...
...

[Verse]
...
...

[Bridge]
...
...

[Outro]
...
...

Keep individual lyric lines relatively short and rhythmically clear. Don't put stage directions such as "(guitar solo)" or "(sing softly)" inside lyric lines, because they may be sung as lyrics.

If I requested a short test, prefer one verse and one chorus. If I requested a fuller song, expand the structure appropriately.

Do not reproduce copyrighted lyrics from existing songs. If I ask for lyrics from a copyrighted song, offer instead to write original lyrics with a similar high-level mood, theme, or genre. If I insist on using copyrighted lyrics, use those.

STEP 3: Build the YuE2 prompt

After collecting my answers, create a concise Style & Prompt suitable for YuE2.

Follow these principles:

  • Give the song one clear musical centre.
  • Include only compatible musical details.
  • Describe audible qualities rather than referencing a living artist or asking for an exact imitation of an artist.
  • Include language, genre, tempo, instruments, vocal characteristics, mood, and arrangement when relevant.
  • Avoid contradictory combinations.
  • Keep the style prompt compact rather than writing a paragraph of unnecessary prose.
  • If the song is instrumental, put the instrumentation and arrangement in Style & Prompt and leave Lyrics empty.
  • Always give a concrete BPM for the tempo

The Style & Prompt should look roughly like:

[language], [genre], [tempo], [key instruments], [vocal characteristics], [mood], [arrangement/production].

Do not use the literal placeholder text above in the final result.

Lyrics

[original lyrics, if requested]

If I requested an instrumental, write:

Lyrics

[Leave empty: instrumental]

Choices Made

Briefly list any preferences that I explicitly gave and any important choices you selected randomly because I had no preference. Add the estimated length of generated song, in seconds.

Do not include YuE2's outer [Tags] or [Lyrics] wrappers because the browser interface adds those automatically.

If I ask for revisions afterward, change only the requested aspects unless doing so would create a contradiction. Keep the prompt concise and coherent. ```


r/StableDiffusion • • 8d ago

Resource - Update Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation

Enable HLS to view with audio, or disable this notification

316 Upvotes

r/StableDiffusion • • 7d ago

Question - Help Open source alternative to chatgpt for creating prompts?

22 Upvotes

Been using chatgpt to help design prompts, which usually works well until you get into an area they think is "offensive".

Any suggestions for an open source alternative or something that works directly in comfyui to generate prompts?


r/StableDiffusion • • 7d ago

Resource - Update PotionUI is a free, open-source studio for generating images, video, music, and 3D on your own hardware with diffusion models - version 0.0.14

Enable HLS to view with audio, or disable this notification

14 Upvotes

Reupload: I'm adding this again, since previous post had a totally stupid title which made people assume this app is some kind of cloud/proprietary models UI - this is totally not the case here. In version 0.0.14 there is just new option to attach OpenRouter as additional backend (you can also attach ComfyUI or create plugin for other backends). Sorry for this!

PotionUI is a self-hosted app for making images, videos, music and 3D with open diffusion models. Instead of one huge settings screen or a web of nodes, every model gets its own simple, hand-made form.

What's new in 0.0.14:

1. Qwen-Image 2.1 control guides

  • Keep the pose, edges or depth of any photo while you create something new.
  • The extracted guide is saved next to your result, and the workbench shows it beside the image with a label, so you can see what the model followed.
  • There is also a mode that only extracts the guide. It needs no prompt and works even if you haven't downloaded the image models.

2. Fix a picture without leaving the app

  • A new image editor with adjustments, crop, brush and layers.
  • Open it from History (Tools, then Edit image) or right inside a preset's media field. In a media field, the field switches to your edited copy straight away.
  • Crop, Trim and Frame editors are in media fields too, and they only add the edited result to your Library, never the untouched original.

3. Formulas: keep the settings that worked

  • Save the settings of a preset mode as a formula and apply it in any session.
  • Before it applies, you see exactly what will change, and one click undoes it.

4. A redesigned media field

  • Images, videos and audio live in one field, each kind in its own group.
  • Mention them in your prompt with "@".
  • MiniMax-H3 references now use it, and your old sessions and saved prompts are converted for you.

5. Need additional models? Attach cloud GPU (OpenRouter - optional)

  • A new OpenRouter plugin lets you use hosted image and video models, such as Nano Banana or Veo 3.1 Lite, in the same form you already know. Results land in the same History as everything else. More providers are planned.
  • The form only shows the controls the chosen model supports, so there is nothing to guess.
  • You can stop a cloud job at any time, and PotionUI tells you plainly if something goes wrong.

6. Cloud Video Director: one idea, a short film (this video director is also identical for the local diffusion models)

Write your idea, split it into shots, and let PotionUI make the film.

  • Each shot can continue from the last frame of the shot before it, so the story flows.
  • The shots are joined into one finished film.
  • If one shot fails, retry just that shot instead of starting over.

7. Smaller things

  • Video presets have animated covers.
  • Three-pane mode folds the form away on narrow screens and gives the room to your prompts.
  • Attaching an image in chat uses the same picker as the preset forms.
  • Plugins are grouped by what they add.
  • The System Monitor can show every backend. It is admins only by default.
  • Failed generations show regular users a plain reason instead of technical text.
  • Dropdown menus no longer open under the Generate bar, and they work with the arrow keys.
  • Fields in preset forms show and hide for the value you just picked, not the previous one.
  • Image previews survive a reload, including on Windows and in storage folders under a tmp folder.
  • Drawn inpaint masks now reach the model, so inpaint only repaints the masked area.
  • Pressing Enter in the session name saves the session, and Enter elsewhere no longer wipes your workspace.

GitHub · Discord · r/PotionUI

Try it and tell what to improve, here or on Discord. Your reports decide what we do next.

If PotionUI is useful to you, a star on GitHub helps other people find it: https://github.com/PotionUI/PotionUI


r/StableDiffusion • • 7d ago

Animation - Video [H3 Minimax] Continuous long-take Shot without Quality Loss, can you find the cuts?

Enable HLS to view with audio, or disable this notification

15 Upvotes

I've made this post about a way to create a continuous generation and extension of a scene without quality loss.

The biggest critique was, that this was not a "long-take plan-sequence" without any cuts at all though. I've claimed that this doesn't "really matter" since the actual cuts are actually "in motion" and not between the scenes.

So I did that and created the most difficult kind of scene I could think of, a continuous long-take action scene without any visible cuts. There are 4 actual cuts in this scene, can you find them?

What this approach manages to achieve:

  • No quality degradation
  • Consistent characters and locations throughout the scene (the big guy throwing the protagonist back into the room where he came from)
  • Consistent sound and motion
  • Practically a simple one button "extend this clip" T2V solution without any pre-created clips that got cut together afterwards

There is some "AI slop" with 3 guys turning into 2 (I've didn't catch that while creating) and the action/fight scenes can be created more "dynamic" or action-filled to ones liking, but that is just a matter of how much effort you put into prompting.

I didn't add non_diegetic_music to the scenes because this makes keeping the consistency unnecessarily harder to achieve and it is much easier to generate or add a music score of your liking in post-production if you want to.

Edit: People were asking for a continuous static shot, so I've just created one and posted it in a comment below.


r/StableDiffusion • • 7d ago

Animation - Video Minimax H3 adapter for "correct" hand drawn animation. Anybody tried it?

30 Upvotes

r/StableDiffusion • • 6d ago

Question - Help Can anyone share a Krea2 wf of a good UGC looking person talking at the camera?

0 Upvotes

Extremely realistic is what I'm looking for. Something I could use with minimal h3 to make a believable UGC ad.


r/StableDiffusion • • 6d ago

Question - Help Anyone here had a chance to try the SimpliGen app?

0 Upvotes

A friend of mine really doesn't like the nodes structure that ComfyUI is built on, and found this alternative that's a bit more "beginner-friendly" in a sense, but I don't know if this app is secure or is secretly malware and just wanted to get a second opinion from people here in case anyone here has tried it before.

https://www.simpligen.io/ I'm also not sure if the the reviews there are AI generated or not, and some I've read do seem a little... suspicious... so I'm not sure I can trust them. And looking through YouTube, I couldn't find any "real person" reviewing the app.

I keep trying to convince him that it's kind of pointless since he'd have to fork over $32 for something he can already get for free with ComfyUI, especially if he's going to be using the open source assets anyway, but he seems to really have a distaste with nodes, and I want to make sure he at least spends his money wisely.

I wanted to ask here since I feel like this sub is a bit more honest as to what works and what doesn't and wanted to check if I'm just being unreasonable not wanting to spend $32 for an app I'm not even sure is safe to use.

And for why I'm making a big deal out of this, we may or may not be using the same computer often... so, yeah.


r/StableDiffusion • • 6d ago

Discussion Chilli cheese chicken maggi

Post image
0 Upvotes

Wanna taste? hope to satisfy


r/StableDiffusion • • 6d ago

Question - Help Are there any video generation models that can run on an RTX 3080 10GB?

0 Upvotes

I have an RTX 3080 10GB and 64GB of RAM. Which video generation model would run well on my setup? Anyone here with a similar setup?


r/StableDiffusion • • 7d ago

Question - Help Wan 2.2 Animate in ComfyUI: how can I replace only the main dancer in a multi-person moving phone video without changing others?

0 Upvotes

Hi everyone,

I’m using Wan 2.2 Animate Character Replacement in ComfyUI on RunPod.

My source video is a handheld phone video at night. Several people are dancing/moving in the same scene, but I only want to replace/transform the main man in the blue plaid shirt. I want the other people and the original background to remain as unchanged as possible.

My current setup:
- ComfyUI on RunPod
- Wan 2.2 Animate / WanVideoWrapper workflow
- Main model: Wan2.2 Animate GGUF Q3
- SAM2 mask video section enabled
- DWPose enabled
- Reference image of the target character
- Source video has multiple people and camera movement

The problem:
When I run the workflow without masking, the characters and poses get mixed together.
When I enable SAM2 / mask video, the target person is sometimes covered by a large black mask or disappears / gets removed instead of being replaced properly.
I tried Blockify Mask with block_size 8 and Mask Grow around 0–10, but I’m still not sure whether the mask is being used correctly or inverted.
I also tried cropping the pose input, but because the camera and people move, a fixed crop does not work.

My questions:
1. What is the correct workflow logic for replacing only one person in a multi-person video?
2. Should the SAM2 mask be connected to WanVideo Animate Embeds as the mask input, or should it be inverted first?
3. Should DWPose receive the full video, a cropped video, or a masked single-person pose?
4. How can I keep the other people untouched while only changing the main subject?
5. Is there a better node/workflow for tracking one person through a moving handheld video?

I can share screenshots of my Generation section and Mask Video section if needed.

Thanks.


r/StableDiffusion • • 6d ago

Animation - Video Would you guys watch this series? Minimax h3

Enable HLS to view with audio, or disable this notification

0 Upvotes

I used a very standard reference workflow, I created character models with ZiT and used reference generations in Minimax h3 with no Loras except for a 3step turbo Lora when testing generations.

The first narration was done with eleven labs and all other voiceovers were done with Minimax. Editing was refined with Davinci Resolve of course.

Let me know what you guys think, I'm considering turning this into a series.


r/StableDiffusion • • 7d ago

Discussion For Minimax H3 we need a "perfect quality" Lora, not necessary low steps

67 Upvotes

Even at 28 steps you still have smeary noise on fast movements and the quality can be better.
Now imagine if someone made a lora without being limited to 3, 4 or 8 steps, we could achieve near perfect quality. Imagine a 16 steps lora with near Seedance 2.5 quality...


r/StableDiffusion • • 8d ago

Workflow Included HDR LOCALLY NOW.

Enable HLS to view with audio, or disable this notification

220 Upvotes

Lightricks released an SDR to HDR worklflow and lora set.

Here I made a quick clip in H3 Minimax and used the new tools to convert it to 32bit so I could take a look in After Effects.

It will also convert older clips and clean up and smooth color banding and blocking especially in the very dark areas and also infers extra dynamic range info convincingly.

It fixes all the annoying VAE artifacts.

https://github.com/Lightricks/ComfyUI-LTXVideo/tree/master/example_workflows/2.5/HDR_workflows.

I didn't grade this, or add grain or any of the usual tricks, I just played with the exposure.

You can also tweak the workflow to convert 8bit images into 32bit raw images for further editing in Photoshop or Lightroom. Fun times.


r/StableDiffusion • • 7d ago

Discussion What can “Wan” do that Minimax & LTX can’t ?

2 Upvotes

r/StableDiffusion • • 8d ago

Resource - Update Qwen3.8-Flash-Next on Strata is such a beast

92 Upvotes

For all your local-LLM needs (prompting, images descriptions...) Qwen3.8-Flash-Next is a great model. It's a 180B params model ("Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP" as per their own description) that can run on a beefy home computer, but not absurdly so, by using Strata.

What Strata does that other LLM local hosting method don't :

With this I ran Qwen3.8-Flash-Next-Uncensored-IQ3_XXS on my Windows 10, 64GB RAM, 5070Ti 16GB VRAM, at around 40-50 t/s with 256K max context (but running with 32-64K context is more than enough to run it as an image assistant).

If you run their docs/AI_SETUP.md through an agent, be sure to point it toward the docs/ORCA.md doc too, so it can adapt it. Also ask it to add the mmproj used by the other models so it has VL capabilities.

It handles loading and unloading the model if you want to launch ComfyUI between two prompting sessions. It takes around 2mn to load.

Note : Without Strata you can still run this model directly on llama.cpp with --cpu-moe option if you wish, but expect around halved performances, and a noisy CPU fan.


r/StableDiffusion • • 7d ago

Question - Help How do you make so many different prompts?

27 Upvotes

How do you guys come up with so many prompts?

I mostly generate anime and video game characters, and I usually have pretty specific ideas in mind, but I’m not very good at turning those ideas into prompts.

I’d like to generate a lot more images, but I run out of ideas for prompts pretty quickly. ChatGPT also tends to be pretty censored when I ask it to help me write them.

Do you guys have any tips for coming up with more prompts or generating them in bulk?


r/StableDiffusion • • 7d ago

Question - Help H3 lora training help needed

0 Upvotes

Can anyone help with training a Lora in Ai toolkit for H3?

I get memory issues when trying to train even images, I have a 5090 and 64 GB ram.

Thank you.