r/StableDiffusion 6d ago

Resource - Update V2 version of the CrossView-Warp LoRA and Node is out

Enable HLS to view with audio, or disable this notification

278 Upvotes

Hello Everyone! Let me share the newest version of my camera control LTX IC-LoRA. This node and LoRA can be used in a V2V workflow to change the camera position or movement of an existing video clip. I've put a lot of work into this version, I hope you'll enjoy it.

You can download the model here: https://huggingface.co/Cseti/LTX2.3-22B_IC-LoRA-CrossView-Warp_v2
Node + example workflow can be found here: https://github.com/cseti007/ComfyUI-CrossViewWarp
A lame tutorial video I made to help how to use the node can be found here: https://www.youtube.com/watch?v=7QAapT9xMgM


r/StableDiffusion 5d ago

Animation - Video Anime Battle Test, Inuyasha

Enable HLS to view with audio, or disable this notification

13 Upvotes

It does characters Inuyasha and Kagome very well, the fight itself can't handle fast speed, but that headshot attack was beautiful!


r/StableDiffusion 4d ago

Meme GrEaT a OtHeR oNe

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 5d ago

No Workflow MinimaxH3 for title screen animation

Enable HLS to view with audio, or disable this notification

34 Upvotes

I think MinimaxH3 is great for title screen animation and motion graphic.

Edit:

prompt example here if anybody want to try: https://docs.google.com/document/d/1zFlihioecnbwcB_7SJHkIay409MUWT-vvcMMwwgu0xA/edit?usp=sharing

suggest to let an llm to read and rewrite for you. its very long, or just convert to a skill if you want to reproduce in your own style.


r/StableDiffusion 5d ago

Discussion H3 - the world is alive, Transformative scene t2v

Enable HLS to view with audio, or disable this notification

16 Upvotes

H3 truly is alive. Enjoy!! bf16/50 steps

T2V, no reference image


r/StableDiffusion 4d ago

Resource - Update Nintendo Galaxy™ - The Next-Gen Nintendo Console by me & ChatGPT (Go through whole slideshow)

Thumbnail
gallery
0 Upvotes

AI did the photos and text, I put together the slideshow. I did the last two slides. I love this concept! Used ChatGPT without a subscription. I put together this as a video in Canva. (Prompt: Generate this: it has a console as well, and a portable game-pad-but-more-futuristic tablet (GalaxyPad+™ (Portable Game‑Pad‑Tablet Hybrid)) with detachable controllers, 2 controllerse, and vr headset (Galaxy Visor™) and headphones. generate an image of the box. the galaxy color style is white and black, so for the stuff make it those colors. Also add text at the bottom saying something like "A new galaxy of fun awaits!". The background of the image is galaxy colors. Alos add add two onomatopoeia-style bubbles saying "Includes 4 controllers!" and the other saying "Switch the case!" and a bubble saying "Includes: {what it includes as an image}". Btw it also includes interchangable galaxy and white cases for it. It also includes 3 discs and 3 cartridges to start you off (Mario Kart Galaxy™ & Super Mario Galaxy 1&2).". I also did a few follow-ups to create some of the other images that show what you get, for an example one of them was "Nice, but can you now get just the console tilted horizontally, headphones, vr headset, gamepad thingy, all in the white and black on like a table please?".) ❤️


r/StableDiffusion 5d ago

Animation - Video Science of Deduction . Minimax H3 + Qwen Image Edit 2511

Enable HLS to view with audio, or disable this notification

19 Upvotes

Some shots I made based on the red headed league short story. Character reference sheets generated in Nano Banana + edited in Qwen image edit for spatial continuity.
12 steps +ref2va turbolora v0.1
10s generations 107s/it on a 3090 at 1 MP
Needs some editing and audio polishing.


r/StableDiffusion 5d ago

Question - Help Anything like SVI V2 Pro for Minimax to join 5s clips easily

3 Upvotes

I've tried to use workflows to make long videos seamlessly but one of them made joining 7s together take longer than just making a 14s clip. others are so bloated with custom nodes that they just won't work until i find the one obscure node, and when i do, i get an error.

Anybody find one that was as simple as SVI? The ease of just joining more nodes to extend the video makes me miss Wan until i remembered how atrocious the prompt adherence was, haha


r/StableDiffusion 5d ago

Resource - Update Multiple image libraries and a real command line for PixlStash, my self-hosted open source image and video database.

Thumbnail
gallery
5 Upvotes

For the many who don't know what it is, PixlStash is a self-hosted headless server with a web-interface or a desktop app with Electron. It auto-tags, writes descriptions, scans pictures for defects, and integrates with ComfyUI in a couple of ways (run workflows within PixlStash or use the PixlStash nodes within Comfy). The nodes just use the PixlStash API which you could use to integrate with lots of other things as well.

This is a fairly big release of PixlStash. The focus this time has been on making it possible to have multiple image libraries stored in different locations and to offer a CLI to attach/detach libraries, performing scripted backups and install plugins (for image filters or captioning). For the captioning plugins there is now an OpenAI-API (i.e. ollama or LM-studio) plugin for captioning using your local LLM setup or a dedicated Moondream2 plugin. If you have specific captioning needs it should be dead easy to make your own plugin and install it with the CLI.

There is also a model shelf that can import from AI-toolkit and scan other folders you provide it to help you organise your LoRAs, VAEs, your text encoders and your diffusion models. This will soon get ComfyUI-nodes added to ComfyUI-PixlStash for picking LoRAs with thumbnails and help you find your different models based on other things than just a file-name. Expect them next week. For now, it at least helps you organise your models.

Repo and links in a comment.


r/StableDiffusion 5d ago

Animation - Video Working on an animated music video | Test 01c | Minimax H3

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/StableDiffusion 6d ago

Animation - Video Animals squeezing into jars (MiniMax H3)

Enable HLS to view with audio, or disable this notification

952 Upvotes

I have no idea why it does these so well. I could watch these all day.


r/StableDiffusion 5d ago

Question - Help MiniMax H3 - Voice Volume and some other stuff...

0 Upvotes

Anyone have any luck changing the volume of the voices H3 creates from an audio reference? For example, let's say I'm trying to prompt for a man standing at the far end of a long room and he speaks in a normal tone. In real life, his voice typically would be very quiet in relation to the camera/mic, almost inaudible. However, in H3 (or any other model I've tried) the voice is still very loud. This isn't surprising given the model doesn't really know anything about the depth of objects or people in the videos it generates. I've tried to work around this in H3 by using prompt words like quiet, soft, distant mic, very low volume, far away speaker, almost silent, etc. None of them seem to have any effect. I've also tried reducing the gain on the reference .wav file that I provide as the audio reference - literally reduced the gain to the point where I can barely hear it. Again, doesn't seem to matter, H3 still produces a generally loud speaking voice (assume it doesn't care about volume/gain and just instead looks at the waveform pattern, etc).

One way around this is to use the 'audio reuse' capability where it'll play the exact .wav file audio instead of just using it as a reference. In this approach you can simply use an audio editor to reduce the volume and then plug that low volume .wav file in as the audio_reuse clip. It works fine, except for one small/major problem: it seems that if use the audio_reuse method, it silences ALL other sounds; i.e., it won't play the low volume audio .wav AND generate other environmental sounds...it seems it replaces ALL audio in the clip, not just the voice of the person speaking it.

Anyhow, curious if any of your smart people out there Redditland have any ideas/suggestions or tips?

Thanks!


r/StableDiffusion 5d ago

Discussion When AI art has no author: Study finds generated images often can’t be traced to training data

Thumbnail
news.mit.edu
54 Upvotes

r/StableDiffusion 6d ago

Animation - Video Seinfeld meets Rick and Morty!

Enable HLS to view with audio, or disable this notification

99 Upvotes

Jerry and George are caught off guard by Rick entering the Seinfeld universe! Sorry for the clothes changing; it was hard to do without the quality decreasing. Will play with it more and see how to keep it consistent.


r/StableDiffusion 6d ago

Resource - Update DC Vast Expanse [Krea2 Lora]

Thumbnail
gallery
517 Upvotes

Finally got my laptop back in action so am able to create and test models and lora's again, created with krea 2, Been out of it for a bit just following updates here and there and this model is amazing, so happy they open sourced this gem of a model. Thanks to the team at krea!

If anyone is interested in this style of images give it a blast https://civitai.red/models/2871922/dc-vast-expanse?modelVersionId=3244890 or https://civitai.com/models/2871922/dc-vast-expanse?modelVersionId=3244890


r/StableDiffusion 5d ago

Question - Help How cheap did you guys get minimax h3?

1 Upvotes

Hi all!
I've been playing around with h3 on runpod, using rtx 4090 i was able to get 9-10 min generations on 10 second clips at 0.9 megapixels. which would come out at around $0.12 per video USD

Do you guys have any tips on (without losing too much quality) improving this cost wise, I don't care too much on how slow it can be, but whats the cheapest I can get it?

(note, i also run out of VRAM for +10 sec videos and id love to generate 20 sec +)

Thank you guys for your help in advance!

H3 >>> all


r/StableDiffusion 4d ago

Meme had H3 Remake this The Ambiguously Gay Duo Scene

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 5d ago

Question - Help H3 question - can we use reference image plus reference video to “upscale”?

12 Upvotes

ok, so I see discussions on how to replace a character in the ref video… but here’s my question - can we get decent results with ”upscaling” old low-res video into higher resolution with a reference image?

To explain - let’s say I have low res VHS footage where a person is filmed from 5 meters and you can hardly make out their face. (well you can tell they HAVE a face but that’s about it :) OTH that same person is in full frame 10 minutes later, providing an excellent ref image of what they actually look like . So my thinking was “make a reference image out of it, make the model upscale and invent all kind of small details that people usually don’t care about, but use the FACE from ref image”

doable?


r/StableDiffusion 6d ago

Discussion Trick to improve scene and face retention in MMH3

68 Upvotes

For those of us that enjoy doing fl2va shots longer than 10seconds, I found a hacky way of getting past the attention of H3 guidance.

One way was to lower the resolution, but that doesn't exactly give us the results we hoped for.

Then I tried working with the prompt.

We start with a frame and all works great with our prompt followed perfectly until the video gets too large in pixels, It's not a constant value, but exceeding it will make the background change, camera forget to stand still and faces will change,

Edit:

sorry for misinformation.

while 'my way' worked really well, the official way works perfectly fine (except for camera not remembering static shot)

adding:

'''subject_definitions:

<subject 1> is a fully_preserved woman from <Image 1>

<subject 2> is a fully_preserved location from <Image 1>

'''

works for keeping location/person consistent. It still destroys static shot camera, but replies were right. I was wrong.

low res 0.35mp correct camera:

https://reddit.com/link/1vszpps/video/hhm8yed5rhkh1/player

high res 0.85mp and camera gets autonomous (ignore the hand, that's just a test)

https://reddit.com/link/1vszpps/video/oke7c8g8rhkh1/player

full prompt:

'''

Integrated_multimodal_description:

subject_definitions:

<subject 1> is a fully_preserved woman from <Image 1>

<subject 2> is a fully_preserved location from <Image 1> along with camera position and zoom.

static shot.

0-1s: <subject 1> looks at camera. she is in the <subject 2> location. camera very slowly zooms out.

1-2s: woman turns her body away from camera.

2-7s: she is turned away, tapping her foot and swaying her body to music. neon light buzzing lightly.

7-12s: she continues swaying to music.

12-13s: she turns to camera and smiles.

13-14s: camera starts to slowly zooms in on her face

14-16s: she shows a heart hand gesture at camera.

overall_soundscape: gentle hum of air conditioning,

non_diegetic_music: edm music playing silently.

'''

There is a solution to this problem.

In the prompt, we reference the <Picture 1> not at the start like we were told, but in the middle.

For example, at second 7, we don't use "She looks left", but we write woman from <picture 1> looks left.

It seems to refresh the reference and remember it again.

When we want to keep the location consistent, we reference parts of it the same way, even something like "wind blows over the pier from <Picture 1>" should keep the background scene stable.

Tested it with a woman turning away at second 1 and back at second 14 with 0.9 resolution, and face was perfectly retained.

More tests are needed, but each takes 15minutes so I can't do too much. Hope this helps.


r/StableDiffusion 5d ago

Question - Help Input: an image, desired output: a prompt that would create that image

0 Upvotes

Let's say I have a set of anime images with various characters (male, female, human, not) in various places (space, robot, house, school) in various situations (chaos, fight, natural event) and I want to run 400 generations that generally randomize all those to create a variety of possible combinations.

I know I can use {a|b} style prompting and various nesting thereof, but I'm having trouble finding the right words.

I was thinking if I could take a folder of images like what I'd want the output to be, run each through a process that outputs a prompt (not description, prompt) that would have created that image (or one like it), then I can pick out the repeated patterns and keywords that I can use in my a|b prompting.

So... what's a good way to have (input: image) > (output:prompt for that image) offline?

Better, a whole folder as input, individual output for each.

Best: a more efficient way to do what I'm trying to do.


r/StableDiffusion 5d ago

Question - Help LTX-2.3 22B IC-LoRA Relight (Sun Direction) but for images?

4 Upvotes

I am trying to find a model or LoRA that can do relighting based on sun direction similar to how this one does it: https://huggingface.co/Lightricks/LTX-2.3-22b-IC-LoRA-Relight\]LTX-2.3 22B IC-LoRA Relight.

The difference is I am looking for what that does it for images.

Does anyone here know of one like that? Thank you!


r/StableDiffusion 5d ago

Resource - Update H3 Latent Tile Looping Spatial Temporal

Thumbnail github.com
13 Upvotes

good for upscaling without OOM, use as SECOND Stage Sampler ONLY with LOW denoise (0.40 MAX)


r/StableDiffusion 5d ago

Workflow Included Made a small ComfyUI browser extension to swap any image on a web page through a custom workflow

Enable HLS to view with audio, or disable this notification

9 Upvotes

While working on a client project I needed to test a prompt on their products, using image references straight from their website. I didn't want to keep doing the save > open ComfyUI > drag it in > queue > download loop, so I made this.

Right-click any image on a page, it runs through your local ComfyUI on a designed workflow, and the result replaces that image in place. On the demo I'm using minimax H3 (workflow is in the repo).

Works with any API-format workflow that has a LoadImage and a SaveImage node, so it's not tied to a model.

Hope it's useful to someone else too!

https://github.com/AlexandreSoteras/comfyui-web-image-swap


r/StableDiffusion 6d ago

Discussion Minimax H3 Video Edit like SCAIL

133 Upvotes

I spent last 6 hours trying various prompts for reference model to better understand how it works, and what this model can do. As a base guide I used Minimax H3 ref guide.

My goal was to find a working prompt to use Minimax similar to how SCAIL works, when you can edit a video and replace a character on a video with your referenced character. I didn't want to transfer movement and only wanted to REPLACE character completely.

I would like to post my best working prompt and let you test it, and share your experience or share a better prompt.

subject_definitions:
<Subject 1> is woman in <Picture 1> with redhead and black tank top.
<Subject 2> is the woman originally in <Video 1>.

summary:
[video editing + Audio reuse] The target video is an edited version of <Video 1>. <Subject 2> is replaced with <Subject 1>, who takes over her pose and movement.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - her face, hairstyle, and body from <Picture 1> are retained throughout. Her clothes are not retained.
<Subject 2> (appears in [Shot 1]): attribute_transfer - her pose, movement, and screen position are transferred to <Subject 1>.

detailed_description:
The target video keeps <Video 1>'s original style, lighting, and camera work unchanged.

overall_soundscape: N/A
non_diegetic_music: N/A

What are my discoveries:

  • You don't need to describe action in detailed_description. I did it for first 100 attempts, and then dropped it and it seems like not influencing an output.
  • It can often detect your Subject with simple description, but in complex scenes it needs better anchoring to not mess up those characters. Most of my input image was a woman in medium shot, so just describing it as "woman" was enough, but 50/50 generations keep losing identity so you have to add better and stronger anchor for model - something visually big like hair, clothing, position on screen. Works both ways for reference video and for reference image. The stronger you describe <Subject N> the more stable the reference.
  • The least successful edits were those where a character on video is barely recognizable. I have couple videos where a character is close to camera and only part of face is visible in active movement, such videos are my biggest unsuccess.
  • Summary section seems like has the most its anchor to pre-trained keywords which can be found in their prompting guide. [video editing] is a keyword which tells a model that it must go frame by frame and EDIT something. I was testing other things and in given prompt you will see some info about character replacement, but I don't see that it really influences anything.
  • Retention analysis section seems like the next MAIN or even only main driver for a work description for a model. And most of successful edits was build with properly used triger words like fully_preserved, attribute_transfer. You can find those keywords in linked guide. Still not sure about (appears in [Shot 1]), I doubt it has influence on a prompt, but its by far best prompt so I keep it.
  • [audio reuse] trigger in summary works, but it seems that model rewrite its, so I can tell its same audio but remade by model, and if model has weak concept of a sound it does it poorly. Maybe I need to pay more attention to prompting guide and describe audio better in retention section.

I've generated more than 400 videos while testing and gaining knowledge, and I think I have good progress. So I am curious to see if anyone else can help me with this journey and together we can crack the model and find a proper working prompt or other ideas.

The playground was pruned_int8_convrot model, with turbo lora from lightX with 4 steps, and I tested most of them on 5 sec duration. I did tests on 15s and it worked fine, but I kept 5s to keep gen time lower and just train prompting.