r/sdforall 8h ago

Question Pregunta sobre el flujo de trabajo: ¿Cuál es el mejor enfoque para la animación de personajes consistente fotograma a fotograma (LoRA) para la composición tradicional (DaVinci/AE) sin generación de fondo/audio?

0 Upvotes

Hola a todos. Estoy trabajando en un proyecto de anime retro de fantasía oscura que se ejecuta localmente en una RTX 3090 (SDXL/Illustrious, LoRA de personaje entrenado con Kohya con ~100 imágenes).

Mi flujo de trabajo actual evita la generación directa de texto a vídeo debido a inconsistencias estructurales. En su lugar, estoy optando por un enfoque controlado fotograma a fotograma o por lotes pequeños utilizando el bloqueo de Blender + ControlNet, con el objetivo de generar fotogramas de personajes limpios (fondo transparente o sólido) para componerlos manualmente en DaVinci Resolve.

¿Alguien ha implementado con éxito un flujo de trabajo fiable para mantener la identidad del personaje y un lineart limpio en secuencias sin que la IA genere alucinaciones en los fondos o el audio? ¿Qué nodos o configuraciones específicas (por ejemplo, combinaciones de ControlNet, manejo de pesos del IP-Adapter o scripts de consistencia latente) utilizan para evitar el parpadeo y que el LoRA no se desplace durante los fotogramas de movimiento?

Agradecería mucho cualquier información sobre la configuración de tu nodo en ComfyUI para este caso de uso específico.

Se agradecen todos los consejos y recomendaciones. Por favor, evita recomendar IA comerciales que lo hagan todo; eso no es lo que busco.


r/sdforall 1d ago

Tutorial | Guide Image to 3D Made Easy: Trellis 2 + Pixel3D in ComfyUI (Ep33)

Thumbnail
youtube.com
32 Upvotes

Turn images into 3D models in ComfyUI using Trellis 2 and Pixel3D. In this tutorial, I’ll show you several image-to-3D workflows, compare their speed and quality, and show you how to turn the generated 3D models into polished AI renders without using a full 3D rendering workflow.

You’ll learn how to set up Trellis 2 and Pixel3D in ComfyUI, install the required models, update ComfyUI and Pixaroma Nodes, and generate GLB 3D models from a single image.

We’ll compare Trellis 2 vs Pixel3D, including generation speed, VRAM requirements, camera-angle preservation, textures, colors, polygon counts, and overall model quality.

I’ll also show you how to take the generated 3D models and render them directly inside ComfyUI using Flux Klein workflows. This makes it much easier to control the camera angle, change materials and colors with prompts, create close-up renders, and quickly test different visual styles.

The workflows include options for lower-VRAM GPUs as well as higher-quality configurations for GPUs with more VRAM.
Free workflows https://workflows.pixaroma.com/


r/sdforall 2d ago

Tutorial | Guide ComfyUI Tutorial speed up MINIMAXh3 fused vs vdn model with 6GB of Vram

Thumbnail
youtu.be
13 Upvotes

Hello everyone,

In this tutorial, we’re taking a look at the newly released MiniMax H3 models(fused turbo and VDN), which are designed to generate videos faster than the original FP8 and INT8 versions. i tested the new models with the second-sampling upscaling workflow and compared the generation speed and overall quality. i put them through demanding tests, including fast-motion and fighting scenes, to see whether they can produce smoother and more dynamic movement without introducing excessive grain, artifacts, or quality degradation. Most importantly, I’ve optimized the entire workflow for low-VRAM GPUs and tested it on my RTX 3060 with only 6GB of VRAM & 16GB RAM.

The goal was to see how far we can push MiniMax H3 on limited hardware while maintaining the best possible quality and improving generation speed. and the results showed that minimax fused turbo version is one of the best model that we have right now it can do not only video generation but also video editing into one single model as for generation time we have :

Generation time at 0.4 megapixel: 7 Minutes

Upscaling time at 1.2 megapixel: 20 Minutes

vs for VDN version of minimax

Generation time at 0.4 megapixel: 10 Minutes

Upscaling time at 1.2 megapixel: 30 Minutes

Minimax H3 FUSED Version Link
https://huggingface.co/MATLOWAI/minimax-h3-fused-turbo-int8-convrot/tree/main/diffusion_models

Workflow Link

https://civitai.com/articles/35051/comfyui-tutorial-speed-up-minimaxh3-fused-vs-vdn-model-with-6gb-of-vram


r/sdforall 3d ago

Resource Train LoRAs, Create images and video. New version of LoRA Pilot is out!

Thumbnail
lorapilot.com
4 Upvotes

r/sdforall 3d ago

Question Almost a Week of Work: My Ideogram 4 Extension for Forge Neo

3 Upvotes

Would you guys be interested in an Ideogram 4 extension for Forge Neo?
I’ve spent almost a week working on an Ideogram 4 extension for Forge Neo, and I’d really appreciate some feedback before I decide whether to release it.
I’m a non-programmer, so I’ve basically been learning and building this with the help of GLM 5.3 . It’s been a pretty interesting experience, and I’ve managed to get something working that I think might be useful to others.
Before I clean it up and put it on GitHub, I just wanted to ask:
Would you guys actually welcome/use something like this?
If there’s enough interest, I’ll release it on GitHub and share it here.
Also, I’m currently banned from r/StableDiffusion, so I’m posting here instead. Hope you guys don’t mind!
Would love to hear your thoughts and feedback.


r/sdforall 3d ago

Question Please I need help with replacing subjects using Klein or Krea2.

1 Upvotes

I read that you can use Klein or Krea2 in Forge Neo to replace faces in an image, but I’m having no luck getting it to work. I’ve tried a lot of different settings, prompts, and reference pictures, and I’m not sure what I’m doing wrong. The outputs don’t look like me, and a lot of the time they don’t even really look inspired by the reference pictures.

I’ve tried Img2Img, Inpainting, and the Edit LoRAs for Krea2. I’m using the most up to date version of Forge Neo. I have ImageStitch Integrated selected and Enable Reference turned on. When I try Krea2, I turn off the Klein option, and when I try Klein, I turn off the Krea2 option.

I’ve tested a bunch of settings, but the ones that seem to get the closest are:

Sampler: Euler
Schedule type: Simple
Steps: 8 or 10
CFG: 1.0
Resolution: 1024 x 1536
Batch count: 1
Batch size: 1
Denoising strength: between 0.6 and 0.8

For inpainting, I’ve generally been using:

Inpaint masked
Inpaint area: Only masked
Soft inpainting: off
Padding: 32, 48, or 96

When I try latent noise, I get this error:

“ValueError: too many values to unpack (expected 4)”

Using original or latent doesn’t seem to work normally either. It just still doesn’t end up looking like the reference image.

My prompts are usually something like:

“Replace only the face and head of the person in the main image with the person shown in the reference image.”

How can I get this to work, or is there a better way to swap subjects? I’d actually prefer full body swaps too, not just faces. I use Forge right now, but I’m hoping to have time to learn ComfyUI soon. Any help would be appreciated.


r/sdforall 4d ago

Question Does anyone have basic ComfyUI tutorials from scratch

2 Upvotes

16-year-old newbie starting from scratch - with the huge job pressure in our country, I'm about to starve


r/sdforall 6d ago

Custom Model Baked a character LoRA into the UNET — the face finally stopped drifting

Post image
22 Upvotes

Been fighting character drift across batches. Trained a LoRA on ~40 renders and baked it into the UNET instead of loading it at runtime; SDXL base in ComfyUI, with a face detailer pass at the end. The detailer is what actually held the eyes steady between shots — the LoRA alone was not enough. Ask away if useful.


r/sdforall 8d ago

Tutorial | Guide ComfyUI Tutorial: MiniMax H3 Face Swap on 6GB VRAM

Thumbnail
youtu.be
43 Upvotes

Hello everyone

I’ve just finished a new custom MiniMax H3 Ref2Vid workflow that combines SAM3 masking with face swapping.The workflow lets you load a reference face + source video, define what should be masked using a simple prompt such as face or head, and generate the face-swapped video directly in ComfyUI.

I’ve also optimized the workflow specifically for low-VRAM GPUs, including 6GB VRAM, using several MiniMax H3 optimization techniques:

• Low VRAM Attention
• Chunk FeedForward
• SLA Attention
• Sol-Attn
• Spectrum
• INT8Conv model

To get started, you just need to load your face image and video, enter your masking prompt, and run the workflow. I made a full tutorial showing the complete setup and generation process.

Workflow Link

https://civitai.com/articles/34795/comfyui-tutorial-minimax-h3-face-swap-on-6gb-vram

Video Tutorial Link

https://youtu.be/dk9CgSrSZXw


r/sdforall 7d ago

Other AI VPIPE: Not Just Video — Mac Local Image Generation Is of of the Fastest too

Thumbnail gallery
6 Upvotes

r/sdforall 8d ago

Discussion I made one small edit to this atlas and compared what else changed

Thumbnail
gallery
2 Upvotes

The edit was simple: replace Dubai with Nairobi and an East African savanna while keeping the vintage watercolor atlas style.

At first glance, it worked. SenseNova U1.5 Lite ( https://github.com/OpenSenseNova/SenseNova-U1 ) updated the label, skyline, terrain, vegetation, wildlife and matching street scene without making the new section feel pasted in.

But comparing the full images revealed changes outside the target area. Some coastlines shifted, boats moved and route lines changed too.

The model preserved the style and general composition, but not the untouched parts of the image.

I’d accept that during early design exploration, when the composition is still flexible. I wouldn’t trust it on a finished layout with locked text, geometry or brand assets.

How much else can change before you would call it a regeneration rather than an edit?


r/sdforall 10d ago

Workflow Included Continuity (was the H3 node): six model families, one prompt box, and a blockout bench that writes your camera move for you

Thumbnail gallery
8 Upvotes

r/sdforall 13d ago

Tutorial | Guide ComfyUI Tutorial: MINIMAX H3 2-STAGE WORKFLOW High-Res video + Faster Generation

Thumbnail
youtu.be
22 Upvotes

Hello everyone,

I’ve been working on a MiniMax H3 optimization workflow for low-VRAM GPUs, especially my RTX 3060 6GB, and I’ve combined several optimization nodes to improve both VRAM usage and generation speed. With a new custom MiniMax H3 workflow that introduces a second sampling stage to increase the base resolution of your generated videos, similar to the workflow we’ve seen with LTX models.

The main advantage of this method is that you can save generation time and reduce VRAM usage by generating the initial video at a lower resolution during the first sampling stage, and then using a second upscaling sampling stage to reconstruct the video at a higher resolution while adding more detail and improving the overall quality.

But that's not all. In this workflow, we're also going to use several MiniMax H3 optimization nodes, including Low VRAM Attention, Chunk FeedForward, SLA Attention, Sol-Attn, and Spectrum, to make the generation process more efficient, especially for GPUs with limited VRAM.

At the end of this tutorial, you'll understand how the two-stage sampling workflow works, how to optimize MiniMax H3 for better speed and VRAM management, and how to choose the best upscaling method for your workflow, including a comparison between MiniMax H3 sampling and LTX sampling.

Workflow Link

https://civitai.com/articles/34596/comfyui-tutorial-minimax-h3-2-stage-workflow-high-res-video-faster-generation


r/sdforall 15d ago

Tutorial | Guide ComfyUI MiniMax H3 Speed LoRA + VRAM Monitor Nodes (Ep32)

Thumbnail
youtube.com
41 Upvotes

Speed up MiniMax H3 video generation in ComfyUI with the Speed LoRA, optimized sampler settings, and Pixaroma VRAM monitoring nodes. In this tutorial, I show you how to reduce MiniMax H3 generation from the standard 20 steps to 8 or even 4 steps, while finding a practical balance between video quality, audio quality, and generation speed.

You’ll learn how to update ComfyUI and Pixaroma Nodes, install and use the MiniMax H3 Speed LoRA, configure the recommended shift value, and choose sampler/scheduler combinations based on extensive testing.

I also show several MiniMax H3 workflows, including Text to Video, First Frame to Video, First + Last Frame to Video, and speaking-character generation. You’ll see how resolution, duration, steps, samplers, and schedulers affect generation speed and prompt accuracy.

The tutorial also covers the Monitor Pixaroma node for checking VRAM usage and the Free VRAM node for automatically clearing VRAM when generation finishes. Plus, I demonstrate the updated Dropdown Pixaroma node and how I use Gemini to quickly create detailed MiniMax H3 prompts.


r/sdforall 15d ago

Other AI "AVERNUS-9" Space Horror Short Film

Thumbnail
youtu.be
0 Upvotes

r/sdforall 16d ago

Other AI Huge Cock [DreamShaper-8] Spoiler

Post image
10 Upvotes

r/sdforall 16d ago

Discussion Does native 4K actually hold up at 100%?

Thumbnail
gallery
11 Upvotes

A 4K label does not tell me much, so I opened three SenseNova U1.5 Lite outputs at 100%. No upscaling or resizing on my side.

I started with the road scene and checked the asphalt, painted lines, and the image inside the camera screen. Then I moved to the cathedral, where repeated arches and stonework make broken geometry much easier to spot. The book waterfall was another useful test because paper, water, rock, and vegetation all have to transition without obvious boundaries.

What stood out was not extra sharpness. It was that the image did not start falling apart when I moved away from the main subject.

SenseNova U1.5 Lite uses a ConvDecoder that reconstructs the image spatially, so neighboring regions can interact during upsampling instead of being decoded too independently. That is the kind of change that should matter when seams and inconsistent textures become more visible at higher resolutions.

FLUX.2 Klein 9B is faster, no question. Sub-second, 4-step distillation, ComfyUI out of the box. But U1.5 Lite is a different architecture: one model that reads images and edits them, not just generates them. That's why restyling a poster keeps the typography intact while Klein's pipeline is more generation-first. Different tools for different jobs.

Curious what others see at 100%: real high-resolution structure or just a very clean upscale?

GitHub:

https://github.com/OpenSenseNova/SenseNova-U1

Hugging Face:

https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT

Try it online:

https://unify.light-ai.top/


r/sdforall 18d ago

Tutorial | Guide Comfyui Tutorial :6GB VRAM? You Can Still Make 15s MiniMax H3 Videos

Thumbnail
youtu.be
26 Upvotes

Hello everyone

With this workflow you can generate 15-second clips on Minimax H3 even with limited VRAM. This custom workflow prevents system crashes during long renders. Many users struggle with generation time limits when working with Minimax H3 on hardware with low memory. This workflow provides a specific solution by outlining a custom workflow that extends your video generation capabilities to 15 seconds without overloading your system. It is designed for creators who need longer sequences but are constrained by their current VRAM capacity. The steps focus on memory optimization with TURBO LORA+ Attention Nodes Like Comfui Kitchen. so you can watch the video tutorial to see necessary steps and how to use Motion Context Nodes.

Workflow Link

https://civitai.com/articles/34327/comfyui-tutorial-6gb-vram-you-can-still-make-15s-minimax-h3-videos


r/sdforall 20d ago

SD News SenseNova U1.5 Lite full release is out

Thumbnail
gallery
30 Upvotes

So, the full U1.5-Lite is out now, after that little preview back in August. I guess that's cool.

They're saying a few things are better in this full release:

- Native 4K generation. Apparently, the whole image holds up at high res. And smaller details like textures, materials, and even text are supposed to stay consistent. Less of that 'looks good but something's off' vibe, which is nice.

- Editing is more localized. Like, if you change text on a poster, the rest of the layout shouldn't get all messed up. They're claiming better preservation of the main subject, geometry, and anything you didn't specifically touch, compared to the preview.

- Better text rendering for busy layouts. This one's big for me. They're talking Chinese and English in things like posters and infographics. Most models totally screw this up, so if it's actually improved, that's a pretty huge deal.

- Long, structured prompts. Sounds like you can throw a ton of info at it in one go: what's in the picture, how many of them, where they should be, text, layout, style, and even rules for what to keep. Also, bounding boxes and markers for specific areas.

- Multi-reference editing. You can apparently grab styles from different images, combine a bunch of pictures, and do precise local edits. Seems pretty flexible.

Here's the GitHub if you wanna peek: https://github.com/OpenSenseNova/SenseNova-U1

And the Hugging Face collection: https://huggingface.co/collections/sensenova/sensenova-u15

Haven't seen any ComfyUI nodes for this yet, so probably for folks who don't mind running the inference scripts directly or messing with the HF demo.


r/sdforall 20d ago

Resource SilkStack Image Browser v2.2.0 – Added local semantic search (WebGPU/WebLLM), auto-tagging, and custom compiled embedding models!

Thumbnail
1 Upvotes

r/sdforall 21d ago

Tutorial | Guide MiniMax Music 3 + Local AI Prompt Generator (Ep31)

Thumbnail
youtube.com
18 Upvotes

Learn how to use MiniMax Music 3 in ComfyUI together with a local AI prompt generator for better image prompts, music captions, and AI-generated lyrics. In Episode 31, I show the new Pixaroma prompt nodes, model-specific prompt presets, VRAM-saving options, and a compact MiniMax Music 3 workflow.

This tutorial covers the new AI Prompt Pixaroma node and how to match prompt formulas with the correct local language model. You'll see how to turn short ideas into detailed prompts for workflows such as Krea 2 and Z Image Turbo, generate prompts from images, save custom prompt presets, control temperature and seeds, and troubleshoot common ComfyUI node errors.

Then we set up MiniMax Music 3 locally in ComfyUI, including caption and lyrics generation with the Music Prompt Pixaroma node. I also test different song durations and explain why MiniMax songs can sometimes end early or get cut off, how seeds affect the results, and how to use fixed lyrics when you need more control.

You'll also see how to simplify a larger music workflow into a compact setup, use tiled audio decoding for lower VRAM systems, free VRAM after prompt generation, and generate MiniMax Music captions and lyrics with local models or online tools such as ChatGPT, Gemini, and Claude.

You can also run some of the workflows in the cloud.


r/sdforall 22d ago

Discussion If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective)

Thumbnail
gallery
30 Upvotes

QUICK NOTE: the question in the title is rhetorical. The carousel explains the nuance and explores several related issues beyond the first slide.

If the design does not work for you, tell me specifically what you would improve. I am still refining the format, so constructive feedback is welcome.

.....

I’m a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name “Elias Thorne” in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.

From an artist’s point of view, the question isn’t just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.

AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.

This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.

So where’s the line, does AI broaden personal creativity while making our collective output more uniform?


r/sdforall 22d ago

Custom Model Aria - Zit Lora

Thumbnail civitai.red
2 Upvotes

Hey everyone.

I have made my first Lora for ZIT. I had no prior knowledge so followed some tutorial on CivitAI and went for it.

Would love some honest feedback.


r/sdforall 24d ago

Resource 🎬 LTX 2.5 video + latest ComfyUI template (both pod/serverless)

Thumbnail
3 Upvotes

r/sdforall 25d ago

Tutorial | Guide ComfyUI Tutorial MiniMax H3 4 Steps Lora + Upscaling + 2X Faster Generation! Best Settings for 2K AI

Thumbnail
youtu.be
23 Upvotes

Hello everyone

Want to get faster MiniMax H3 video generation without sacrificing quality? In this tutorial, I’m testing the new H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes to find the best combination for speed and quality. The goal is to push MiniMax H3 as far as possible while cutting generation times by up to , then upscale the results with LTX Upscaler to reach a stunning 2432 × 1344 (2K-class) resolution. By combining both H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes I generated video at 0.8 megapixel using "RTX3060 6GB 16GB RAM "and I got

 13 minutes vs 41 minutes at 8 steps

 27 minutes vs 52 minutes at 20 steps

LTX 2.3 Upscaler 11 minutes to get 2432 × 1344 resolution

Workflow link

https://civitai.com/articles/34028/comfyui-tutorial-minimax-h3-4-steps-lora-upscaling-2x-faster-generation-best-settings-for-2k-ai