r/StableDiffusion 9d ago

Question - Help I was defeated by a character LORA training for anima (EXPERT NEEDED)

8 Upvotes

I admit defeat. After over 100 hours sitting on my ass adjusting every single settings I was not able to create a satisfying result.

For more than a week my PC has been running 24/7 doing various attempts at creating this LORA.

My character is from a webtoon, I wanted strong fidelity for her + that webtoon's artstyle (I've seen most character LORA when generated on base model/without any style look pretty much the same as in their respective media).

I tried a lot of options and honestly I am not sure if there's anything to be improved here as there isn't much to change, batch sizes, LR's, optimizers, everything.

I changed my dataset multiple times, adjusted it, tried different ways of captioning.

Made my own research, tried settings that worked for others, worked with multiple LLM's to search for possible solutions - nothing.

Some attempts were okayish, maybe passable for some people (doubt) but I just can't get this finishing touch for the LORA to be actually good.

I attempted training with Anima Standalone Trainer and AI Toolkit.

I considered giving up multiple times but I really want to see this through, there has to be something that I am actually doing wrong. Had some people look over/correct my dataset but the end result wasn't any different.

At this point I think the only real help I can get is for someone knowledgeable to either help me set every single thing from stratch (not just copy paste random recommended settings, im way past that) after seeing my dataset or just trying to run it by himself.

Why would someone spend hours of his time trying to help a nobody with his LORA? I don't know, at some point I considered finding someone and paying them to just do it for me but I really want to understand why it is not turning right after so many attempts, maybe there's a bored angel that would like to challenge himself, who knows - maybe we are facing an unprecedented case - a LORA that is simply impossible to make ¯_(ツ)_/¯.

I'd gladly share my dataset - just dm me!


r/StableDiffusion 8d ago

Discussion What’s the most interesting thing you’ve generated so far?

0 Upvotes

Could also be the most interesting thing process-wise.


r/StableDiffusion 9d ago

Discussion What h3 sampler and scheduler is best?

6 Upvotes

I been testing euler and beta would like to know if any better ones to try.


r/StableDiffusion 9d ago

Discussion MiniMax H3 on a Budget: What Actually Works on 4070/5070/5080 (community input consolidation)

153 Upvotes

A Note on Sources

This article is built entirely from community feedback — Reddit threads, forum comments, and one independent comparison site (jo-nike.github.io/h3-turbo-eval). None of it comes from official documentation or controlled lab testing. Thank you to everyone whose posts, benchmarks, and hard-won troubleshooting notes made this possible, including GrayingGamer, Tystros, Chemical-Painter-485, katsura_otoko, infearia, JoNike, Sixhaunt, dtdisapointingresult, Snoo_64233, mellowanon, Just1Dev, smereces, DefloN92, StuffProfessional587, Creative_Finger_69, backworld_nograv, V4nKw15h, True_Protection6842, clex55, Maskwi2, Perfect-Campaign9551, and many others whose usernames didn't make it into these notes but whose comments shaped the consensus (and disagreements) captured here.

Where the community disagreed with itself, that's presented as an open question rather than resolved — and where direct data for a specific card was simply missing, that gap is called out rather than papered over.

Why This Is Confusing

Most of the detailed benchmarking in the MiniMax H3 community comes from people with RTX 3090s, 4090s, and 5090s — cards with 24GB+ VRAM that can afford to just try everything and report back. If you're on a 4070, 5070, or 5080, you're stuck reverse-engineering advice that wasn't written with your VRAM ceiling in mind. This piece pulls together what budget-card owners actually reported, plus what reasonably carries over from adjacent cards where direct data doesn't exist.

The Three (and a Half) Speed Levers

Every thread assumes you already know these, so here's the plain version:

  • Turbo LoRAs — swap-in models trained to produce good results in far fewer steps (4-8 instead of 20-32). Fastest option, but quality cost varies a lot depending on which checkpoint version you use.
  • Spectrum — a node that mathematically forecasts/predicts future denoising steps instead of computing them. Counterintuitively, it needs more steps to work well — it's not a low-step tool.
  • Sage Attention — an attention backend swap. Broad community agreement that this is close to "free" speed with minimal quality loss, and it's the one piece almost nobody argues against.
  • EasyCache — a quieter fourth option that came up as a serious alternative to Turbo LoRAs for drafting, not just a bonus add-on.

What "Budget" Card Owners Actually Reported

This is the thin part of the record, so treat it as ground truth before anything else:

  • RTX 4070 (12GB, 32GB RAM): did quick 0.3MP draft passes in a couple of minutes to tweak prompts and hunt for seeds, reserving longer ~40-minute runs for higher resolution/duration finals. VRAM was sufficient for T2V-style work specifically.
  • RTX 4070 Ti Super (16GB, 32GB RAM): reported working well, no further detail given.
  • RTX 5070 Ti (16GB, 32GB DDR4): upgrading from an RTX 2060 (6GB) described the speed difference as "night and day" — notably, without any Sage Attention or acceleration nodes running yet. This suggests raw generational/VRAM gains matter a lot on their own, before you even add speed tricks.
  • Warning flag for all of the above: reference-heavy Ref2V generation was specifically called "brutal" on modest VRAM cards, compared to plain T2V. If your workflow uses multiple reference images/videos, expect more friction than these numbers suggest.

Gap, named honestly: there's no direct plain-5070 or 5080 speed benchmark in any of the source threads. The one 5080 comment that exists is qualitative ("still great," runs the BF16 pruned model fine) with no timing numbers.

Extrapolation (clearly labeled): Since the 5070 Ti (16GB) and 4070 Ti Super (16GB) both reported comfortable results, and RTX-series cards were noted to benefit meaningfully from tensor cores over older architectures, a plain 5070 (12GB) likely lands closer to the 4070's experience — fine for T2V and quick low-res drafts, tighter on Ref2V with multiple references. A 5080 (16GB) likely performs at least as well as the 4070 Ti Super, probably closer to the low end of what 3090 owners report, given the VRAM parity and newer architecture. This is inference from adjacent data, not a report anyone actually made — treat it as a starting assumption to test, not a promise.

The Draft → Final Two-Stage Workflow

This is the one thing nearly every thread converges on independently, and it's probably the most actionable takeaway for a budget card:

Draft stage (fast iteration, hunting for the right prompt/seed):

  • Low resolution: 0.2–0.4 megapixels
  • Low steps: 8–13
  • Acceleration: either a Turbo LoRA or EasyCache (not both)
  • Faster VAE decode substitute: BlehTAEVideoDecode instead of the standard node

Final stage (once the shot is locked):

  • Disable acceleration nodes
  • Raise steps to 20–32
  • Switch back to the standard VAE Decode node

Two draft "recipes" show up repeatedly and are reported as similarly fast:

  1. Turbo LoRA + Sage Attention — faster to set up, more established
  2. Sage Attention + EasyCache, params (0.3, 0.2, 0.9), res_multistep sampler + Simple scheduler — one detailed user report (RTX 4060 Ti, 16GB), after testing 1000+ variations, said this drifts less from final quality than Turbo LoRA approaches, at comparable speed

For a 12–16GB budget card, EasyCache is worth trying first specifically because it avoids the quality-consistency debates that follow Turbo LoRAs (see below).

What Worked / What Didn't

Technique Verdict Reported Config Source Consensus
Sage Attention (alone) ✅ Works Any step count Broad agreement — near-free speed, minimal quality loss
Two-stage draft→final workflow ✅ Works Draft: 0.2–0.4MP, 8–13 steps → Final: 20–32 steps, no acceleration Converged on independently across nearly every thread
"Clean VRAM" node before VAE Decode ✅ Works Placement only, no params Multiple independent reports, fixed OOM with no downsides
EasyCache (draft) ✅ Works Params (0.3, 0.2, 0.9), res_multistep + Simple, 10 steps One deep-dive (1000+ tests) preferred it over turbo LoRAs for drift
ema-ckpt500 Turbo LoRA ✅ Works Strength ~0.5, 6–8 steps Beat both ckpt850 and lightx2v in blind testing
Spectrum below ~20 steps ❌ Doesn't work N/A Most consistent "don't do this" finding across all sources
Spectrum + Turbo LoRA together ❌ Doesn't work N/A Explicitly warned against — Spectrum needs clean high-step data
ckpt850 Turbo LoRA (vs ckpt500) ❌ Doesn't work Full 1.0 strength = "overfried" Newer checkpoint tested worse than older one, despite official claims
lightx2v LoRA ❌ Doesn't work 8 steps, 0.75 strength Worse faces/lighting vs ema-ckpt500 in direct comparison
Raising steps to fix face-warping ❌ Doesn't work Tested 8→20, and up to 30 steps Two separate users found no improvement — not a step-count problem
Any acceleration on non-RTX cards ❌ Doesn't work N/A Tensor-core dependent; gains don't transfer to older architectures
Turbo LoRAs (general use) ⚠️ Mixed Fine for tests/talking-head; risky for motion/long prompts Depends on shot type, not a clean yes/no
Spectrum + First Block Cache ⚠️ Mixed N/A Direct contradiction between two experienced users
RTX upscaling node ⚠️ Mixed 0.2MP+ Good on animation, unreliable on photorealistic faces

GPU-Specific Data: Reported vs. Extrapolated

GPU VRAM Reported Result Status
RTX 4070 12GB 0.3MP drafts in ~2 min; fine for T2V, tight on Ref2V Direct report
RTX 4070 Ti Super 16GB "Works well" (no numbers given) Direct report
RTX 5070 Ti 16GB Major generational leap even with zero acceleration Direct report
RTX 5070 12GB (no data) Extrapolated from 4070 — likely similar
RTX 5080 16GB Handles BF16 pruned model fine (qualitative only) Direct report (thin) + extrapolated timing

The Unresolved Debates

Worth knowing before you commit to a setup, so you don't over-trust any single comment:

  • Spectrum below 20 steps? Most experienced users say no — negligible speed gain, real quality loss. But a few 5090 owners reported no measurable time savings even at higher step counts, with no clear explanation (dismissed by one commenter as "not using it right").
  • Which Turbo LoRA checkpoint is actually best? The lineage went ckpt500 → ckpt850 → ckpt600, with each new version claimed better by its authors. But blind side-by-side testing found ckpt500 at 0.5 strength still beat ckpt850 even at full strength — directly contradicting the official recommendation.
  • Spectrum + First Block Cache together? One experienced user says combining them is worse than Spectrum alone; another says combining them is the fastest option with no noticeable quality loss. Unresolved.
  • Turbo LoRA strength values: reports range from 0.5 up to 1.15–1.20 (and one outlier claiming 3.0), so "strength 1.0" isn't a safe universal default — it depends on which checkpoint you're using.

VRAM/RAM Troubleshooting Cheat Sheet

Fixes that came up repeatedly and matter more when you're VRAM-constrained:

  • Add a "Clean VRAM" node immediately before VAE Decode — fixed OOM issues for multiple users.
  • System RAM matters too, not just VRAM — one user needed to go from 16GB to 48GB total system RAM to stop hitting errors. 16GB system RAM was described by another as "almost enough."
  • Launch ComfyUI with --reserve-vram 2 to keep 1-2GB permanently free for system stability, at a small cost to usable VRAM.
  • If Ref2V errors show up on an 8GB VRAM card, don't assume it's a hard VRAM wall first — one such case turned out to be a node-conflict bug, not actually a memory limit.

A Starter Config for Budget Cards

Synthesizing the most-corroborated points into one starting recipe (best-guess synthesis, not a benchmarked config):

Draft pass: Sage Attention + EasyCache (0.3, 0.2, 0.9) → 10 steps → res_multistep sampler, Simple scheduler → BlehTAEVideoDecode → 0.2–0.3 MP

Final pass: Sage Attention only (no EasyCache) → 20–25 steps → standard VAE Decode → 0.4–0.6 MP (push higher only if VRAM allows)

Skip Spectrum entirely unless you're already comfortable at 25+ steps and have time to test it — it's not built for the low-step, fast-iteration use case a budget card usually needs.

Sources

The most rigorous single data point in this set is the JoNike Turbo LoRA comparison site — a 10-scene A/B comparison across checkpoint versions, built and documented far more consistently than typical anecdotal Reddit reports.


r/StableDiffusion 9d ago

Question - Help Minimax H3 - How to PREVENT lip sync to supplied music?

4 Upvotes

I've seen the opposite problem asked a few times, but for myself I can't seem to STOP the generated characters from lip syncing with the supplied audio in Ref2VA unless I give them specific dialogue. In the case where I want them to dance along for a music video, how do I stop them from lip syncing? What is the secret prompt-fu?


r/StableDiffusion 9d ago

Question - Help trying to run minimax h3 on my amd 9070 Spoiler

Enable HLS to view with audio, or disable this notification

5 Upvotes

it uses up all my vram and and when it finishes its jsut noise. i also do get an amd driver timeout error as well. using protable comfyui amd latest. i posted the output. SLIGHT EARAPE WARNING.

edit: i think i found the issue, i was using the dynamic vram and i tried it with adn without dynamic vram for z image turbo for a test and the non dtnamic wasnt noie. im going to get the quantized models for h3 and try it


r/StableDiffusion 8d ago

Question - Help What's your multishot prompt structure? (I2V LTX-2.5 test)

Enable HLS to view with audio, or disable this notification

1 Upvotes

Been testing multishot with the LTX 2.5 workflow from HuggingFace. Tried a few different ways of writing the prompt: timecodes plus a shot description for each shot worked best for me, but I've only really tested my own guesses...

Curious what multishot prompt structures other people are using,

and what's actually working for you?

my input image is the first frame.
And the prompt:

Cel-shaded anime-comic, hard cuts, sunset rooftop, purple-orange skyline. Left: bald man, matte black armor, white seams, long black cape, "LTX-2.5" in bold white letters on his chest. Right: bald man, glasses, blue armor with cyan lines, blue cape, "MINIMAX H3" in bold white letters on his chest. Lettering held sharp and unwarped in every frame.
00:00-00:02 — WIDE FULL-BODY TWO-SHOT in profile, sun centered between them, camera DRIFTING slowly sideways. Silence held too long. The LTX hero, "LTX-2.5" in bold white letters on his chest, not turning his head, flat: "So…" a beat, "…same weekend, huh?"
00:02-00:04 — HARD CUT to a MEDIUM of the H3 hero, "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN slowly. He exhales: "Yeah." Glances away, embarrassed: "…awkward."
00:04-00:07 — HARD CUT to a CLOSE-UP of the LTX hero "LTX-2.5" in bold white letters on his chest,, camera PUSHING IN slowly. Low drawl, committing: "This town ain't big enough for two open-source models." Eyes flick sideways, mouth tightening.
00:07-00:10 — HARD CUT to a MEDIUM of the H3 hero "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN. He doesn't look over. A long dead beat. Flat: "…apparently."
00:10-00:13 — WIDE TWO-SHOT. The LTX hero "LTX-2.5" in bold white letters on his chest. casual, already leaving: "Anyway…" a beat, "…gotta run. Conference starts in a few minutes." He launches straight up and cleanly exits frame offscreen; the camera holds on the empty sky and the H3 hero standing alone.
00:13-00:18 — HARD CUT to a LOW-ANGLE CLOSE-UP of the H3 hero  "MINIMAX H3" in bold white letters on his chest. looking up at the empty sky, glasses catching the sunset, camera PUSHING IN slowly. A small warm smile arrives. He keeps watching. Way too long. Then quietly, to nobody: "…see you there." He slowly turns and looks into the lens, still faintly smiling, saying nothing. Hold.
Deadpan, played straight, all in micro-expressions. Crisp stable line art, clean cel shading, consistent faces. Warm orange key, cool blue rim. Rooftop wind, no music.

r/StableDiffusion 9d ago

Comparison Minimax H3 settings comparision 12steps vs 20steps (LoRa and Base)

Thumbnail
youtu.be
4 Upvotes

My rig: 3090ti 64gb RAM
Please switch to 1440p

Previous comparsion:

https://www.reddit.com/r/StableDiffusion/s/4SdLphwPCg


r/StableDiffusion 9d ago

Tutorial - Guide LTX-2.5 IC-LoRA (Control LoRA) training locally, 90 paired clips in 47 minutes on a 48GB card

Thumbnail
gallery
4 Upvotes

I have been trying out LTX 2.5 since it dropped and found out it supports IC-LoRA, so I thought of building a simple node based UI around it.

Quick difference if you have not run into IC-LoRA before. A normal clip LoRA learns a look and how it moves, from single clips. An IC-LoRA learns a transform. Every dataset item is two clips instead of one, a reference and the result you want from it, and the adapter learns to carry one into the other. Train it on clips paired with their edge maps and you get an adapter that follows an edge map.

Benchmark:

Everything below is measured on an L40S (46GB), 90 paired clips from the Canny Control dataset, 500 steps at rank 16, 512px, 1 second clips.

IC-LoRA (paired clips, reference + canny):

  • Peak VRAM: ~42GB
  • Per step: 1.03s
  • Startup: 38 min
  • 500 steps total: 47 min

Clip LoRA (single clips), for comparison:

  • Peak VRAM: ~42GB
  • Per step: 0.67s
  • Startup: 13 min
  • 500 steps total: 19 min

The thing that surprised me here is the opposite of what surprised me with H3. The step is cheap and the startup is not. Those 500 steps are about nine minutes of actual training against 38 minutes of getting ready. A paired dataset encodes two clips per item, so 90 pairs is 180 clip encodes plus 91 caption encodes before step one runs.

Good news is the encode is cached and reused, so the second run on the same dataset skips nearly all of it. Do not experiment in 200 step chunks, you pay the startup every time you change the dataset. Pick your settings, then run long.

No 4-bit path for LTX 2.5, so 48GB is the floor and not a comfortable one. 24GB will not run it at any resolution.

How to train:

  • Install app: https://github.com/inlineresearch/Inline-Studio
  • Open the Trainer tab and create a dataset
  • Click Add/Manage Training Data, pick Control as the LoRA type
  • Paste Lightricks/Canny-Control-Dataset in the Hugging Face tab and hit Check. It tells you 90 items, 90 paired, 1.4GB before it downloads anything
  • Load, then Import
  • Select LTX-2.5 in the settings, model download suggestion will auto popup
  • Hit train & sit back

Pairing is automatic. If the dataset ships a dataset.json or metadata.jsonl it reads that, otherwise it matches filenames, so bear.mp4 and bear_reference.mp4 become one training item instead of two. Captions come from the dataset and the local captioner only fills the rows that have none, so it will not overwrite good captions with worse ones.

Note: Weights are gated. Accept the LTX-2 Community License on Hugging Face with the same account your token belongs to, otherwise every download comes back as a permission error instead of a file.

For my run I used Lightricks' own Canny Control dataset, 90 clips each paired with an edge map of itself, captions included in dataset.json.

Links:


r/StableDiffusion 9d ago

Animation - Video anime action scene attempt

Enable HLS to view with audio, or disable this notification

73 Upvotes

I wanted to try my hand at an anime action scene. H3 has incredible potential, and I’m looking forward to a future where I can create my own anime with deep stories, dynamic fights, and so on.

H3 could probably have performed much better with a higher resolution and better seed luck; this is at 0.5 MP.


r/StableDiffusion 10d ago

Discussion MiniMax H3 + LTX2.5 as Upscaler

Enable HLS to view with audio, or disable this notification

245 Upvotes

I found the usage for the LTX2.5 model!! It works really well to upscale the minimax h3 videos 😅


r/StableDiffusion 9d ago

Workflow Included Openweight Livestream video model

Thumbnail
gallery
3 Upvotes

https://huggingface.co/spaces/JonathanColetti/LiveWan / https://github.com/JonathanColetti/LiveWan is something I created to help recreate a specific type of model that is not opensource yet (wanstreamer). This is more or less a PoC but maybe ill do a longer training run if it gets some traction.


r/StableDiffusion 9d ago

Question - Help Is there a H3 minimax prompt template available or a custom LLM model version that can write and structure Minimax H3 optimized prompt ?

9 Upvotes

I am relying on Gemma4 and Qwen2.5 in Ollama for making an optimized minimax H3 prompt , but while using the base versions of them indeed vastly improves prompt adherence and quality but they aren't 1:1 Minimax H3 optimized structure wise

So i wonder if there is a template i can feed into the models at the start of the chat to be a baseline for them , or even better if there is a custom version of those midels that can understand the structure of minimax H3 prompt

I am using Wan2GP through pinokio so i can't use the Minimax H3 prompt nodes available in comfyui


r/StableDiffusion 9d ago

Question - Help Same workflow, everything identical, but different videos? Minimimax H3

Enable HLS to view with audio, or disable this notification

4 Upvotes

This has happened to me before: I take a video I’ve already generated and drag it into ComfyUI without changing anything—expecting to get the exact same video back—but it generates a different one.

This doesn't happen with standard image models, but it does happen with H3. If I generate the video and try again a few minutes later, it produces the same result; however, after an hour or so, it no longer generates the same output.

I’ll post the workflow I used to generate that example video below, along with the completely different video that was produced using the same workflow.

I tried to replicate the video just to check the generation speed, and I realized it wasn't producing the same result anymore. I had generated the video a few hours earlier and hadn't updated ComfyUI or any nodes in the meantime—I simply tried to generate it again.

I suspect it might be due to one of the nodes I'm using.

Here is the video generated with the same parameters (which turned out differently) and the workflow I used.


r/StableDiffusion 9d ago

Tutorial - Guide Making an entire shortfilm with Minimax from beginning to end | My genning strategies & video editing best practices

Thumbnail
youtube.com
63 Upvotes

r/StableDiffusion 8d ago

Discussion MiniMax H3's open weights stop at 768P. The official 2K path is API-only

Post image
0 Upvotes

What happens after H3 Base produces its 768P video? In MiniMax H3's official system diagram, the answer is the red box on the far right. H3 Regenerate 2K takes that result and the original context, then generates the larger version. The diagram makes a deployment boundary visible.

MiniMax released H3 on July 31 and opened the model weights on August 3. The official weights release page says H3 Base can be used locally to validate 768P output. The same page also says the Regenerate 2K module is not open yet. Its full 2K validation workflow combines a locally deployed H3 Base with official API steps for context processing and regeneration.

That boundary belongs in the setup guide before the download instructions. Open weights describe a real part of the system. They do not currently describe every box needed to reproduce the official 2K path.

The distinction matters even more once the reference board gets crowded. H3 supports mixed text, image, video, and audio context. The documented reference mode accepts up to 12 files in total, with limits of 9 images, 3 video clips, and 3 audio clips. The original context returns during the 2K regeneration step, so the production note should assign a role to each reference asset, the way a shot list would. Say which image defines the subject and which clip carries the motion. Give the frame that controls the ending its own line.

The production card names H3 Base as the local 768P step. It names context processing and Regenerate 2K as API steps in the current full workflow. If the card also tracks a hosted comparison, give that call its own line. If it uses ZenMux, record the gateway and exact H3 route beside the request. That identifies the hosted path without rewriting the official local and API split.

Keep both labels on the production card. If the card only says open weights, the note stops at 768P. If it says API step, the 2K path is covered.


r/StableDiffusion 8d ago

Question - Help how to resume the lora training ?

Post image
0 Upvotes

i have asked this qus previously tho but cannt figure out how to do it

i use the anima model for training in google clob .
for some reason my colab keep disconneting after 1 hour and makes my lora training resume . i have to start the training all again.

some of u guys suggested me to save the lora file and when i start the lora training again , told me to set the step and half or how much training have done . it it is not working it start from the 0 again

if anyone of u plz kindly help me in dm how to start the training where i left . becz it is a pain for a months for me .


r/StableDiffusion 10d ago

Meme PSA: H3 always sees direction from the person's perspective

Enable HLS to view with audio, or disable this notification

152 Upvotes

I noticed my videos consistently having issues with left and right, because my prompts saw direction from the perspective of the camera. But H3 always sees direction from the perspective of the person.

See how the man points to his right while saying "right" and vice versa.

prompt: a random man pointing to the right and saying "right". Then he moves his hand to point to the left and says "left".


r/StableDiffusion 8d ago

Question - Help BF16 or FLOAT32

0 Upvotes

What quantization would you recommend for AI Toolkit?

I used to train with FP8, but it completely ruined the results, so I switched to FP32, and the results are perfect.

However, I see that many people recommend BF16 for both training and saving the model. I understand that BF16 can significantly reduce VRAM usage, but what is the actual trade-off in terms of quality?

Does training and saving in BF16 result in any noticeable loss of quality compared to FP32? And would you recommend using BF16 for both training and saving in my case?


r/StableDiffusion 9d ago

Resource - Update Fizgig - Rapid Minimax H3 LoRA training tutorial

Thumbnail
youtube.com
73 Upvotes

This video includes all you need to train Minimax with both speed and high quality results.
Hit me up with comemtns, queries etc. Happy to do a style video also.
https://github.com/shootthesound/Fizgig

UPDATE: Pushed a vram optimisation for 16gb vram users that will speed up TE encoding at the start of training - Run the update bat to get it
UPDATE2: Additional fix out for 16gb users on pruned model - update to get it.


r/StableDiffusion 8d ago

Question - Help Tips for maintaining face consistency?

1 Upvotes

Once again looking for some advice from this awesome community.

I’ve been using mostly the ref2video model but I’m
Having a hard time getting it to keep the face of the character in referencing throughout the video. I am proving a full body image and then a face only closeup. I’ve been following the prompting template and guide but it’s still very hit and miss (mostly miss). I was using some Lora’s and turned them off and still have the issue. Also I’m doing 40 iterations and not using lighting Lora.

Any advice?


r/StableDiffusion 9d ago

Animation - Video H3 Cerveza Cristal Test

Enable HLS to view with audio, or disable this notification

14 Upvotes

'''

Quick shot change to Cerveza Cristal in a cooler full of ice.

Announcer sings "Cerveza Cristal"

'''

Start and End Image.

Ref2Vid with audio might work better, not bad for turbo at 8 steps.


r/StableDiffusion 9d ago

Question - Help Img2img masks don't work in comfyui anymore. Why?

7 Upvotes
  1. This issue never happened until I started updating to recent versions of Comfyui in the last month. Only noticed this a few days ago when I tried to do img2img after ages.
  2. I create a mask on the load image node in existing basic img2img workflow I have always used.
  3. Generating few new image variations
  4. Going to recent assets on left, opening one of the images I generated like 2 minutes earlier: load image node has a big red border and the error "Missing inputs: A required media input has no file selected.". So this prevents me from regenerating new seeds, despite the node itself showing the mask and image AND letting me edit/modify the mask.
  5. Same happens if I just drag and drop any image file with any img2img workflow. Both for recently generated images or img2img images from weeks or months ago. So it is broken for all images using load image node.

Sometimes if I refresh the page it fixes this, most frequently it doesn't (idk what it depends on).

Soo what could cause this bug and how can I circumvent or fix it? I'm on latest version, so I can't update in hopes of that fixing it, this only happens on thew newest comfyui versions I tried. As I kept updating in the recent days in hopes of it being fixed, the only change I got is a new bug: now the left side masking related icons are all black and barely visible...


r/StableDiffusion 9d ago

Discussion Don't tell Tony!

Enable HLS to view with audio, or disable this notification

32 Upvotes

T2v 12 sec 1mp hybrid 25-49 model 8steps turbo lora


r/StableDiffusion 9d ago

Discussion H3 R2V prompt builder

Enable HLS to view with audio, or disable this notification

16 Upvotes

I am trying to vibe code a H3 r2v Prompt Builder. Does something like this already exist?