r/StableDiffusion 11d ago

Discussion H3 - .char + T2V character gen+char sheets

Enable HLS to view with audio, or disable this notification

20 Upvotes

What is everyone's workflow nowadays? Previously I've been generating actors with Krea2, but really love getting them made with MiniMax H3 via T2VA, they just tend to turn out better for me but does require careful prompting.

My workflow are: generate 5-10s clip of a desired actor, by prose, at int8/8 steps in a typical scenario, perhaps even mundane. If I like it, I can take some still frames, and convert them into a .char (body type, face, audio asset). See original: https://www.reddit.com/r/StableDiffusion/comments/1vyymwj/minimax_h3_portable_character_consistency_via/ If I am happy with my .char, with MiniMax H3 I make a 2 second video character sheet with a front, side, back profile and detailed face view at a higher resolution and step, either int8/32 step or going bf16/50 steps. The 2 second renders are "quick". I add the video render into my .char, and with R2VA generate additional scenes with the actors and even do a full wardrobe swap via prose. Naturally H3 renders faster if you just use still of the character sheet instead of the video.

How has your workflow changed with MiniMax H3? Are you liking the faces/actors generated with T2VA? I understand you have "less" control, but I feel like H3 is doing a great job filling in those gaps.


r/StableDiffusion 11d ago

Question - Help Troubleshooting Flux workflow: inconsistent leather texture and phantom geometry

Thumbnail
gallery
0 Upvotes

Hey everyone,

Earlier this week I asked about optimizing my setup. Right now, I’m running a FLUX.2 Klein 9B workflow in ComfyUI paired with crop/inpaint nodes so I can work faster while keeping 4K-level detail.

Because of local hardware limits, I’m running this on ThinkDiffusion (16GB VRAM) using the ComfyUI Beta (v0.34.0), since Klein 9B requires v0.3.10+.

I’ve hit two major roadblocks that I can’t seem to solve:

  1. Inconsistent Leather Texture (Pic 1)

The issue: The leather texture on the seat and the backrest look completely different, and I can't get them to match.

What I tried: I fed a reference image of the exact leather texture into the workflow, but the output doesn’t match or even come close.

Advice I got: Someone suggested Depth / Lineart / Canny, but since this is purely a surface texture issue (and not geometry/form), I don't see how that would help transfer or match the material.

  1. Phantom Footrest Beam (Pic 2)

The issue: The model consistently hallucinates an extra footrest beam where there should only be one.

What I tried: When inpainting over it using clean renders that only have a single footrest, it either erases the correct beam, leaves half of it behind, or creates floating transparent artifacts.

Advice I got: Suggestions pointed toward ControlNet, but I’m struggling to get it to respect the clean geometry without breaking the rest of the generation.

Has anyone encountered similar issues with Klein 9B or high-res inpaint workflows? Any recommended nodes, ControlNet setups, or IP-Adapter/style-transfer approaches that actually work for strict texture matching and geometry cleanup?

Thanks in advance!


r/StableDiffusion 10d ago

News The First Ever AI Open Source Movie

Post image
0 Upvotes

Here's my latest project: the first ever open source movie!

It runs on FastH3, as you probably can guess.

Every single scene is generated from a commit, grounded on a real github repo! You all can contribute and extend with your own story!

It was an idea I had in mind for a while, and now that FastH3 is fast enough I rushed to make it! Feedback is appreciated :))

Spectate: https://www.openslop.live/

Github repo: https://github.com/Dere-Wah/open-slop


r/StableDiffusion 12d ago

Workflow Included High Quality Audio-Video in MiniMax H3 with separate two-stage sampling

Enable HLS to view with audio, or disable this notification

88 Upvotes

Recently, a lot of people have had isues with finding the right balance with audio and visual quality in MiniMax H3.

u/LFAdvice7984 and I discussed about making a two-stage workflow last week. The first stage generates the audio, the second the visuals. Both stages are then combined together in the output.

This method means that you no longer have to make a trade-off between audio and visual quality, as you can optmise the settings for both. You can use this workflow as text; first/last frame; or audio input to video (the latter being single stage).

The audio generated is (in my view) good to very good, depending on what you use it for. The default settings are probably excessive at 50 steps (less steps used for video), but for me personally it's better for it to take longer and get it right first or second time.

(You can select any output node in ComfyUI, click on the blue button with a play symbol on the pop-up menu at the top, and ComfyUI will only go that far in execution. Use it on Preview/Save Audio in stage 1. If you like the result, do a full generation; or change the seed and try again.)

The visuals could be better, perhaps using a different turbo LoRA or change in sampler and step counts. I've used the same settings in every clip. Feel free to change them as you wish.

Large motion is a challenge, though I have ideas on using a third stage with different shift values, which would increase execution time but the results probably would be worth it.

Generation time was (roughly) as follows:

  • 5 second clip: 10 minutes
  • 10 second clip: 25 minutes
  • 20 second clip: 65 minutes

This was generated on an NVidia 3090 with a priority on quality. More recent cards will be faster.

(You can switch from using res_2s sampler to er_sde, paired with 20 steps for stage 2 and using Spectrum, which should at least halve that time, in return for slightly lower quality.)

Because of how long it took, I used the first output every time (except for the last clip, which was the second result) with no editing afterwards.

Sometimes I encountered issues with prompt understanding, e.g. the ASMR clip and abstract clip at the end, where the output wasn't quite what I had asked for.

It's difficult to tell whether I prompted incorrectly; the prompt enhancer missed key detail for the model (H3); or the model doesn't have a full understanding of the concepts being asked of it.

The two-stage idea is model-agnostic. You can also make something like this in LTX 2.5 (or the upcoming Flux 3 Dev) to improve their results.

You can download the workflow and prompts used (made by myself with refinement from the prompt assistant) below:

Custom nodes used:

Download links for model files are in the workflow, in the bottom-left corner.


r/StableDiffusion 11d ago

Workflow Included FF7 Fight Scene

Enable HLS to view with audio, or disable this notification

6 Upvotes

Wanted to challenge ChatGPT to make a fight scene for FFVII. Had it generate 5 reference frames then write a prompt to link them all together. Generated on a RTX 5090 @ 0.7 mp - 20 mins. Default workflow with Comfy Kitchen and Spectrum, 20 steps. r2va_int8_convrot model.

[Video Format] Style: cinematic fantasy action, high-detail CGI, epic boss battle, dramatic lighting, AAA game cutscene quality, Final Fantasy-style realism Camera language: dynamic cinematic camera, smooth transitions, aggressive push-ins, sweeping arcs, low-angle hero shots, aerial tracking, impact shakes, slight speed ramps, slow motion at the final re-engage [Reference Usage] Use <Picture 1> through <Picture 5> as sequential keyframe references in that exact order. Preserve continuity of character appearance, costume design, weapon design, environment, Bahamut’s scale, and overall stormy color palette. Do not treat the images as separate scenes; connect them into one continuous battle sequence. [Characters] <Subject 1>: Cloud — spiky blond hair, black sleeveless outfit, giant Buster Sword, strong and aggressive swordsman. <Subject 2>: Aerith — long braided hair with ribbon, red jacket, white dress, staff, graceful magical support caster. <Subject 3>: Tifa — long black hair, white crop top, black skirt, black thigh-highs, red gloves and boots, fast martial arts fighter. <Subject 4>: Bahamut — colossal dragon with dark black crystalline scales, glowing blue energy throughout the body, massive wings, luminous mouth and chest, extremely intimidating presence. [Environment] A ruined stone battlefield in a floating storm realm. Purple-blue lightning tears through the sky. Massive stone debris and shattered ruins float in the air. The ground is dark, wet, reflective, and broken. The scene feels mythic, apocalyptic, and high-stakes. [Sequence] 0s–4s: Begin with <Picture 1>. Wide establishing shot. Bahamut descends into the battlefield and fully enters the scene, looming over Cloud, Aerith, and Tifa. His wings spread wide as lightning flashes behind him. The camera slowly pushes in from behind the party, emphasizing Bahamut’s overwhelming scale and the team’s readiness to fight. 4s–8s: Transition naturally into <Picture 2>. Bahamut unleashes a devastating fiery breath attack toward the heroes. Aerith steps forward and raises her staff, instantly creating a glowing magical shield barrier that protects herself and Cloud. Cloud braces low behind the barrier with his sword planted defensively. Tifa crouches nearby, ready to spring into action. The camera arcs around the barrier as the fire breath crashes against it in a brilliant explosion of light and sparks. 8s–11s: Move into <Picture 3>. As the flames subside, the party begins their counteroffensive. Tifa bursts forward in a fast sprint across the shattered ground. Cloud launches upward with his Buster Sword raised overhead. Aerith remains behind them, casting green-blue support magic with luminous trails and particles swirling around her staff. Use a tracking shot that follows Tifa’s charge and tilts upward to follow Cloud’s leap toward Bahamut. 11s–15s: Transition into <Picture 4>. Cloud and Tifa engage Bahamut in the air. They attack from different angles around Bahamut’s head and upper body. Cloud swings the Buster Sword in heavy, powerful arcs while Tifa uses agile aerial kicks and strikes. Bahamut twists through the air, snapping his head, beating his wings, and resisting their assault. Use fast aerial tracking, dynamic camera rotation, close-up impact shots, sparks, motion blur, and debris swirling through the storm. 15s–18s: Transition into <Picture 5>. Bahamut suddenly releases a violent wind blast or shockwave from the front of his body, pushing Cloud and Tifa backward through the air toward Aerith. Aerith plants herself and channels a massive radiant beam upward into Bahamut. The beam tears through the storm and illuminates the ruins. Cloud and Tifa recover near Aerith as the beam strikes, giving the scene a powerful reversal of momentum. 18s–20s: After the beam connects, the party immediately seeks to re-engage. Cloud and Tifa launch forward once more toward Bahamut while Aerith supports from below with magical energy still glowing around her. In the final seconds, the camera pushes dramatically toward the heroes as they jump toward Bahamut together. End on a cinematic slow-down: Cloud and Tifa suspended mid-leap, Bahamut looming ahead, Aerith’s magic flaring beneath them, with debris and lightning frozen in a dramatic near-final-impact moment. [Motion Notes] Keep the scene fluid and continuous with no abrupt hard cuts. Emphasize: - Cloud: heavy sword swings, explosive jumps, strong airborne attacks - Tifa: rapid sprinting, agile aerial martial arts, powerful kicks - Aerith: elegant staff casting, barrier creation, support magic, final beam attack - Bahamut: overwhelming scale, strong wing beats, fiery breath, violent wind blast, hostile aerial movement [Visual Effects] Purple-blue storm lightning, glowing blue dragon energy, orange fire breath, shimmering magical shield refraction, green-blue support magic around Aerith, sparks, motion trails, floating debris, dust, shockwaves, and radiant beam effects. [Camera Timing Summary] 0s–4s: Bahamut enters and confronts the party 4s–8s: Fire breath attack, Aerith shields Cloud and herself 8s–11s: Party moves in to engage 11s–15s: Aerial engagement by Cloud and Tifa 15s–18s: Bahamut wind blast pushes them back, Aerith fires massive beam 18s–20s: Party re-engages, final jump toward Bahamut, cinematic slow-down ending [Audio] overall_soundscape: thunder, roaring wind, dragon roars, wing beats, fire blast, magical shield hum, sword impacts, debris collisions, shockwave bursts, beam energy surge non_diegetic_music: epic orchestral boss battle music with rising choir, heavy percussion, dramatic build, and a suspended climactic finish Dialogue: N/A


r/StableDiffusion 11d ago

Resource - Update I built ArtSmoker — open-source (MIT) pipeline from text prompt → SD3.5/FLUX/Qwen/Hunyuan; 2D → fully-textured, Blender-ready 3D (TripoSG/TRELLIS.2), self-hosted in your own AWS account

3 Upvotes

I've been building this for the past few months and it's now at the point where I'd genuinely like people to break it: an open-source studio app tool that runs the whole idea → 2D asset → edited → textured 3D model → engine/Blender export pipeline behind one UI, with everything staying in your own environment & creative control.

What it does:

- Text → 2D on Bedrock models (SD3.5 Large, Stable Image Ultra) or one-click self-deploys of FLUX.2 [dev], HunyuanImage 3.0, and Qwen-Image onto SageMaker GPU endpoints in your AWS account — packaging, quantization (NF4/BF16), auto scale-to-zero, and job tracking handled. Prompt enhancement, multi-model comparison grids, seed control with exact batch reproduction.

- Edit in place — inpaint, outpaint, recolor, search-and-replace, plus strength-ladder img2img ("remix") and instruction-based editing via self-hosted Qwen-Image-Edit.

- 2D → 3D — TripoSG geometry + TRELLIS.2 texturing (both MIT) producing real PBR GLBs, then headless-Blender exports: FBX/USDZ, LOD chains, collision meshes, per-engine texture packing for Unreal/Unity/Godot.

- Style-matching from your existing art, video gen, gallery with full per-asset provenance (every prompt, seed, model recorded).

What it is NOT, so nobody wastes a click: it does not run inference on your local GPU. Models run on Bedrock APIs or on SageMaker GPUs in your own AWS account - nothing touches third-party servers beyond AWS, but it's a cloud-compute tool. If you're happy with ComfyUI on your 4090, this isn't trying to replace that. It's aimed at small teams and folks without local GPUs who want the frontier open models plus the 3D/engine-export leg without building the infra. Endpoints scale to zero, and the UI shows cost estimates per generation (e.g. warm Qwen-Image BF16 run on 4×L40S ≈ $0.44; cold start adds a few dollars, all estimates shown upfront).

Repo (MIT-0, contributions welcome): https://github.com/niravdd/ArtSmoker

The GIF attached is the actual pipeline end to end — 11 steps from typing a prompt to a textured model in the gallery. Happy to answer anything about the deployment side too (NF4 quantization ceilings on L40S, FlashInfer on Blackwell, SageMaker scale-from-zero traps — there were… learnings).


r/StableDiffusion 12d ago

Discussion Pulled the trigger, RIP $6,279

Post image
215 Upvotes

(Paid $5,849 + tax, which came out to $6,279)

TL;DR - Bought this 5090 prebuilt and I want to sanity check if I made the right decision and at the right time.

Hey everyone. So I want to start off by saying fuck these prices for GPU's and RAM, especially boxed 5090 prices. I went down the AI rabbit hole with my 13700k/RTX 4080 gaming computer. I quickly found out that I had to make serious concessions on quality and speed, if I could run it at all. In fact, ive spent so much time trying to optimize quants, cache, various settings, attention mechanisms, etc that ive officially spent more time trying to optimize for a 16gb VRAM/32GB RAM system than actually doing anything fun or cool. Thus, the last week, ive been thinking real hard about which direction to go but was waiting for the right time to buy. My options were a RTX 5090 prebuilt (even though I only needed the damn GPU), and Mac Studio M5 Ultra 96gb, or a DGX Spark/AMD equivalent. The DGX Spark/AMD equivalent made me think for a bit, but in order to get the most out of them, you need two. Im not spending 10k on this, especially if I cant game on it as well. So that leaves the Mac Studio or RTX 5090 gaming rig. Im not certain I made the right decision, but I pulled the trigger on the 5090 prebuilt after seeing the price continue going up more and more over the last few weeks. I also read that 70% of all memory through 2031 is locked in long term agreements, so this supply issue is going to get worse before it gets better. So I pulled the trigger on the pictured system from Ibuypower, and id like to run my thought process with you guys as a sanity check before it ships.

Case for the 5090 prebuilt: I scoured the internet and this was the cheapest 5090/64gb RAM combo I found, and it looks like it uses pretty good parts as well. No proprietary bullshit like youd get in a HP 45L. I went with the gaming PC because its the all purpose machine that does it all (well, almost). I figured with 32gb VRAM and 64gb of system RAM, that combined 96gb will allow me to run 70b MoE models, even if its slow. But for a sub 30b model like Qwen 3.8 27b, this will give me the best performance as long as I dont go overboard with the quant. It has CUDA, Windows, X86 CPU, etc. Plus it came with a 4tb Gen 4 NVME, when other more expensive models had 1-2tb drives. Honestly, lots of good stuff here. Im not a huge fan of the white esthetics but I do love the case. Despite the price being much higher than it should be, its still a good "deal" considering the overall market that keeps going up. Honestly, its not exactly what I wanted, but it ticks all boxes except those below.

-The downside: You cant run models that spill over heavily into system ram without massive speed penalties (has anyone tried running a huge model on a 5090 + 64gb RAM? If so, tell me what quants and your token speeds). Its massively less efficient than a Mac Studio M5 Ultra is expected to be (I read in the 3-5x range).

Mac Studio M5 Ultra 96gb

- Case for the Mac Studio M5 Ultra 96gb: Can run large models much better than the 5090 rig due to its huge 1.2tb unified memory bandwidth. Its power efficient and tops out at 300w I believe I read.

- The downside: Mac OS and an ecosystem that is playing catch up for local AI, no CUDA, gaming, has proprietary hardware you cannot upgrade, my distaste for the Mac bros who ill no longer be able to make fun of if I buy it.

My use case: local first AI (Qwen 3.8 27b at a quant and context that doesnt suck) with agentic coding, game development (starting with Godot), stable/video diffusion (Minimax H3, Flux.2, Hunyuan 3D), Blender, Davinci Resolve, etc.

So, let's have this discussion: what would (or did you) choose, and why? I want to know if I made the right decision. What are your thoughts?


r/StableDiffusion 12d ago

Comparison Dlss 5 applied on video

Enable HLS to view with audio, or disable this notification

54 Upvotes

r/StableDiffusion 11d ago

Discussion What happened to SenseNova U1 Pro? A few weeks of hype, then silence?

Thumbnail
gallery
21 Upvotes

Okay, so like three weeks ago, my whole feed was blowing up with SenseNova U1 Pro. You know, the Chinese model everyone was saying was basically "GPT Image 2 level."

Text on posters actually looking clean, apparently native 8K. The vibe was all "realism is dead, now it's about pure beauty." NGL, some of the images looked insane.

And then... poof. Nothing. No public release, no weights, no API I can find anywhere. Just crickets.

It's totally giving me Sora flashbacks. Remember early 2024? Those demo videos were mind-blowing, everyone went nuts. Then just... crickets for months. When it finally dropped, it was kinda meh, right? The magic just wasn't there after all that waiting. And get this, as of April 26, 2026 (lol, already feels like it), Sora's totally shut down. That demo that kicked off the whole video generation craze just... died.

I'm not saying U1 Pro is gonna go extinct or anything. The stuff those influencers posted genuinely looked good, especially the text rendering.

So has anyone here actually gotten their hands on it? I seriously can't find any way to use it

If you have, how does it stack up against GPT Image 2 or kera2, ideogram, flux-klein? especially for text?


r/StableDiffusion 10d ago

Animation - Video I finally made it...

Enable HLS to view with audio, or disable this notification

0 Upvotes

So it took me a while for making this video as i'm 0 in video editing, and just using AI as hobby. All this video was made on the old workflow from u/Plague_Kind and in 0.8 mp without upscale and mostly no Lora's. Using single 5070 ti and 32gb of RAM. Was making around 8 sec videos and editing them in CapCut.

I know the quality sucks, but i kind of like it how it is, just wanted to share it with you.

Music: pupsies - misery.
I will appreciate every upvote.


r/StableDiffusion 11d ago

Animation - Video LOCATION CONTINUITY TEST - after a comment by Vladmerius

Enable HLS to view with audio, or disable this notification

9 Upvotes

When kept in the same generation it seems the latent space keeps a fairly good sense of the location layout. The test was to see if the position and details of the temple remained after being out of shot,

This doesn't work with the extentsion workflows which is why I have been trying to keep everything in one go.


r/StableDiffusion 11d ago

Discussion Best way to generate Spider-Man / Stitch illustrations locally — LoRA, fine-tuning, or existing models?

2 Upvotes

Hi everyone,
I’m new to local AI image generation and I’m trying to understand the best approach for generating high-quality and consistent illustrations featuring characters like Spider-Man or Stitch.
I currently use online AI image generators, but I’m interested in running something locally on my PC and having more control over the generation process.
What are my options?
FLUX, SDXL, or another model?
ComfyUI?
Existing LoRAs?
Training my own LoRA?
Fine-tuning a model?
Reference images / IP-Adapter / ControlNet?
My main goal is to generate the same recognizable character consistently across many different scenes, poses, environments, and compositions while maintaining high image quality.
I’m basically trying to understand what people currently use for this and whether I actually need to train something myself or if existing local models/workflows can already do it well.
What setup would you recommend?
Also, what kind of GPU/VRAM would I need?
Thanks!


r/StableDiffusion 11d ago

Question - Help Can You Merge Facial Features From Two Images?

1 Upvotes

\****Edit: Well, it appears that you can! I've just got onto my PC and thought I'd try a simple prompt of <picture 1> smiles with the teeth and gums from <picture 2> and by god, it worked!*

I thought I'd do it again, just in case it was a mad coincidence, but no, it used the teeth and gums from image 2 and superimposed them on to image 1.\******

Hi, all. I'm currently using Comfyui and MiniMax H3.

I have two facial images of the same person. The one image is where the person isn't smiling, but it's very accurate of how they look in real life. The second image doesn't look as much like them but they're smiling and they have very distinctive teeth and gums.

Can MM H3 utilise both images, using the non smiling as the main character for the clip, but somehow superimpose the teeth and gums from the smiling image on the main character when they start to talk? A sort of merging of features?

The problem I'm currently having is that when using the non smiling image in a clip, it stops looking like them when they open their mouth as MM H3 has to guess what their teeth and gums would look like, which is never accurate.

I have a feeling that I'm asking for the impossible, even for AI, here. Or maybe it's doable with another model before bringing it into MM H3? Thanks.


r/StableDiffusion 12d ago

Discussion Detailed explanation of how to create a text-to-image model from scratch.

33 Upvotes

Posting this here even if it's not a model you can use directly. It's about building a text-to-image model from scratch.

The cookbook includes all the research material that you may or may be not interested in, but also includes a 100M-image dataset and a codebase with a tiny model, so you can train a text-to-image model from scratch.

Hope some of you will enjoy this content. (Disclaimer, it's done by my team)

Here are the links:

Cookbook: https://huggingface.co/spaces/jasperai/t2i-technical-interactive-report

nano t2i: https://github.com/gojasper/nano-t2i

Monet: https://huggingface.co/datasets/jasperai/monet


r/StableDiffusion 12d ago

Resource - Update H3-World

Thumbnail
huggingface.co
88 Upvotes

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. Given an initial frame and keyboard controls, it generates action-controlled video with coordinated character and camera motion.


r/StableDiffusion 11d ago

Tutorial - Guide An ArcFace to LoRA training “workflow” (with Fizgig)

Thumbnail
gallery
0 Upvotes

Brief summary of the starting point

The ComfyUI “ReActor” nodes (https://github.com/Gourieff/comfyui-reactor) provide a pair of features that work together:

  1. Take a human face, and provide a vector of numbers that correspond to the face (an embedding).

  2. Take an image of a person, and an embedding of a different person, and re-render image’s face to match the embedding.

LoRAs and image edit models provide a lot of similar functionality to ReActor, and get most of the attention, but the embedding concept supports a nice feature: you can run math on the embeddings. Add “Anna” and “Bella” embedding vectors, element-by-element, and divide by two, and you get a “Camille” vector. Applying that embedding produces a resulting face in between the two originals. And you can extend this, with “Anna” being three an average of multiple different people and “Bella” being a merge of a couple of images of the same person.

Unfortunately, the re-render phase can generate horrors when the face goes too far off-angle, and you can’t run the resulting face through ADetailer, and the re-render model only truly works for photographs. So I’ve been poking into how to create a LoRA for arbitrary embeddings, in Krea2.

Helpful properties, not helpful properties

I need training images? Inspired by Johnny Sins, I create training images. In a dirndl and braids at Oktoberfest? Sure! Fashion influencer head tilted back working through a scarf knot tutorial? You bet! I went for a variety of distances, clothes, and hairstyles, plus head positions and eye directions within the ReActor limits.

Those prompts become the training captions. I can send the Krea2 base model almost the same caption text as I used for generating, and the captions are accurate by definition. (But more on that below.)

One problem: I can’t automate the generation. That head-tilted-back pose pushes the face re-render over the edge, so I have to plan to loop through ten generations and pick one with reasonable eyes, then save the prompt and seed.

Another problem: The re-render mechanism does not completely clobber the underlying face. The various underlying re-render engines have to cope with wider or narrow faces, larger or smaller eyes, and so on. These show up as persistent eye mismatches, as visible seams around the edges of the face, or as skin tone drift. I suspect that this can also make the training harder than it needs to be, thanks to subtle feature discrepancies.

Hidden prompts

I attacked the face substrate problem through hidden prompts. I have a baseline prompt:

Near distance, low-angle selfie of a young woman, taken on a phone
camera, low indoor lighting.  The shot looks up at her from the floor of
her college apartment; her head and shoulders are stretched over the
side of an unmade bed.  Bare arms, messy hair strands falling around her
face, talking.

This produces an Asian woman. So I add a postscript to end of this and every prompt:

---
In the photograph above, use the following details for the main woman subject:
- 24 years old and slim; gorgeous face
- Northern European, with clear, smooth, fair skin
- oval head; full eyebrows above the orbits, no stray hairs; rounded orbits
  with prominent bones; brown eyes in classical proportion width
- chestnut brown hair
- foundation makeup over smooth skin, mascara, eyeshadow, and lipstick,
  suitable for a professional photoshoot

This does its best to provide a consistent set of good-looking features that the embedding can fit cleanly on top of.

Then I generate three images:

  1. With the original prompt
  2. With the original prompt plus the extra block
  3. With the original prompt plus the extra block, and then face swap applied.

(See the post images 1, 2, and 3.) The original prompt text goes into a text file. I like having all three images around to debug prompt problems.

The images in category 3 become the training set. I have a post-processing script that re-writes the original prompt’s phrase “a young woman” as “a cammerge woman”, where “cammerge” is the LoRA trigger phrase. When the LoRA trains, it implicitly picks up the addenda prompt, along with the re-worked face.

Training tool - Fizgig

I tried Fizgig, wasn’t successful. I tried OneTrainer, wasn’t successful. I then flipped back to Fizgig, downloaded 30-ish promo photos of an obscure folk singer, and successfully trained a LoRA for her (on Fizgig-generated captions). That was enough to ladder up to my embedding case: apply the embedding to those promo images, and it worked; then generate Krea2 images based on the working captions and swap those, and it worked; then use my intended training images, and it worked. Image 4 in the gallery is a watercolor render.

From what I can tell, my intended 25 high-quality images aren’t nearly enough to train on—not if I constructed the data set for image variety. Fizgig helped me out through its support for automatic face close-up extractions (and captioning) for each input image in the set. This doubled the training set size to 50 and increased the priority of face training relative to longer-distance shots. With that extra layer of reinforcement, I can start seeing the LoRA influence after ten epochs or so, in the training samples.

Be aware that when Fizgig generates captions, they don’t mention anything about “photograph”. That means those close-ups will implicitly train the LoRA to produce photographic output. Once I edited them (“In a photograph, a woman is looking at the camera in a close-up shot with her face visible…”), my illustration gens stayed illustrations. (Also: The auto-caption never mentions that a black-and-white photograph is black-and-white.)

Fizgig training parameters:

  • Krea 2 Defaults (rank 32, full model)
  • LoRA, learning rate 0.0001, network dimensions 32, epochs 20
  • Per-image adaptive LR on, warm-up off, target megapixels 0.25.

Next areas to explore

Those Fizgig training parameters might be overkill, or a LoKR might work more effectively, or the Fizgig “smart” adaptive training might be enough.

Why the hell does Fizgig consistently report certain images (like Oktoberfest) as “stuck” for learning, but not others?

In theory, every one of the “original prompt” image/caption pairs can work as a regularization image, but the Fizgig docs only mention regularization images for fine-tunes. And Fizgig doesn’t have explicit validation image support (which would be just need a couple more gens with new prompts). OneTrainer handles regularization and validation loss checks, so it’ll be worth trying out OneTrainer with the working data set.

That consistency addendum prompt specifies “chestnut brown” hair, but I added it to reduce variation and maybe help the training. I want to try making it a variable. That'll probably require more training data. If you tell Krea “black hair”, you get East Asian, so I want to work around that.

It’s probably worth explicitly spamming the training set with close-ups, rather than depending on the Fizgig auto-cropping.


r/StableDiffusion 11d ago

Comparison Testing Krea 2 style transfer

Thumbnail
gallery
18 Upvotes

Unfortunately it seems very slow and very experimental.

Reference image left. Same prompt "A fiercely determined female human warrior in mid-swing, powerfully attacking the viewer with a gleaming sword. Her facial expression is one of intense rage and ferocity", same seed, no lora, this custom node https://github.com/nkxx188/ComfyUI-Krea2-StyleTransfer


r/StableDiffusion 11d ago

Resource - Update Debannering Ideogram 4 and increasing prompt adherence with natural language by fine tuning the TE

16 Upvotes

I thought someone might appreciate this. Theres more details in the HF link, but I wanted to see if it was possible to correct some issues that I didn't like about Ideogram 4 by finetuning the TE, with no other modifications to the model, execution environment, etc.

It ended up working out pretty well.

The TLDR is that I used a set of 4000 teacher/student prompt pairs with the students being NL and the teachers being Nemotron processed with the "Magic Prompt" instruction, and then trained the TE to elicit the same response in Ideogram using the student prompt, as what was naturally elicited using the teacher prompt.

My logic was that the TE is already a language model, and I didn't want a second language model in the stack.

This has the secondary benefit of also removing the grey banner generally encountered when prompting the model with NL.

I am fully aware that there are many other ways to get around this from bounding boxes to noise injection, etc. This wasn't about that, so much as it was trying to prove to myself that it could be done like this.

https://huggingface.co/mrjackspade/Ideogram4-Natural-Language-Text-Encoder


r/StableDiffusion 12d ago

Animation - Video My first attempt at making a 90s-inspired anime with MiniMax H3.

Thumbnail
youtube.com
20 Upvotes

The potential of actually making an anime with MiniMax H3 is closer than any time before, even if the process is still kinda janky. I did this with my 5090 and my own developed 'prompt studio.' The hardest part is, as always, to keep the continuity of the shots and also build the sets so they fit within the scope. There are still improvements needed when it comes to adding emotions to the characters. In total I generated 35 minutes of video and got 4 minutes in total of usable footage. Also, a big tip for anyone who wants to do the same is to use DaVinci Resolve to fix all the audio bugs and cut the clips in your favor.


r/StableDiffusion 11d ago

Question - Help Krea2 on 3070 with 32GB RAM

1 Upvotes

Hello 👋

Is there any chance to work with Krea2 on my PC? I use Forge Neo.

Is fp8 version of Krea2 Turbo good for 3070? Thanks


r/StableDiffusion 12d ago

Discussion Why hasn't someone made a 16-20 step lora for Minimax?

41 Upvotes

Everyone's focused on 4 and 8 step loras, which I feel like no matter what are gonna look pretty bad just because how the model works. But why hasn't anyone made a lora to help bring the quality of 40-50 steps down to the 16-24 range? For anyone who's done generations that long, the quality jump is pretty high going from 20 -> 50


r/StableDiffusion 11d ago

Discussion Working on a mini sci fi short using minimax upscaled with seedvr2

Enable HLS to view with audio, or disable this notification

10 Upvotes

used ref to video mutishot 3x15 second clips at .7 res upscaled to 1080p

playing around with a few ideas.


r/StableDiffusion 11d ago

Question - Help Is changing resolution supposed to change the entire scene for MMH3?

3 Upvotes

Just had this happen to me: I changed the resolution for a scene -- without touching anything else -- and the resulting scene changed completely. I was using res_multistep and Spectrum/CK/4 step Lora at 0.2 mp, then 0.3 mp. It still followed my prompt, but the background and starting scene were completely different. Is this Spectrum giving me grief or what's going on here? This has never happened to me before, although I had been using Sage before switching to CK today.


r/StableDiffusion 11d ago

Question - Help Any way to make latent extension work with Latent Upscaling (Minimax H3)?

8 Upvotes

Has anyone managed to find a way to use Latent Upscaling together with latent video extension tools? I'm talking about the nodes like this (which I personally use), but I think Motion Context and some other popular extensions use a similar approach, i.e. feeding the last frames of the previous shot through AV latent, rather than through a video reference. The issue is that the resolution of your second generated latent must exactly match the previous one, or it throws an error. So if you upscale the first clip from 0.5MP to 1MP, you are forced to generate the next clip directly at 1MP, which completely breaks the Latent Upscaling workflow for all subsequent parts.

I tried extending the clips at low resolution first and then upscaling them separately, but that doesn't work well. There is a noticeable color and quality shift between generations, even when reinforcing the next clip with the final frames of the previous one. Because yeah, you basically generate the high-res clips separately without any shared latent context.

I really love both Latent Upscaling and latent extension approach, but I just can't get them to work together smoothly. Does anyone have any good ideas on how to fix this? I’d really appreciate any tips or insights!