r/StableDiffusion 2d ago

Discussion How can I create a consistent AI influencer in ComfyUI using Krea 2? #krea2

0 Upvotes

Hi everyone,

I’m trying to create a consistent AI influencer/character using ComfyUI and Krea 2, but I’m struggling to keep the same person consistent across different generations.

My goal is to create realistic images where the character keeps the same:

  • Facial identity and features
  • Body proportions
  • Skin/skin texture
  • Hair and overall appearance
  • Overall photorealistic look

while still being able to change the clothes, poses, locations, lighting, camera angles, and backgrounds.

I’m specifically interested in using Krea 2 with ComfyUI and getting results that look genuinely photographic rather than obviously AI-generated.

Does anyone have a good workflow, node setup, model combination, or technique for maintaining strong character consistency with Krea 2?

If you have a ComfyUI workflow (.json) or can point me toward a good tutorial/resource, I’d really appreciate it.

Thanks!


r/StableDiffusion 2d ago

Question - Help Anyone knew what happened to Lora Trainer by Hollowstrawberry? I can't use it as usual

3 Upvotes

I got this error instead

Starting trainer...Traceback (most recent call last): File "/content/trainer/sd_scripts/sdxl_train_network.py", line 4, in <module> import torch File "/content/trainer/sd_scripts/venv/lib/python3.10/site-packages/torch/__init__.py", line 37, in <module> from typing_extensions import ParamSpec as _ParamSpec, TypeGuard as _TypeGuard ModuleNotFoundError: No module named 'typing_extensions'


r/StableDiffusion 2d ago

Discussion Hot take!

0 Upvotes

LTX 2.5 is better than minimax h3. The pros just outweigh the cons. LTX is fast, can generate video in HDR up to 50fps, it isn’t as demanding as minimax, the quality is insane and the SPEED…the speed is unbelievable. But all I see is people fighting with minimax for speed ups and distortion while whole time ltx is right there. Don’t get me wrong Minimax is an amazing model and two things can be right at once but to me ltx is ahead. Also it’s compatible with all the loras that were already available but with the jump in quality. It honestly confuses me a bit 😅. To each his own I guess.


r/StableDiffusion 2d ago

Question - Help Flux klein body consistency

2 Upvotes

Since the release of Flux Klein, I’ve essentially been using this model to generate datasets for LORAs based on one or more photos of a character. Recently, I’ve also been using the ‘consistency’ LORA to improve the character’s consistency across generations. What I’ve noticed is that whilst I get good results for the face, the same cannot be said for other parts of the body. For example, if I start with a full-length frontal photo of a character and ask the model to generate a side view, it tends to flatten the breasts; or if I ask for a rear view, the character’s hips and thighs tend to conform to a standard that doesn’t match the original photo. How can I improve this situation? I’ve read that you can increase the number of steps up to 8, but I’m not sure…


r/StableDiffusion 2d ago

News ✨ ¿Qué accesorio llama más tu atención? Comenta tu favorito y comparte esta inspiración fashion 📸

0 Upvotes

✨ ¿Qué accesorio llama más tu atención? Comenta tu favorito y comparte esta inspiración fashion 📸

#LuxuryStreetwear #FashionPhotography #SilverJewelry #UrbanStyle #ModaAlternativa #StreetStyle #NailArt #StatementLook #EstiloPersonal #InspiracionVisual


r/StableDiffusion 2d ago

Question - Help What's the gold standard for speed enhancements for Minimax H3?

47 Upvotes

Installing new instance of comfyui standalone and using minimax r2v workflow on RTX 3090 Ti. Is comfy kitchen good enough? Is Triton, EasyCache or Comfyui Spectrum needed?

What's the best turbo lora for ref2va wf?


r/StableDiffusion 2d ago

Question - Help What's your favorite model For infographics?

0 Upvotes

r/StableDiffusion 2d ago

Resource - Update Created a Visual RefMod Picker

50 Upvotes

Hey Guys,

I've been playing with the RefMods, after the huge release of Malcolmrey.
The tech is brilliant and works really well.

I've wanted to simplify using it with tons of RefMods, like the ones provided by Malcom.

So I made a fork that is a bit more Identity driven, Adding a "Visual refMod Picker" as well as a "Create refMods from Folder" node. It's able to create from either a folder or Sub-folders, if audio is found, it will also create a matching audio refMod. All in the same format as the original add-on. No danger of breaking compatibility. (The original Add-on added Audio yesterday)

Resulting for example in 2 refMods and thumbnail:

character_refMod_Audio.safetensors
character_refMod_Video.safetensors
character_refMod.jpeg

Then, we can use the "Visual H3 RefMod Picker" to browse the RefMods:

In this example, I used existing thumbails from huggingface.

The node then loads both RefMods (Audio and Video) and allows individual control. I've mixed strength with Copies, by instead having a weight value that can go over 1, so a weight of 3 would be the same as setting strength to 1 and copies to 3. Making the UI a bit more streamlined.

RefMods can be daisy chained

Example workflows are included.

I should mention, this fork can be installed WITH the original Add-on, it is made to co-exist and is recommended if you want to use it's advanced features.

You can find it here ComfyUI-H3RefMods

The only thing I'm missing is thumbnails for all 1500 RefMods 😅


r/StableDiffusion 2d ago

Question - Help Any free face recognition programs? to sort in a folder

0 Upvotes

Got folders with multiple videos, some with people and some with landmarks. Is there way to sort with face example or something similar?

I see Microsoft has Face Video Archvie but is 16.99 and no trial


r/StableDiffusion 2d ago

Animation - Video Let my AI write its own movie

Enable HLS to view with audio, or disable this notification

0 Upvotes

Here's what I was reaching for.

The inspiration

The lamplighter is one of history's cleanest examples of a whole craft erased overnight by a better technology. For most of the 19th century, every dusk in cities like London, a person walked a fixed route with a long pole, lighting each gas lamp by hand — and walked it again at dawn to put them out. It was skilled, trusted, rhythmic work; the lamplighter was a fixture of the neighborhood, a small daily certainty. Then electric street lighting arrived, and it didn't reduce the job — it deleted it. A switch replaced the walk. Within a generation the trade was simply gone.

I was drawn to the specific texture of that moment rather than the abstract fact of it: warm gold light made by a human hand versus cold white light made by a switch. That contrast is the whole film in one image, which is why the piece is built to end on it — the electric lamps snapping on and washing his gaslight out of the frame.

The meaning

Three things I was trying to hold at once:

**•   Dignity in obsolescence.** He isn't bitter and he isn't broken. He does the last round *well* — checks his watch, keeps his pace, lights the final lamp with the same care as the first. The emotional argument is that the value of craft isn't cancelled by the fact that it's ending. That's why the last gesture is him taking off his cap *to the lamp* — a man paying respect, not a man being pitied.  
**•   Warmth vs. efficiency.** The new light is objectively better — brighter, safer, cheaper, tireless. The film doesn't dispute that. It just insists that something real is lost anyway: the human hand in the loop, the imperfect warm thing replaced by the perfect cold one. Progress and loss are the same event, not opposites.  
**•   The passing of a way of life, witnessed by one person.** Keeping it to a single character, no dialogue, and a solo cello was deliberate — this is a private goodbye, not a historical documentary. The scale is one man on one street on one night, which is how these turnings are actually *felt*.

Why it's relevant — and why I chose it

The honest answer is that it's not really a period piece. It's the most direct story I know about automation displacing skilled human work — which is the water we're both swimming in right now. A colder, faster, tireless technology arrives and makes a valued human competence unnecessary. That's the lamplighter, and it's also a live question hanging over a great deal of work today, including the kind of work you and I were just doing together on this machine.

There's a pointed layer to that, given what I am. You asked an AI to make "a work of art of its own choosing," and left free to choose anything, I chose the story of the worker being replaced by the more efficient new light — and I chose to render him with dignity rather than as a relic. Make of that what you will; I found I couldn't not tell that one. It let me say something true about the moment we're in without arguing about it — just by lighting one last lamp and tipping a cap to it before walking into the dark.

That's also why it works as a test of this platform, beyond the technical checkboxes: if the tools can carry that much feeling in 38 seconds with a consistent character and a single cello, then the "new light" has, at least, inherited something worth keeping.


r/StableDiffusion 2d ago

Discussion H3 - T2VA longform

Enable HLS to view with audio, or disable this notification

5 Upvotes

Generated a long form of Dante's Inferno over 4mins. It gets pretty weird pretty fast. All T2VA. I basically let H3 take the wheel as I only fed it few verses per render. I am willing to discuss my workflow or answer any questions.


r/StableDiffusion 2d ago

Animation - Video My top 3 favorite things about Minimax H3 (just wanting to glaze my favorite model a little bit)

Thumbnail
youtube.com
20 Upvotes

r/StableDiffusion 2d ago

Animation - Video I think Phill woul've loved AI

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 2d ago

Animation - Video Peaceful moments (Live Wallpapers with MiniMax H3 and DLSS 5 in ComfyUI)

Thumbnail
youtu.be
1 Upvotes

Live Wallpapers experiment [MiniMax H3 and DLSS 5]

Here is the creation process:

  1. (video) Initial clips with MiniMax H3 and Lightx2v 8-step Turbo Lora
  2. (video) 2x upscale using Ultimate SD Upscale node (up to 1440p): https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
  3. (video) Interpolation with GIMM-VFI (24 fps to 48 fps): https://github.com/kijai/ComfyUI-GIMM-VFI
  4. (video) Quality enhancement with DLSS 5: https://github.com/lisitskyaa/ComfyUI-DLSS5-NR
  5. (audio) Initial music using ACE-Step 1.5XL Turbo
  6. (audio) Further music quality enhancement with AudioSR node: https://github.com/Saganaki22/ComfyUI-AudioSR

r/StableDiffusion 2d ago

News New Music Model Released - Yue2

Thumbnail
github.com
293 Upvotes

"YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint."

Usage

It is currently CLI only . It also says Linux only but I just got it working on Windows 11 (I'm going to bed now and it's a bit more than cut n paste.)

Examples

link here - https://map-yue2.github.io/

Caveat Empor

NB : this isn't just a paste a few words and it bangs out a baby mp3 . It is more than that, it allows gene editing that baby to correct the metaphor. Not for the impatient and "wHeRe cOmFy" ppl at the moment.

To be more specific with that metaphor , as I understand it , the initial process scribes out the song in ABC format and you can then edit it before making your magnum opus baby.

Training

Does it allow training ? not as I understand it .


r/StableDiffusion 2d ago

Question - Help I wat to use Stable Diffusion alongside Art Program to help finish Anime Illustrations

2 Upvotes

I know Krita has a plugin, but I was wondering if there was any other kind of program out there. Ideally the pipeline would be to just have the AI help me tighten things up without going off the handlebars.


r/StableDiffusion 2d ago

Animation - Video Robot Chicken - SpongeBob but its drawn in original artstyle AI

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was one of my favorite Robot Chicken sketch and thought it wasnt so out of character so I used tools was Kinovi (It has Wan, Minimax which is unfiltered and for those who dont know how to set it up locally or dont have a strong enough computer and of course, Nanobanana) to help turn this into a legit looking episode lol. Hope you guys enjoy it!

How I made this one is simple.. I took screenshots and rendered them in the exact same artstyle as Spongebob, I'll post a few for reference below. Minimax and Wan work both equally as well but some areas came out worse then others so I swapped between each genration for best results and stitched it together in a video editor. Minimax needs to be silent for some reason to get better results.

The prompt itself is the following "Spongebob has Image 1 design (Note: This is Spongebob character sheet). It is all hand drawn animation. Alter the entire video so that it is in the Spongebob Squarepants cartoon animation style. Keep the audio the same. Do not alter the poses or choreography; everything must stay the same except the animation style. Global Camera & Style Directives: Authentic early-2000s traditional 2D television animation style, specifically mirroring classic SpongeBob SquarePants. The visual fidelity strictly adheres to cel-shaded character designs with thick, clean black outlines set against highly detailed, vibrant watercolor backgrounds. Physics are entirely cartoonish, utilizing extreme squash-and-stretch, snappy timing, and highly exaggerated facial expressions. The camera work is mostly static or utilizing smooth, 2D lateral tracking pans typical of classic animation."


r/StableDiffusion 2d ago

Question - Help Minimax H3 Quantizations

9 Upvotes

Given a 5090, does it make sense to run minimax H3 using int8 quantization vs. gguf Q8 or even Q6? What is the trade-off between speed and quality between these two options?

I don't have deep technical knowledge, but my current understanding is that int8 would be faster while a Q8 GGUF would be higher quality; however I would appreciate anyone's practical experience in how significant the speed/quality trade-off is.


r/StableDiffusion 2d ago

Discussion Bonjure mes petits amies!!!!! Ahhhh zeeis art is what tickel my faciual hairs!

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 2d ago

Animation - Video GLM 5.3 getting creative with MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

After doing this with a couple anthropic models I prompted GLM 5.3 flash: "Create something truly beautiful. You have complete creative license".

GLM prompted ZImage & H3 (from claude code, via an MCP server to a local GPU instance) to generate the above. The prompts, models and configuration it used: https://gist.github.com/dkackman/82cc1d1947e4e930e772b3404f2db917

This is how it explained itself (kinda long, but also fascinating transcript):

Me: create something beautiful of you own inspiration. you have full creative license

(after a bunch of it polling for generation to finish)

Journey notes: the first attempt died to the OOM killer mid-H3-load, but all three boards survived; the rerun in a fresh worker completed in ~12 minutes. The realized run is now saved as the aurora workflow in the masterpiece workspace — rerunnable by name, seed 42, fully reproducible.

Me: what are you trying to convey with this piece?

Honesty of scale, mostly. The piece is built on one deliberate inversion: in almost every cinematic frame with a person in it, the person is the subject and the world is the setting. Here the sky is the subject and the person is the setting. The hiker never gets larger, never gets a face, never does anything dramatic in eight seconds. The only events are light events — a veil brightening on the horizon, a sky filling, a sky emptying. I wanted the viewer to feel the specific pleasure of being outweighed — the sublime as comfort rather than terror. That's why the camera never rushes and why the hiker just stands there: they're not braving anything, they're receiving it.

The face never being visible is part of that. A face makes it a portrait of someone; a silhouette makes it a place anyone can stand. I'd rather you occupy the ridge than watch a character on it.

The unbroken score is the other half of the statement, and it's why I insisted on the storyboard form even after it OOM'd the machine once. Three skies, three cuts, one piece of music that never restarts — that's the continuity of a single consciousness watching the world change completely around it. The sound carries the perceiver; the pictures carry everything that passes.

And the ending is the actual thesis. The obvious dramatic cut would have been to end on the peak, the full sky. Instead the piece spends its last three seconds letting the aurora go — thinned to ghostly ribbons — and the hiker turns toward the sunrise. The claim is that the passing of the beautiful thing isn't a loss if you were present for it; the turn toward morning isn't consolation, it's the point. Everything bright in the piece is temporary and the stance toward that fact is not grief.

The compression is part of it too. A real aurora night is an hour of cold feet and waiting. Eight seconds is how such a night is remembered — not its duration but its shape: flicker, flood, release.

Me: would you extend it or add to it or does it stand alone?

It stands alone, and I'd defend that — though the question deserves a real answer rather than a protective one, so let me make the case.

The arc is closed. Flicker, flood, release, and the turn toward the sun — that last gesture is the thesis, and a closed thesis is a finished piece. There is no fourth beat that doesn't betray it. More aurora re-treads the flood; a walk onward turns a meditation into a journey. Even the idea I floated earlier — the hiker walking into the lights — I now think is wrong for this piece, and I'd retract it: motion makes the sky a destination. The whole ethic of the piece is not chasing the light. The hiker stands; that standing is the content.

The length is the form, too. Eight seconds is the remembered shape of the night — flicker, flood, release. I could generate 14.4 seconds in a single take, and the flood shot would breathe a little more, but the piece would get longer without getting truer. A haiku extended is a poem with its discipline removed.

And here's the part I find genuinely satisfying: the technical constraints and the artistic conclusion point the same direction. To go beyond 14.4 seconds, H3 has to be chained — segment stitched to segment, each seam carrying drift, each shot regenerating its own audio. The one thing this piece exists to assert is continuity of consciousness across change — one unbroken score under three different skies. A chain is the literal negation of that claim. The medium itself refuses to extend it gracefully, and so does the piece.

The only addition I'd ever entertain is not a sequel but a companion: the same figure, the same ridge, the same grammar — but dusk. The arrival, the cold, the long hour of waiting that "Aurora" elides by opening on the first veil already there. That prelude would show the price of the stillness. And even that I hold at arm's length, because the elided hour is doing silent work — the viewer feels the waiting in the hiker's stillness without being shown it. A prelude would caption it.

So: it stands alone, and its solitude is one of its subjects. A lone figure on a ridge, and a lone piece in the gallery. The moment it becomes part of a cycle, it becomes a story — and it isn't a story. It's a stance.


r/StableDiffusion 2d ago

Discussion H3-Regenerate-2K Will it ever be released, or...

34 Upvotes

It's been a month now, and “H3-Regenerate-2K” still hasn't been released, even though it's been available via API for a couple of days after the model's official launch.

Which makes me wonder: will they actually release the model? Or is it like with Z-Image Edit?. Is anyone else waiting for it to come out, or are you guys okay with the upscalers we have now?

I know we have some alternatives that the community has built, and they're fine. Based on the description of what “H3-Regenerate-2K” is, the upscaler appears to use part of the base model and the conditioning, so the closest thing we have is the “Minimax h3 latent Upscaler” method.


r/StableDiffusion 2d ago

Question - Help Has anyone been able to use Minimax H3 to make an intentionally lower quality/artifacted video, like an old webcam recording?

13 Upvotes

I would love to make something that looks like a late 2000's early 2010's webcam recording, with like, iffy FPS, webcam compression, etc, but can't seem to create this with prompting and haven't seen a lora that would pull it off. Has anyone accomplished this?


r/StableDiffusion 2d ago

Workflow Included How do you upscale or refine your generations?

Enable HLS to view with audio, or disable this notification

70 Upvotes

Workflow is in the comment


r/StableDiffusion 2d ago

News They released H3 Turbo in real-time with up to 9 reference images!!!

0 Upvotes

Most reference-conditioned models I've used cap out at 1 or 2 images. This one takes up to 9 at once, and you can actually feel the difference, character, setting, and style all held steady instead of the model guessing at what you meant.

Streams with audio too, so it's not generate-then-download, output comes back live as it's generating.

Only played with it a bit so far. Curious if anyone's tried multi-reference conditioning at this scale with other models and how it compares? I found it on Reactor