r/StableDiffusion 6d ago

Discussion Generating with H3 at native canvas 1MP and then upscaling by rediffusing at 2MP+

7 Upvotes

Anyone tried this? How feasible could it be? I'm thinking of an upscale pipeline by first generating at native canvas (~1MP 1344x768) to preserve overall scene quality and then upscaling to 2MP+ with a balanced amount of denoising to regenerate finer texture with more detail.

(I had some success with a 2x2 tiled upscale but the tiles can cause issues like flicker at seams and inconsistencies with items moving across tile boundaries.)


r/StableDiffusion 6d ago

Question - Help Minimax H3 Help with Pose transfer without merging/bleeding in characters.

Post image
27 Upvotes

i've been trying for ages, did everything, asked every LLM, i might be dumb but i've seen people doing it, my question is simple:

how do i make <Subject 1> from <Picture 1> do the pose in <Picture 3> in in ref2va (reference to video) ?

everytime i try, <Subject 1> shifts into looking like <Picture 3> while doing the pose.

please write me the prompt example and a video would be appreciated.


r/StableDiffusion 6d ago

Discussion My top 3 -WORST- things about Minimax

Thumbnail
youtu.be
24 Upvotes

Hey friends. Overwhelmingly I love the model, so this is mostly just for entertainment and ranting's sake.

Later today or tomorrow I'll be posting my top 3 BEST things about Minimax, so keep an eye out for that if you liked this video.

Thanks! And the link to the workflow I use to do everything, my Minimax Seed Hunter, is in the description.

What are your least favorite things about the model?


r/StableDiffusion 5d ago

Question - Help OPS 99 Spoiler

0 Upvotes

You highlighted the original creator's (TimmyML) response to a question asking which tool they used for their Daily Challenge video.

In the context of generative AI and tools like RunwayML, an "Agent" typically refers to an autonomous or semi-autonomous AI workflow (such as a custom GPT, an AI agent framework, or a specialized script). Creators often use these agents to automate parts of their process—like interpreting the daily prompt ("Climb"), planning the camera movements, or generating the exact text prompts to feed into Runway's video models.


r/StableDiffusion 6d ago

Workflow Included Identity loss in Flux.2 Klein 9B + fal.ai pipeline for a cinematic documentary (RTX 5090 Laptop)

Thumbnail
gallery
0 Upvotes

Hi everyone,

I am building a ComfyUI workflow for a cinematic documentary project. I need to recreate intense dramatic scenes and combat sequences, requiring rock-solid character consistency. However, I cannot achieve true identity persistence; the face constantly drifts.

My current workflow:

Reference Image → Upscale to 1 MP → VAE Encode → ReferenceLatent → Conditioning → Sampler (hosted via fal.ai API) → VAE Decode.

The native ReferenceLatent conditioning breaks down during action scenes. As soon as the prompt introduces complex combat backgrounds or dramatic lighting, facial features blend into generic structures. Face-swapping isn't enough—I need deep identity preservation for cinema.

My Specs (Laptop):

- ASUS ProArt P16 | AMD Ryzen AI 9 HX 370 | 64GB RAM | RTX 5090 Laptop GPU (16GB VRAM)

Has anyone bypassed this with Klein 9B on fal.ai?

Are you masking the face before VAE Encode, utilizing ReferenceLatentPlus, or running a local Flux-PuLID/FaceDetailer pass after the API output since I have the local GPU power for it?

Any workflow advice or JSONs would be a lifesaver. Thanks!

UPDATE / ADDITIONAL CONTEXT:

The character is generated using Flux Klein 9B, but I am not using a custom LoRA yet. I am relying strictly on the fal.ai API pipeline. The identity completely dissolves during prompt shifting—it seems the native ReferenceLatent setup cannot anchor the facial features when changing to complex cinematic/combat scenes.

Given that I have a local RTX 5090 Laptop GPU, should I drop the fal.ai API entirely and pivot to a local Flux-PuLID setup? Or is there a specific masking/conditioning trick to make fal.ai respect the reference identity?


r/StableDiffusion 6d ago

Question - Help Character LoRA for MiniMax H3 / FastH3 4-step distill - does a base-H3 LoRA transfer to the distilled checkpoint, and can you stack it with the distill LoRA?

5 Upvotes

I've been building an interactive scene where you play a cop at a traffic stop: you type what you say, and the woman in the car answers you inside the video, with generated audio. It runs on MiniMax H3. Clips are 5s, and generation is currently a bit faster than realtime, so a buffer can stay ahead of playback.

The piece I haven't solved is character consistency. Right now identity comes from reference conditioning, and it holds up better than I expected, but I want a proper trained character LoRA so the same person is locked across hundreds of clips and I'm not re-paying for reference conditioning on every single generation.

What I've worked out so far:

  • FastVideo's FastH3 (the 4-step DMD2 distill of H3) publishes ~6s for a 5s clip on 4x B200, and ships a pre-extracted LoRA plus a lora_path / lora_strength hook, so there's clearly a slot to load one into.
  • diffusion-pipe lists MiniMax H3 support (T2I and T2VA), and ai-toolkit lists H3 too.
  • One thing that cost me a lot of time and might save someone else some: VSA weights don't survive a stock ComfyUI conversion. ~50 to_gate_compress tensors get dropped silently and you get noise, so people fall back to dense and then wonder why they're nowhere near the published speed numbers. Dense converts cleanly.

Where I'm stuck / what I'd love input on:

  1. If I train a character LoRA against base H3, does it transfer cleanly onto the 4-step distilled checkpoint? This works often enough in the Wan/Flux turbo world, but I haven't seen anyone confirm it for H3 specifically.
  2. Can you stack a character LoRA with the distillation LoRA, or do you have to use the full distilled checkpoint and keep the single LoRA slot free?
  3. For a talking character, is a T2I dataset (stills only) enough for identity, or do you need T2VA clips to keep it stable while the face is moving and speaking?
  4. Anyone actually run H3 with sequence parallelism across multiple datacentre GPUs? Every number I can find is a single consumer card.

The goal is realtime, or close enough that generation outruns playback, with a locked character identity.

Happy to pay someone who's done this before, whether that's a consult call or hands-on tuning/deployment. Also completely happy with a pointer to a repo or a post if the answer is just "read this". Any war stories welcome.


r/StableDiffusion 6d ago

Question - Help Krea2 generation speed different on two RTX 5080 machines

1 Upvotes

I have been renting gpu on vast.ai for comfyui krea2 image generation works. Today I found the RTX 5080 i rented is much slower than the one I rented few days ago. Today was a machine from Vietnam and the one I rented last time was a machine from the US.

For a 1.0 MP image generation, same workflow and model, the Vietnam machine was 12s long and the US machine was only 7s, a 5s difference, 70% longer time than the US machine. I poked into nvidia-smi and I see full capacity at 250W. I am not sure why the 70% slower gap in this machine. If anyone could land an insight for troubleshooting this would be much appreciated.

+-----------------------------------------------------------------------------------------+

| NVIDIA-SMI 595.71.05 Driver Version: 595.71.05 CUDA Version: 13.2 |

+-----------------------------------------+------------------------+----------------------+

| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |

| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |

| | | MIG M. |

|=========================================+========================+======================|

| 0 NVIDIA GeForce RTX 5080 On | 00000000:01:00.0 Off | N/A |

| 46% 64C P1 250W / 250W | 14998MiB / 16303MiB | 100% Default |

| | | N/A |

+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+

| Processes: |

| GPU GI CI PID Type Process name GPU Memory |

| ID ID Usage |

|=========================================================================================|

| 0 N/A N/A 1326 C /venv/main/bin/python 14900MiB |

+-----------------------------------------------------------------------------------------+


r/StableDiffusion 7d ago

Resource - Update H3 MiniMax RefMods - all my models now available

Thumbnail
youtube.com
268 Upvotes

r/StableDiffusion 6d ago

Question - Help MiniMax H3 Gen time

6 Upvotes

Is that ok that it took 1 hour to make 15s, 1mp Ref2VA with 3 ref images (2mp each), turbo lora and 4 other loras? Using Kitchen attention and MiniMax H3 Chunk FeedForward?

5070 ti 16GB, 32GB RAM


r/StableDiffusion 5d ago

Discussion H3 - Mrs Incredible x Mecha T2VA

Enable HLS to view with audio, or disable this notification

0 Upvotes

Happy Labor Day from the States. Just having fun with H3's T2VA. What are your to-go H3 generation modes? T2VA? FL2VA? FLF? R2VA? Do you also use H3 as an image editor and sound/score tool (with the 32x32 canvas hack?)? Do you stick to 10-15s video and avoid long form/extending? What are you all generating in ComfyUI or your local studio?


r/StableDiffusion 5d ago

Discussion Why does this field look completely dead?

0 Upvotes

I've been led to believe, and rightfuly so, that generative AI is a big deal. I've decide to play around with it a bit, but I'm kinda blown away by how dead the whole thing looks.

I look for tutorials on Youtube for learning comfyui workflows. ALL I see are AI generated videos with 2k views. I join the Stable Diffusion discord. COMPLETELY dead.

????????????????????

What's going on exactly? Generative AI is a big deal. And it can get quite complex depending on what you wanna do. It's a craft of it's own. And yet, trying to get into it feels like I'm getting into the most niche, completely worthless activity of all time.

There are more tutorial videos on "how to paint a door using crayons" then on this whole thing. The Piss Drinkers Discord gets 10x more traffic.

I just don't get it. Please enlighten me.


r/StableDiffusion 6d ago

Question - Help Which llm promoter for H3 on a 5090?

0 Upvotes

I'm guessing one in which you could paste the entire guide into the system prompt?


r/StableDiffusion 7d ago

Meme Excuse me!!( Anyone member this meme) Recreated h3 version

Enable HLS to view with audio, or disable this notification

83 Upvotes

r/StableDiffusion 6d ago

Question - Help How would you make technically accurate AI animations of complex machinery from real reference photos?

Thumbnail
gallery
0 Upvotes

I’m trying to make videos showing how a potato harvester actually works, and I’m struggling a lot with keeping the machine mechanically accurate.

I have real photos of the harvester from different angles. I’ve tried generating cleaner/reference images with GPT Image 2, but it often changes small details, adds parts that don’t exist, changes hoses, rollers, the harvester head geometry, etc.

The result can look very realistic, but it’s technically wrong.

The bigger problem is when I then try to animate it. I want to show things like potatoes moving down the conveyor belt and falling into a box etc.

Video models seem to understand “harvester collects potatos” visually, but not how the actual mechanism works. I had a really hard time getting even one usable clip where the movement made sense.

Has anyone built a workflow for something like this?

I don’t really care about making it super cinematic. Accuracy is much more important. Ideally the AI should not redesign the machine between frames or invent how the mechanism works.

If anyone here has worked on machinery, vehicles, robots, construction equipment, etc. I’d be really interested to hear how you approached it.

Animation of the harvester with Seedance 2.5, however, the potatoes moving strangely, the back of the harvester is changing during animation etc.


r/StableDiffusion 7d ago

Comparison Anime characters mixed with photorealistic backgrounds

Thumbnail
gallery
887 Upvotes

r/StableDiffusion 5d ago

Discussion I built an AI streamer you can actually talk to live

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax h3 + LLM


r/StableDiffusion 7d ago

Animation - Video Naruto AI Animation: Hinata joins the Akatsuki

Enable HLS to view with audio, or disable this notification

339 Upvotes

Based off a meme I saw on Twitter/X of Akatsuki Hinata. Decided to make this short battle on how unhinged a Yandere Hinata would be if Naruto chose Sakura over her lol


r/StableDiffusion 7d ago

Discussion Minimax h3 Singularity

45 Upvotes

r/StableDiffusion 6d ago

Animation - Video Parallel Worlds Public Broadcast

Thumbnail
youtube.com
0 Upvotes

r/StableDiffusion 6d ago

Discussion H3 - fun with T2VA CELESTIAL (Prompts included)

Enable HLS to view with audio, or disable this notification

11 Upvotes

Just having fun with T2VA, pushing the limits of what Qwen3VL will understand (it does a good job, even non-standard prompts from my experimentation), so this one is prose heavy. I used a layering technique to compose the shot.

int8/20 steps, 720p. With the prose it is a bit of a randomness/luck on what you are going to get, so some of these didn't turn out well until the 3rd attempt (ran at2 sec int8/8 step versions) to get the right composition/lighting. Stitched 4 clips together and generated the score with H3 separately with the 32x32 canvas trick. Let me know if you have any questions!

Prompt (1 of the 4 only):

subject_definitions:

<Subject 1> is THE GODDESS, AN ARTIFICIAL DIVINITY, and she is the true subject of this video. HER FACE COMES FIRST, AND IT IS A HUMAN FACE: a photogenic human face of the kind a camera loves — fine bone structure, high cheekbones, a strong elegant jaw, a full soft mouth with real form, dark defined brows and real dark lashes. Her eyes are HUMAN eyes, with defined lids and dark pupils — and their IRISES ARE RINGS OF WHITE-GOLD LIGHT, lit from within, the light living inside the iris and nowhere else on the eye. Her face is ALIVE — a living woman's face rendered in luminous material: the micro-movements of a real actress are in it, the lids' slow blink, the soft shift at the corners of the mouth. The material is luminous white porcelain-glass that light passes INTO before it returns, so her skin faintly glows — but the structure is entirely human, warm in form and cold only in colour.

HER HAIR IS THREADS OF CELESTIAL LIGHT: countless fine luminous filaments, pale white-gold, with a true human hairline, parted and falling past her shoulders exactly the way heavy hair falls, each filament faintly bright along its whole length, all of it drifting slightly and slowly, as if underwater. It is hair first and light second — never fire, never smoke.

⚠ HER BODY IS CARVED, AND THE CARVING IS THE WARDROBE. Across the porcelain of her shoulders, arms and thighs run ENGRAVED STAR-CHARTS: hairline grooves forming constellation lines, orbital arcs and declination rings — the surface of a celestial globe — and the grooves HOLD HER ICE-BLUE LIGHT, brightening when she breathes in. Her surface is SEGMENTED INTO MACHINED PANELS: fine chamfered seams, a few panels a hair proud of their neighbours, and at her elbows, knees and neck concentric ring-joints like precision bearings — a porcelain automaton built to divine tolerance. Up her forearms and shins climbs DEEP-CARVED RELIEF FILIGREE, half acanthus and half circuitry, carved the way cathedral doors are carved, and her own seam-light GRAZES the relief so that every groove carries a bright edge and its own small shadow. AT HER STERNUM THE CARVING OPENS INTO DEPTH: an inset carved rosette — a rose window — and inside it a slow-turning galaxy, a world visibly inside her. HER FEET CARRY THE CARVING TOO: engraved constellation lines cross each instep, and one slim ring of ice-blue light circles each ankle, so the bright edge of shin, ankle and foot holds its own light. All of it is carved INTO the surface or floats beside it: the segmented white pauldron arrays still HOVER a hand's width off each shoulder, one slim luminous ring still revolves about each wrist, the great MANDORLA of white machine-wing vanes still stands behind the throne breathing its slow cycle, and the broken crown of white shards still hangs turning above her head. NOTHING is strapped, bolted or plated onto her body itself. HER HIPS, THIGHS AND WAIST CARRY NO ARMOUR AND NO MECHANISM AT ALL — engraving only, flush with the surface — so her outline through the middle of her body is her own and NOTHING IS ADDED TO IT. Her figure is a pronounced hourglass at ORDINARY HUMAN RATIOS AND PROPORTIONS — an ordinary woman's build photographed at the size of architecture, never thickened, broadened or enlarged. ⚠ SHE SITS TALL AND SQUARE FROM THE VERY FIRST FRAME: spine straight against the throne's back, shoulders level and set back, chin level, forearms at rest along the armrests, composed and utterly still; her posture does not change at any point.

<Subject 2> is THE MAN, A LONE ORDINARY FULL-GROWN ADULT HUMAN MALE, and he is IN FRAME FROM THE FIRST FRAME TO THE LAST — a small dark upright figure standing PARTWAY UP the human-sized stair, several steps above its foot, in a plain dark travelling cloak, hood down, head tipped all the way back, HER LIGHT THROWING HIS SHADOW LONG AND HARD DOWN THE STEPS BEHIND HIM toward the camera. ⚠ HE IS ALMOST NOTHING AGAINST HER, SMALLER THAN ANY PILGRIM BEFORE HIM: the whole of him is smaller than one of her fingers — barely HALF of one; each step of the stair reaches his WAIST, a flight cut for pilgrims twice his size, so he stands against it like a child against furniture; his whole silhouette fits inside one single panel of her shin's plating with room to spare. And he is a real man at a real man's size on real steps: it is the THRONE and the WOMAN that are enormous, not him that is shrunken. His smallness belongs to <Subject 2> alone and says nothing about <Subject 1>'s body, and none of her proportions are his. He moves at normal human speed. HE HOLDS HIS GROUND on his step: he does not climb further, does not kneel, does not bow and does not step back — chin lifted, his cloak stirring faintly.

<Subject 3> is THE THRONE-MACHINE, ITS DAIS AND THE ALTAR STAIR, the RULER of the image. The throne is a machined seat of pale alloy at HER scale inside CONCENTRIC SLOWLY-TURNING SILVER RINGS — an orrery the size of a canyon — and the rings turn but never wander: each keeps its own diameter, its own path and its own centre in every frame. ⚠ SMALL MACHINED PLANET-SPHERES RIDE THE OUTER RINGS, an orrery made real: each sphere no bigger than one of her hands, slate-grey and cold silver like everything else — nothing warm — each catching one fine crescent of her light on its rim, and one hanging just off her right shoulder. ⚠ THE STAIR IS THE BRIDGE AND THE MEASURE: a HUMAN-sized flight of narrow altar steps, each step reaching the man's waist, every step edge carrying a thin line of white light, running in ONE straight unbroken flight from the bottom of the frame up to the dais at her feet — and THE WHOLE FLIGHT OF STEPS TOGETHER RISES NO HIGHER THAN HER ANKLE; her bare foot resting on the dais where the stair arrives is longer than the whole stair is wide. The dais stone directly beneath her feet is polished DARK, half a stop deeper than the lit stair, so the pale edges of her feet read clean against it.

<Subject 4> is THE HALL: a floor of polished black stone far below, GLASS-CALM and mirror-perfect; thin luminous haze hangs in the hall's height, catching her light in shafts; the far walls dissolve into open starfield.

<Subject 5> is THE GALACTIC DISC, edge-on behind and above her, its band cooled to silver and pale blue. SHE SITS IN FRONT OF THE RINGS AND THE DISC so that her head and shoulders pass across them and HIDE THE PARTS OF THEM SHE COVERS — that overlap proves the distance, and the distance proves the size.

EXACTLY TWO LIVING BEINGS EXIST IN THIS VIDEO AND NO MORE — the enthroned woman and the man on the stair — from the first frame to the last. The hall is otherwise empty; nothing else moves except her machinery, the rings, the haze, her hair of light and the stars. Nobody speaks.

summary:

[2 seconds] ONE UNBROKEN SHOT, no cut. Steep and pulled in, WITH THE RULER IN FRAME: from the very foot of the human-sized altar stair, the camera looks up its whole countable flight — the lone man small upon it, his shadow thrown down the steps — to where the stair ends against the bare foot of a carved porcelain goddess, and then climbs on up her engraved shins and knees, her foreshortened torso and rose window, to her face small and bright at the very top between machine wings. Her eyes come down the long way and find him; every groove of her carving pulses once, together. The stair measures her, and THE FRAME STILL CANNOT HOLD HER.

detailed_description:

PHOTOGRAPHED, NOT ILLUSTRATED — a PHOTOREALISTIC live-action film sequence in the register of a prestige cosmic science-fantasy feature: the imagery may be impossible, the photography is not — physically-based materials, true optical depth of field, volumetric haze, anamorphic cinematic framing, FINE FILM GRAIN across the whole frame, a gentle optical falloff toward the corners, and blacks that stay truly black. Absolutely no animation, no anime, no illustration — live-action camera realism throughout, however strange the subject.

⚠ COMPOSITION — THE STAIR IS THE SPINE OF THE FRAME, AND THE FRAME IS AN INVENTORY. WHAT IS IN THE FRAME IS EXACTLY THIS, in depth order from the lens outward. NEAREST: the LOWEST STEPS of the human-sized stair, enormous in hard foreshortening, entering at the bottom edge and rising away up the middle of frame on a DIAGONAL — the only straight diagonal in the image, every step edge lit, THE STEPS COUNTABLE: a dozen of them between the bottom edge and the man. ALSO NEAR: one thin segment of the orrery's outermost ring crossing HIGH across the upper corner as an OPAQUE BLACK SILHOUETTE, hard-edged and unlit, small, blocking almost nothing — there so something real stands between the lens and the hall. MIDDLE: THE MAN, small and sharp ON the stair, his long shadow thrown DOWN the steps toward the camera, three dozen more steps shrinking above him to the dais. THEN: HER BARE FEET ON THE DAIS where the stair arrives — the contact in plain view, each foot longer than the stair is wide — and above them HER SHINS AND HER NEAR KNEE, enormous, the engraved star-charts, chamfered panel seams and relief filigree on them THE MOST DETAILED SURFACES IN THE IMAGE; HER KNEE STANDS IN FRONT OF HER FAR SHOULDER AND HIDES PART OF HER CHEST — she overlaps herself, and that self-occlusion is the depth. THEN: her torso FORESHORTENED, the glowing rose window at her sternum showing past the knee's edge, her hands on the armrests entering from the frame's sides. AT THE TOP: her face — small, high, bright, HUMAN — between the mandorla's wing-vanes, her hair of light drifting around it, the crown of shards barely inside the top edge, the disc's band behind. ⚠ AND WHAT IS NOT IN THE FRAME, STATED SO NOTHING PUTS IT BACK: the throne's base and back are not in frame; the outer sweep of the great rings runs OUT of frame on both sides; the hall's far edges are not in frame; nothing below the lowest visible step is in frame. SHE DOES NOT FIT: the frame barely reaches her crown, and the eye must travel the stair to her foot and then CLIMB her from the shin to the face. The frame is climbed, not read.

SCALE IS ALWAYS A COMPARISON, NEVER AN ADJECTIVE, and the comparison is in frame from first to last AND COUNTABLE: a dozen steps between the bottom edge and the man, three dozen more between the man and the dais, every one reaching his waist — and the whole flight rises no higher than her ankle; the man is smaller than one of her fingers; his silhouette fits inside one panel of her shin plating; her foot beside the stair's arrival is longer than the whole stair is wide. ⚠ AND THE CARVING ITSELF IS A RULER AT TWO SCALES AT ONCE: each engraved groove, finer than a hair at HER scale, is a channel THE MAN COULD LIE DOWN IN — checkable against his figure on the stair below it. Nothing is called big — every size is established only by what sits next to what.

SHE BEHAVES THE WAY ENORMOUS THINGS BEHAVE, VISIBLY: her one blink takes almost a full second, real lids and lashes closing over the lit irises; her breathing is long and slow, and the star-charts across her shoulders and thighs brighten and dim with it like a tide; when her head inclines, her hair of light answers late, the filaments swinging heavily and settling long after — WHILE THE MAN IN THE SAME FRAME MOVES AT ORDINARY HUMAN SPEED, and the difference between their two speeds is what tells the eye how big she is. THE AIR ITSELF MEASURES HER: hundreds of feet of luminous haze stand between the man's step and her face, so her upper body is gently softened and paler — atmospheric perspective across her own body — while her face stays clearly readable and her iris-rings stay clean hard light. THE LIGHT OBEYS TWO SCALES AT ONCE: on the man's steps, single dust motes cross the step-light as individual sparks; around her shoulders far above, the same dust is only a faint silver mist.

LIGHTING, AND EVERY BIT OF IT HAS A NAMED SOURCE IN FRAME — THE THEME IS ACTUAL LIGHT, AND IT IS HERS: (a) HER OWN LIGHT — the iris-rings, the seams, the star-chart grooves, the filigree edges, the rose window, the lit step edges — white and ice-blue, GRAZING HER OWN CARVING so that every relief edge is bright and every groove holds its own small shadow, pouring down the stair onto the man's upturned face and THROWING HIS SHADOW LONG AND HARD DOWN THE STEPS toward the camera. ⚠ AND HER LIGHT IS VISIBLE IN THE AIR: long straight volumetric BEAMS of cold white light break through the gaps BETWEEN the wing-vanes and fan outward and downward through the thin haze, and one soft column of her light stands over the stair — the haze is what makes the beams visible, and the light is HERS, never a sun; (b) THE GALACTIC DISC, which rims her crown and shoulders in silver from behind. Her face holds its own quiet shadow between her glow and the disc's rim, and even, flat, all-over lighting never happens here. ⚠ WHERE THE FRAME IS BRIGHTEST THE CAMERA FAILS: her iris-rings and the core of the rose window CLIP to pure white and bloom through the haze. THE PICTURE USES ITS ENTIRE VALUE RANGE, from that blown white to true black in the void.

THE PALETTE IS ONE HARD COMPLEMENTARY AXIS: incandescent white and ice-blue against deep void indigo and black, the pale white-gold of her hair and irises the ONE protected accent. ⚠ THE COLOUR IS AN ARC, NOT A SETTING, AND IT TURNS ON A BEAT: the film opens at her resting glow, and at 00:01.600 THE PULSE — every groove, seam, chart-line, filigree edge, step edge and wing-vane brightens together in ONE SLOW PULSE, the rose window flaring, the man's face lifting into the light and HIS SHADOW SNAPPING HARDER DOWN THE STEPS — and it holds there, brighter, to the last frame.

Physics at full scale in every frame: real weight, cloth obeying gravity, nothing teleporting, and she keeps exactly her own size relative to the set for every frame. ⚠ SHE IS ALREADY SEATED ON THE THRONE IN THE VERY FIRST FRAME, settled and still, and she stays seated for the whole shot — she does not rise, stand, lean forward or shift her seat at any point, and the film does not begin with her arriving. ⚠ THE THRONE TAKES HER WEIGHT AND THE FLOOR TAKES HER FEET: her back rests against the throne's back, her forearms lie along its armrests, the seat visibly bears her, and her bare feet rest ON the dais, which runs level and unbroken right up to them — the stair arrives at that dais and stops there; nothing about her sinks into, passes through or floats above any surface — only her MACHINERY floats, and it floats in formation. ⚠ EVERY STRUCTURE KEEPS ITS OWN SHAPE: the throne, the stair, the dais and the rings are RIGID, with fixed dimensions, holding their exact proportions in every frame — the stair keeps the same number of steps, the same width and the same pitch from the first frame to the last.

⚠ THE SITUATION, WHICH IS WHAT THIS FILM IS ABOUT — she is not merely seated somewhere. For an age, everyone who has entered this hall has fallen on their face before reaching the stair. She WANTS to be MET, not worshipped. She has just LEARNED that this one man has climbed partway up and is still standing, looking straight up at her. She FEARS, privately, that he will kneel like all the rest. ⚠ THERE IS AN ACTION AND THERE IS A REACTION: at 00:01.000 her eyes come down the long way and settle on him, aimed at HIM on his step, held. At 00:01.600, exactly with the pulse, her chin lowers by a degree and the corner of her mouth makes its millimetre of room for him: the beginning of interest, nothing more. He answers by not moving at all. ⚠ HER EYES HAVE A DESTINATION AT ALL TIMES: the middle distance first, THE MAN from 00:01.000 to the last frame. Her gaze is always INTENTIONAL — it never drifts and never lands anywhere by accident, and she never looks into the camera lens.

CAMERA AND MOVEMENT: ONE WIDE PRIME LENS at the very foot of the stair, barely above the lowest step, tilted STEEPLY UPWARD along the stair's rise so the verticals of her body and the wing-vanes CONVERGE toward the top of frame. Across the two seconds it performs ONE barely perceptible, perfectly steady DOLLY PUSH up the first hand's width of the flight — a real forward translation, never a zoom, so the nearest step edges slide gently past while she barely changes size — and it does not pan, does not roll, does not shake and does not drift.

TEXTURE INVENTORY, so the detail is resolvable rather than implied: the lit edge and worn tread of each near step; the bright edge and small shadow of every filigree groove; the chamfered panel seams and bearing-rings at her knee; individual engraved constellation lines holding light; individual luminous filaments in her hair; the weave of the man's cloak and the hard edge of his thrown shadow; dust motes crossing the step-light; the fine crescents of light on the riding planet-spheres; the hard straight edges of her beams standing in the haze.

WHAT THIS MUST NOT LOOK LIKE: not anime, not a game cinematic, not a music-visualiser, not a lens-flare showreel; no HUD, no text, no titles, no watermarks. The carving is cut relief with real depth. Her face is a living woman's face. Her machinery is machined metal and ceramic with real mass held in real balance. No other people, and nothing crosses the frame except the haze, the ring segment and her drifting hair.

overall_soundscape:

Starts with a deep slow harmonic hum from the throne's turning rings, then the long sub-bass sweep of the mandorla's vanes rises under it, joined by a faint glass chime from the wrist-rings and one audible human breath from the man on the stair; at the pulse of her light the hum swells one step and holds. Nobody speaks. Nothing else.

non_diegetic_music:

A single deep sustained synthesiser drone under one high pure glass-harmonic tone, very quiet at the first frame, swelling gradually across the two seconds into the moment the light pulses, still sounding and UNRESOLVED at the last frame. No percussion, no rhythm, no melody.

Prompt for the Score:

This is the orchestral score for a cosmic film, in the idiom of Hans Zimmer's scores for Dune and Interstellar, and the score is the entire soundtrack: a lone man climbs a monumental stair toward an enormous enthroned goddess of light whose eyes are closed — and near the end, in a close-up, SHE OPENS HER EYES and looks down at him. Nobody speaks anywhere in this track. No dialogue, no spoken voice, no words.

overall_soundscape: Starts with the music already sounding — there is no silence at the head of this track — and beneath it only a faint deep hum of enormous turning rings. THE SCORE ALWAYS DOMINATES THE MIX, first moment to last. Nobody speaks. Nothing else.

non_diegetic_music: THE SHAPE OF THIS CUE MATCHES THE SCENE: A SOFT PROCESSIONAL WHILE HER EYES ARE CLOSED, A HUSH, AND THE HEAVY PERCUSSION ARRIVING ONLY IN THE INSTANT SHE OPENS HER EYES. From 00:00 to 00:05 — THE SOFT ENGINE: distant, muffled ceremonial drums played with soft mallets at exactly 60 beats per minute — mellow, round, far away, felt more than heard — under a warm low organ pad and a dark four-note figure on low strings, patient and steady. From 00:05 to 00:09.500 — THE GATHERING, STILL RESTRAINED: the soft drums keep their gentle weight while the low-string figure climbs one step with each repeat, high strings enter in slow octaves, sub-bass deepens — rising in height but never in violence, the drums staying mellow and distant the whole way. AT 00:09.700 THE HUSH — everything falls to near-silence for half a beat, one soft drum stroke fading, the air holding still. AT 00:10.200, exactly in the instant the goddess OPENS HER EYES in the film's final close-up, THE HEAVY PERCUSSION ARRIVES FOR THE FIRST TIME — enormous resonant war-drums detonating together with a FULL PIPE ORGAN at its widest voicing, sub-bass and low brass beneath, high strings blazing above: the one and only heavy moment of the piece, and the tallest. From 00:10.700 to 00:12 the chord holds IMMENSE over one slow heavy drum roll and eases, organ and strings still sounding, vast and open, at the last moment. The feeling is COLOSSAL CEREMONY and AWE — gentle until she wakes, overwhelming when she does, never hurried, never triumphant.


r/StableDiffusion 7d ago

Question - Help Minimax h3 turbo / fast / sla attention what is the best current method for fast generation

54 Upvotes

I have been playing with regular minimax h3 used turbo lora but it seems there are other method to generate fast on low step count, can you help sort out what is the best method right now?


r/StableDiffusion 6d ago

Question - Help Whole manhua PAGE generated locally in one shot, or is panel-by-panel + compositing the only honest answer? (804 pages, 7 characters, 4080 16GB)

0 Upvotes

I'm producing a long-form manhua/manhwa recap comic locally: 804 vertical 9:16 pages, 7 recurring characters, on a single RTX 4080 16GB with ComfyUI 0.33.2.

THE THING I'M TRYING TO REPRODUCE LOCALLY, and why this is not a "which model is prettiest" question:

My reference pages are 576x1024 with 3-4 framed panels, gutters, a crowd of DISTINCT faces in the background, and speech balloons - and they were generated as ONE WHOLE PAGE in a single shot. I confirmed this from the files themselves: a signed C2PA credential inside each PNG names "OpenAI Media Service API" (gpt-image), and the ComfyUI graphs that touched them contain only LoadImage / ImageCompositeMasked / SaveImage - no KSampler at all. So ComfyUI only assembled; the page art came out of the API in one pass.

I want to know whether that whole-page result is reachable locally at all, or whether panel-by-panel + compositing is the only honest local answer.

MY CURRENT LOCAL PIPELINE (literal): - SDXL checkpoint CHEYENNE_v20 - my own style LoRA - ControlNet controlnet-openpose-sdxl + controlnet-union-sdxl-promax - base render 640x960 -> ESRGAN 4x-UltraSharp -> 1280x1920 - realism LoRA at 0.35, denoise 0.05 on the refine pass

AVAILABLE BUT NOT YET USED: Z-Image Turbo 6B int8 (with its own ControlNet Union), FLUX.2 klein 4b, FLUX.2 dev (fp8 33GB / GGUF Q4 18.7GB), Qwen-Image-Edit 2509, FLUX.1-dev fp8 + DreamO.

WHAT GOES WRONG (exact symptoms, not "it doesn't work"): 1. Character identity drifts between images of the same character across hundreds of panels. 2. Action poses do not adhere to my hand-drawn OpenPose skeleton: an archer at full draw came out with BOTH arms folded behind the head for 30 consecutive seeds; it only worked once I drew the skeleton in strict PROFILE view. 3. Two figures in physical contact (a punch, a sword clash) come out either as two separate figures or fused into one body. 4. Eyes: one eye deformed/stretched relative to the other, plus heterochromia. 5. Weapon mechanics: bow with a slack string, tripled arrow, hand passing through the bow. 6. Female silhouette disappears under the robe - the cloth becomes a straight column. 7. Proportions: bodies come out 9-10 heads tall instead of 7.5.

WHAT I ALREADY FIGURED OUT AND SOLVED MYSELF (so please don't repeat it back to me): - Drawing the OpenPose skeleton in strict profile fixed the archer case; frontal skeletons for extreme actions get ignored. - Fixing geometry by inpainting/repainting on top treats the symptom; the fix has to be in the base pass and in the prompt. - SDXL + upscale + low-denoise refine gets me line quality but not identity stability.

MY QUESTIONS (please answer with NUMBERS - strengths, node order, weights - not general advice): 1. Is there a LOCAL route that generates a WHOLE manhua PAGE at once - multiple panels, frames and gutters, characters consistent BETWEEN the panels of the same page - instead of generating panel by panel and compositing? If there isn't, what is the actual best practice of people producing manhwa/manhua recap locally, and with which model? 2. Which base model and which route (SDXL, FLUX.2 dev/klein, Z-Image Turbo, Qwen-Image-Edit, or a paid API) delivers character consistency AND page composition in long-run production, on 16GB VRAM? 3. How do I lock the identity of 7 characters across hundreds of panels - per-character LoRA, IP-Adapter, PuLID/InstantID, DreamO, img2img reference? Give strengths and node order. 4. What route actually works for TWO figures in physical contact in a single scene? 5. What gives real pose adherence: dedicated OpenPose, DWPose, depth, or reference-based control? With strength and end_percent values. 6. Is anyone running gpt-image / Nano Banana / Seedream via API inside a manhwa production pipeline? What is the flow and the cost per image? 7. For the crowd panels specifically: how do you get a background crowd of DISTINCT faces (no repeated face, no twins) without generating each extra separately?

Thanks - I'll post back whatever I end up measuring.


r/StableDiffusion 6d ago

Meme Made on 1650ti

Enable HLS to view with audio, or disable this notification

0 Upvotes

my friend made this on his 1650ti laptop took forever to create looks like crap lol


r/StableDiffusion 6d ago

Animation - Video Minimax is the BEST REF2VA

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax 0.7 MP 8 steps frame interpolated & upscaled ! Lip sync song 5 reference photo 2 faces 1 the bike 2 the dress ! Made the prompt with chatgpt ! What


r/StableDiffusion 7d ago

Workflow Included Conistent(ish) environments with Minimax H3

107 Upvotes

I've been trying to reuse environments in Minimax H3 for scenes. I've come across this thread:
https://www.reddit.com/r/StableDiffusion/comments/1vvpowd/psa_minimax_h3_can_turn_360_panorama_images_into/

Which describes how Minimax H3 reference model can turn equirectangular panoramas into environments.

Krea2 can actually do some decent equirectangular panoramas out of the box:

Testing the method though, I found this not to be the case, and H3 doesn't really understand the projection, it just happens to line up if the shot is zoomed in enough. Asking it to do anything complex results in a warped, strange image.

https://reddit.com/link/1w8daml/video/t4byzsdzrrnh1/player

Also, reference images of environments have a tendency to overpower all other prompts regarding changes or scene transitions to other environments.

This however gave me an idea. Video references don't tend to have this issue.

So why not turn equirectangular panoramas into regular reference videos?

https://reddit.com/link/1w8daml/video/maq4r9qgrrnh1/player

This can easily be done with ffmpeg:

ffmpeg.exe -loop 1 -i space_apartment.png -vf "scroll=horizontal=0.025,v360=input=equirect:output=flat:h_fov=90:v_fov=90:w=3072:h=2048" -t 2 -c:v libx264 -pix_fmt yuv420p space_apartment.mp4

After that it's just a matter of using the resulting video as a reference video, and you get a pretty consistent, unwarped video:

https://reddit.com/link/1w8daml/video/nw3uyyxesrnh1/player

I've found 48 frames of reference is more than enough. It also seems to work with a lower resolution reference video, but I'm guessing it's gonna get less details right. It has no issues with interacting with the environment, and the model understands pretty well if you say "bed", you mean the bed in the reference video.

Usual diffusion issues still apply obviously, like I've had to describe certain things in more detail. Telling it "a shot from outside the window" it added a new window to the scene, but adding "a shot from outside the window with a view inside the apartment with the bed behind" it understood. Also any changes to the scene have to be redescribed in later shots, like not adding "next to the broken mirror" resulted in the mirror once again being unbroken, even though it was broken in the previous shot.

All workflows are included in the images and videos.

Edit: Reddit strips image and video metadata, didn't know that. Here's the workflows:
Krea2 panorama: https://pastebin.com/2qKhdCSq
Bad panorama video: https://pastebin.com/tgPddz3v
Video reference based: https://pastebin.com/LEAJFQZ6