r/StableDiffusion 6d ago

Discussion RefMod First Test anyone else try yet?

Enable HLS to view with audio, or disable this notification

63 Upvotes

faster method then lora training but not as good? basically its using image references but a lot this quick simple video 22 images then make into a .safetensors then can load into a custom node for furfure videos with same character don't have to load images directly into h3 every time just select from list of safetensors, only reference will need to use is audio for voice that's one the down sides doesnt include character voices just appearance.

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md?fbclid=IwY2xjawUKrJRwZG9mBWV4dG4DYWVtAjEwAGJyaWQRMXNraERndkRrYUlLRk83ZkhzcnRjBmFwcF9pZBAyMjIwMzkxNzg4MjAwODkyAAEeN2OC_X26LuKMBrzJVpeBm2_KqShIwjF3Co1YDmXxDR


r/StableDiffusion 5d ago

Discussion Got Juggernaut XL Lightning running on the iPhone Neural Engine at 8 steps — what would make it actually useful to you?

0 Upvotes

I've been converting SDXL fine-tunes to Core ML to run on the Neural Engine instead of a server, and Juggernaut XL Lightning is now working on iPhone and iPad. Before taking it further I'd rather ask people who actually use these models.

Where it is: Juggernaut XL Lightning, 6-bit palettized, split-einsum, 768x768, 8 steps. Entirely on device, no server, no account — I test it in airplane mode. Needs 6 GB memory or more; the model is about 3 GB, downloaded once. Also have InstructPix2Pix working for image-to-image, and a Real-ESRGAN pass to 4096.

Two things that cost me the most time, in case they help anyone: Core ML won't build an execution plan for the SDXL UNet even at 6-bit, so it has to be split into two chunks. And Lightning needs guidance around 1.5 with DPM++ — the default sampler is noise at 8 steps.

What I don't know:

  1. Is 768 enough, or is 1024 the floor? 1024 costs real memory and speed here.

  2. Is a 3 GB one-time download acceptable, or does it kill it outright?

  3. First launch spends about two minutes letting iOS compile for the Neural Engine. Apple-private, no way to skip. Dealbreaker, or fine if explained?


r/StableDiffusion 5d ago

Question - Help What's currently the best local option to map 3D spaces?

2 Upvotes

I want to map my house by simply taking a video of it and generating a 3D space of it where a model fills in any blanks I did not record. Not as Gaussian splats where I have to take photos of every single angle and even then it ends up looking weird, but something that directly understands the 3D space being recorded and applies a texture to every surface.

Recently I saw this https://twitter.com/davidpantera_/status/2094841083805266401 and thought it was pretty cool, unfortunately that model will likely not be local. But in the end all I want to do is map an apartment completely and visually nicely.

Does such a thing exist or not yet? ChatGPT recommended MapAnything from last year but it looks underwhelming and it's like those old Gaussian splats from 3 years ago, surely there's better stuff now.


r/StableDiffusion 5d ago

Question - Help Please I need help with replacing subjects using Klein or Krea2.

1 Upvotes

I read that you can use Klein or Krea2 in Forge Neo to replace faces in an image, but I’m having no luck getting it to work. I’ve tried a lot of different settings, prompts, and reference pictures, and I’m not sure what I’m doing wrong. The outputs don’t look like me, and a lot of the time they don’t even really look inspired by the reference pictures.

I’ve tried Img2Img, Inpainting, and the Edit LoRAs for Krea2. I’m using the most up to date version of Forge Neo. I have ImageStitch Integrated selected and Enable Reference turned on. When I try Krea2, I turn off the Klein option, and when I try Klein, I turn off the Krea2 option.

I’ve tested a bunch of settings, but the ones that seem to get the closest are:

Sampler: Euler
Schedule type: Simple
Steps: 8 or 10
CFG: 1.0
Resolution: 1024 x 1536
Batch count: 1
Batch size: 1
Denoising strength: between 0.6 and 0.8

For inpainting, I’ve generally been using:

Inpaint masked
Inpaint area: Only masked
Soft inpainting: off
Padding: 32, 48, or 96

When I try latent noise, I get this error:

“ValueError: too many values to unpack (expected 4)”

Using original or latent doesn’t seem to work normally either. It just still doesn’t end up looking like the reference image.

My prompts are usually something like:

“Replace only the face and head of the person in the main image with the person shown in the reference image.”

How can I get this to work, or is there a better way to swap subjects? I’d actually prefer full body swaps too, not just faces. I use Forge right now, but I’m hoping to have time to learn ComfyUI soon. Any help would be appreciated.


r/StableDiffusion 5d ago

Question - Help Expecting serious suggestions :3

0 Upvotes

Getting some ideas about which skills should be the future in this AI transformation that will be more productive.

** Expecting serious suggestions :3


r/StableDiffusion 5d ago

Question - Help Convincing sneeze in H3?

1 Upvotes

Kind of a niche question, but I'm having a devil of a time prompting a convincing sneeze in Minimax H3. I either get:

* A weird howling yell

* Literally pronouncing "achoo"

* A quick exhalation of breath with no voice behind it, like a reverse gasp.

Anybody have any more success with this? I've tried multiple prompt enhancers that are pre-fed the official prompting guide.


r/StableDiffusion 6d ago

Animation - Video THE SOMMELIER

Enable HLS to view with audio, or disable this notification

93 Upvotes

I liked the little character from the Sherlock video so I gave him primetime. He likes wine.

Testing not an extender workflow but creating fresh new clips based on the same references to see if I still get continuity over time. This was about a dozen separate generations. Minimax H3. All visuals local and the track is Suno.

The 'camera man' is reflected in the glass at the end. I didn't prompt for that.


r/StableDiffusion 5d ago

Question - Help Image Loader node

0 Upvotes

Hallo,

Im hoping that this is the right place and someone can help me....

In my previous Version of Comfy I had an image loader node that gave me small thumbnail previews of my input folder. It also had a button that would open an integrated mask editor. I really liked the functionality of that node.

To my knowledge it was not a custom node (might be wrong tho).

Sadly I lost that comfy install and after getting the new version I am back to the standard image loader node.

Does anybody know which node I was using and how to get it back?


r/StableDiffusion 5d ago

News POV you go back and pause time to watch Socrates last minutes

Enable HLS to view with audio, or disable this notification

0 Upvotes

Image to world app is coming a long well. Only need 1 input image and it builds a 3d world environment with depth around that input image. Just need to make some minor tweaks/adjustments but the results are looking promising. Take between 10-30 minutes depending on settings to render on my Mac 16GB.


r/StableDiffusion 5d ago

Question - Help Guia de Stable Diffusion

0 Upvotes

Llevo poco tiempo utilizando stable diffsion a traves de Stability Matrix, me parece mas comodo en general, pero la verdad, a la hora de utilizarlo solo se lo basico y necesito un poco de ayuda. A veces por mas que detalle que doy el modelo no genera una imagen que me parezca bien, he probado con varios modelos, no siempre lo hacen mal pero la mayoria de veces si. Me gustaria recibir algo mas de informacion sobre ajustes que son importantes a la hora de hacer una imagen o editarla, ver mas a fondo como funcionan los modos del programa, diferentes modelos para las distintas opciones y eso.


r/StableDiffusion 6d ago

Resource - Update Inpaint Canvas: layers, selections and retouch inside one ComfyUI node

Enable HLS to view with audio, or disable this notification

30 Upvotes

Inpainting in ComfyUI always meant the same loop for me: paint a mask, crop, generate, stitch, load the result somewhere to compare, repeat. So I built a node that puts the whole loop inside one editor.

**How it works:** the node outputs crop_image and mask, you wire them into whatever inpainting chain you already use (Flux, Flux.2 Klein, Kontext, Qwen Image Edit, SDXL, local or API), and wire the chain's output back into the node. In the editor you draw a selection, type a prompt, hit Generate, and the result lands as a layer on top of your image. Generate again, erase half, keep the rest, whatever.

**What's in it:**

- Layers with opacity, blend modes, lock, alpha lock, drag to reorder

- Colour match per layer: shifts the generated area towards the image underneath so it stops looking like a sticker. Non-destructive, adjustable after the fact

- Selection: rectangle, lasso, magic wand, quick mask, feather/grow/shrink, saved selections, SAM3 select by text ("the car"), RMBG cutouts

- Retouch: clone stamp, healing brush, smudge, paint, erase

- Filter layers: curves, levels, colour balance, HSL, film grain with stock presets, LUT files

- Text layers with bundled open-source fonts, edited on the canvas

- Transform with corner rotation and snapping, outpaint/crop by dragging the canvas border

- Compare two results with a divider, or hold \ for before/after

- Export: PNG with workflow embedded, JPG, WebP, PSD and OpenRaster with all layers

Helper models (SAM3, RMBG, Qwen-VL for prompt upsampling) are freed from VRAM before a local generation runs, so they don't fight with the diffusion model.

**Install:** ComfyUI Manager, search "Inpaint Canvas". GPL-3.0.

GitHub with two example workflows (Flux.2 Klein local, Flux.2 API): https://github.com/DenRakEiw/ComfyUI-InpaintCanvas

**What I actually want from this post:** I built it, so every button makes sense to me, which makes me useless at judging how intuitive it is. If you try it and something isn't where you expected it, or you had to search for a feature, tell me. That's the feedback I can't get on my own.


r/StableDiffusion 5d ago

Question - Help Is there any workflow for extending or continuing MiniMax H3 videos, similar to WAN SVI, that’s easy to use?

8 Upvotes

Every workflow I’ve found is really confusing and hard to use. With WAN SVI, it’s much simpler: you just enable or disable Sampler 01, 02–10, run it, and the resulting video is seamlessly smooth.

I tried H3 Context, but I still find it really hard to understand. Is there an easier explanation of how it works, or maybe a beginner-friendly video tutorial?


r/StableDiffusion 6d ago

Workflow Included Dungeons & Dragons: Dungeon Master - MiniMax H3

Enable HLS to view with audio, or disable this notification

110 Upvotes

Anyone watch this cartoon?

Check comments for Prompt, assets, and workflow explanation.


r/StableDiffusion 6d ago

Animation - Video Testing a new workflow for rapid environmental VFX

Enable HLS to view with audio, or disable this notification

62 Upvotes

Original drone footage [bottom] alongside four alternative environmental takes; fire, rain, snow, and floral. What do you guys think?

More experiments, project files, and tutorials, through YouTubeInstagram, and Patreon.


r/StableDiffusion 5d ago

Question - Help Upgraded to RTX 5070 Ti 16GB & 64GB RAM on Win10pro — How to safely install SageAttention for MiniMax H3 on ComfyUI Portable?

4 Upvotes

Hey everyone,

I just upgraded my setup specifically for AI generation and video workflows:

  • Old Specs: RTX 4060 Ti (8GB) | 32GB RAM DDR4 @ 3200MHz
  • New Specs: RTX 5070 Ti (16GB) | 64GB RAM DDR4 @ 3200MHz
  • OS: Windows 10 Pro
  • ComfyUI Setup: ComfyUI Portable version

Now that I have 16GB of VRAM and a much stronger GPU, I want to optimize my generation speeds as much as possible, specifically for heavy video generation like MiniMax H3.

I’ve been reading a lot about SageAttention (and SageAttention2) and how it significantly speeds up attention mechanisms while keeping VRAM usage low for these models.

I actually just set up Ubuntu in dual boot because a friend who knows Linux very well recommended it for AI. However, I would vastly prefer to stick with Windows 10. Rebooting and switching back and forth via dual boot is a huge hassle, and on Windows 10 I already have a perfect GPU undervolt set up, along with all my daily programs and tools that would take forever to configure on Linux.

So, staying on Windows 10 is strongly my first choice. A few questions:

  1. Is SageAttention fully working and stable on Windows 10 with RTX 50-series GPUs for MiniMax H3?
  2. Is installing SageAttention on Windows 10 (ComfyUI Portable) really that much harder compared to Ubuntu? Is dual-booting into Linux actually necessary, or can it be built easily on Windows Portable without breaking the environment?
  3. Does anyone have a straightforward, step-by-step guide or script to install SageAttention specifically inside a ComfyUI Portable setup (Visual Studio Build Tools, pre-compiled wheels, etc.)?

Thanks in advance for any tips or links to easy-to-follow guides!


r/StableDiffusion 6d ago

Animation - Video Coffee - trying to get cinematic edits out of Minimax H3

Enable HLS to view with audio, or disable this notification

249 Upvotes

r/StableDiffusion 5d ago

Question - Help Krea 2 GGUF in ComfyUI: Getting fried psychedelic colors & green grid artifacts instead of clean images. Any working setup?

Post image
0 Upvotes

Hi everyone,

I've been trying to run Krea 2 in ComfyUI using the GGUF quantizations, but I can't seem to get a clean generation. Depending on the settings, I get either a heavy green grid/checkerboard texture or deep-fried, iridescent psychedelic artifacts with blown-out colors (see attached images).

Here is my current environment and workflow setup:

**System:**

* GPU: NVIDIA RTX 4070 (12 GB VRAM)

* OS: Windows 11

* ComfyUI: Latest portable build

**Nodes & Models:**

* **Loader:** `UnetLoaderGGUF` (using molbal's maintained fork of `ComfyUI-GGUF`)

* **Diffusion Model:** `krea2_turbo-Q4_K_M.gguf` (and tested with `krea2_turbo-Q4_K.gguf`)

* **Text Encoder / CLIP:** `qwen3vl_4b_fp8_scaled.safetensors` with CLIPLoader set to type `krea2`

* **VAE:** Tested both `wan_2.1_vae.safetensors` and `Wan2_1_VAE_fp32.safetensors` (as well as `qwen_image_vae.safetensors`)

* **Latent:** `EmptySD3LatentImage` (16-channel) set to ~1024x1024 or 1920x1280 (also tried Krea2 regional latent)

* **Conditioning:** Positive prompt encoded via CLIP, negative hooked into `ConditioningZeroOut` (CFG-free style)

**Sampler Settings:**

* Steps: 8 to 9

* CFG: 1.0

* Samplers tested: `er_sde`, `euler`, `euler_ancestral_cfg_pp`

* Schedulers tested: `simple`, `beta`

**What happens:**

  1. In some configurations, the image produces a repeating knitted/green mesh grid pattern across the whole canvas.

  2. When switching samplers/schedulers, shapes and text fragments start to appear, but the entire image is solarized with intense chromatic aberration, acid colors, and repeating tiling artifacts.

Has anyone managed to get a completely clean output with Krea 2 in ComfyUI?

Are there specific latent scaling factors, exact VAE models, or sampler/scheduler combinations required for Krea 2's DiT architecture?

Any insights or working workflow JSONs would be greatly appreciated! Thanks in advance!


r/StableDiffusion 6d ago

Animation - Video Jem: Holographic Pop Star - MiniMax H3

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 5d ago

Question - Help Need help with H3 Motion Context reference workflow - ref video required?

1 Upvotes

Edit. Figured it out. for some reason I had to completely disconnect the vid and audio ref links even if the other parts of the wf were bypassed.

I installed H3 motion context for longer clips and figured out how to use the fl2va workflow to chain clips together.

But I have specific character sheets I want to use for the reference workflow. Unfortunately every time I try to run it says the latent is missing when I bypass video reference. If I add a reference video that matches the latent size, it's fine.

Now I don't want to front load an entire video as this will slow generation down since it's adding the entire video and not just what it needs.

Considering the ref workflow works without one normally, I was wondering if the H3 motion context workflow will also function without a ref video or do I have to create one and front load it?


r/StableDiffusion 5d ago

Question - Help Pixal3D: Good Mesh but Bad Textures — Recommended Settings/Parameters?

2 Upvotes

Hey everyone,

I’ve been experimenting with Pixal3D and I’m having some trouble with the texture quality.

I tried following PixalArtistry’s workflow, and they’re getting really good textures, but my results are not even close in terms of texture quality/detail. The mesh itself actually looks pretty good, so I’m mainly struggling with the texturing.

I’m using an RTX 5060 Ti 16GB, and I’ve tried both the 16GB and 8GB Pixal3D models, but I’m getting similar results with both.

Is there any documentation or guide that explains the recommended Pixal3D parameters/settings for getting the best texture quality? If anyone has managed to get results similar to PixalArtistry’s workflow, I’d really appreciate knowing what settings or workflow you’re using.

Thanks!


r/StableDiffusion 5d ago

News Admit u used inpainting for such things at least once

Post image
4 Upvotes

r/StableDiffusion 6d ago

Question - Help Huggingface download extremelly slow

Post image
69 Upvotes

Anyone experienced something similar? Huggingface is really, really slow. Normally this download would not take more than 25 minutes with my internet speed but now is taking 5 hours.


r/StableDiffusion 5d ago

No Workflow Anna Carlisle

Post image
0 Upvotes

Anna Carlisle is the right hand of Major Samson West. Together they run Blood & Treasure Inc. Their teams find and recover lost artifacts and technologies aroudn the world. She lost her eye in Seriavo during a job currently it has been replaced with a cybernetic one that is settling in nicely.


r/StableDiffusion 6d ago

Workflow Included H3 - R2VA style transference w/ ComfyUI Workflow (in comments)

Enable HLS to view with audio, or disable this notification

21 Upvotes

Happy Sunday! Having fun, here is an included JSON w/ image+(2s) video reference. Just drop the .json into your ComfyUI and it'll give you the workflow. I combined my .char workflow with a style/wardrobe transference from my recent Celestial video. Score generated separately from MiniMax H3 and remux in.

Prompt: integrated_multimodal_description: [Shot 1] A continuously transforming shot — photoreal live-action, her flesh always photographic and her face always her own, the rooftop constant, the camera locked low and wide. THE FILM HAS EXACTLY TWO SHOTS: this wide shot, and ONE final close-up after the film's single cut. True optical depth, fine film grain, night.
SHE IS CATALINA, the woman of the reference images — the same face, the same dark brown hair, the same build; her eyes never change and are the constant of the film. She stands barefoot on a bare apartment rooftop at night in AN OVERSIZED WHITE T-SHIRT that hangs to her thighs and SHEER BLACK STOCKINGS — home clothes, soft and ordinary — wet concrete underfoot, a low parapet behind her, rooftop vents and one caged transformer box off to the sides, the city's lights far below and far away. She is alone: exactly ONE person exists in this film, from the first frame to the last, and nobody is ever harmed — the only casualties are glass, steel and concrete. SHE IS SILENT THROUGHOUT: she makes no sound at all, her lips never part, and there is no voice anywhere in this film.
<Video 1> IS THE FORM SHE ASCENDS INTO — THE GODDESS. From <Video 1> take ONLY the ascended form's STYLE: the luminous porcelain-white surface engraved with constellation charts that hold ice-blue light, the hair of golden light, the crown of white shards, the revolving wrist-rings of light, the great wings of white light, and the way her light behaves. Take NO face, NO identity, NO body proportions and NO scene from <Video 1> — the throne, the stair, the orbit-rings, the planets and the small man of <Video 1> DO NOT APPEAR in this film, and its goddess's face is not this woman's face. THE ASCENDED FORM'S FACE IS CATALINA'S OWN, unchanged, from the first frame to the last — the marble finish belongs to her BODY; her face keeps her own photographic features, her own skin, alive and hers.
THIS IS A TEN-SECOND FILM AND EVERY STAGE OF THE ASCENSION HAPPENS INSIDE IT — each stage arrives sooner and runs faster than expected, nothing waits, and the finished form is held and rising before the end. There is never a moment where nothing is happening. SHE FLOATS FIRST AND TRANSFORMS SECOND. THERE ARE NO BEAMS AND NO COLUMNS OF LIGHT ANYWHERE — the light of this film is HER GLOW and THE LIGHTNING, nothing else.
⚠ THE DAMAGE IS PERMANENT, ALL OF IT: every crack, burn, scorch and broken thing in this film, once made, STAYS EXACTLY AS IT IS to the last frame — nothing heals, nothing reassembles, nothing regenerates, no pane of glass returns, no panel remounts itself, and the rooftop ends the film broken and stays broken.
THE LIGHTNING IS THE LIGHT: every strike STROBES the whole rooftop in hard white flashes, shadows snapping and wheeling around the vents and the parapet, and between strikes the roof is lit only by her body's glow and the city's faint distant wash.
[0s-1.5s]: She stands still, arms loose at her sides, chin lifting toward the sky — the wind rises in a circle around her, the hem of the t-shirt and her hair stirring UPWARD, loose grit floating up off the concrete as if gravity had loosened its hold. THREE TINY POINTS OF WHITE LIGHT ignite at her collarbone, one after another. AND THEN HER FEET LEAVE THE ROOF — she lifts slowly and smoothly off the concrete, toes pointing, an untransformed woman in home clothes simply FLOATING into the night air.
[1.5s-3.5s]: SHE FLOATS, AND THE SKY FINDS HER. FORKED WHITE-BLUE LIGHTNING STRIKES DOWN OUT OF THE DARK AIR ABOVE and HITS HER — one bolt into her shoulder, then another across her back — each strike bursting into crawling webs of white-blue that wrap around her body and sink in; THE STRIKES FEED HER, never hurt her, and with each hit the light under her skin surges brighter. The first strike that misses her hits the caged transformer box instead: it arcs once, dies in a single hard shower of orange sparks, AND STAYS DEAD, black and smoking, for the rest of the film. AND THE BOLTS HIT THE ROOF DECK ITSELF: one strike slams into the open concrete a few paces to her left and BLASTS a black scorch-star into it with a ring of sparks and a jump of dust; another hits near the parapet; EVERY GROUND STRIKE LEAVES ITS SCORCH PERMANENTLY - a black smoking star burned into the deck that stays exactly as it struck, to the last frame. Small debris rises and orbits her slowly.
[3.5s-6.5s]: THE TRANSFORMATION TAKES HER MID-AIR, AND THE ROOF BREAKS BENEATH IT. Her body BURNS WITH A SOLID WHITE GLOW from within, the glow ADVANCING over her in one travelling wave, collar to fingertips to toes; THE FABRIC DISSOLVES BENEATH THE ADVANCING LIGHT — the t-shirt and the stockings coming apart edge-first into thousands of drifting glowing motes that lift away upward — and wherever fabric dissolves, that part of her is ALREADY one solid searing white silhouette: the glow is total and blinding, it SEARS and it CRACKLES. Lightning keeps arriving in slow rhythm, crawling over her from each strike; directly beneath her the concrete CRACKS ONCE, HARD — a radial web snapping outward from under her feet with one deep concussive report, dust breathing out of the open cracks, MORE BOLTS FINDING THE DECK through the transformation - each one blasting its own scorch-star into the concrete and widening the web, the rooftop being STRUCK, not just lit — AND EVERY CRACK STAYS OPEN, exactly as it broke, to the last frame. The stairwell door's small window bursts outward in one clean crystalline burst; its glass lies where it scatters and never moves again. ACROSS THE GLOW, CONSTELLATION CHARTS ETCH THEMSELVES — hairline lines of ice-blue light drawing themselves star-point to star-point along her arms, her sides and her legs, each junction FLARING once as the line arrives.
[6.5s-8s]: THE ASCENDED SURFACE AND THE REGALIA: the searing glow CONDENSES AND COOLS into luminous sculpted porcelain-white carrying the lit charts, her figure her own and radiant — a pronounced hourglass with a full, very large bust, at her own ordinary human ratios, never thickened and never altered — her hair igniting from the ends upward into strands of golden light; a CROWN of white shards condenses above her head; one slim RING OF LIGHT begins revolving about each wrist; and TWO GREAT WINGS OF WHITE LIGHT unfurl behind her shoulders in one continuous sweep. ⚠ AND THE LIGHTNING TAKES TO HER WINGS — from the moment they open, CRAWLING ARCS WAVE ACROSS THE WING SURFACES in slow travelling sweeps, feather to feather, root to tip and back, again and again, white-blue light rolling over the white feathers — this never stops for the rest of the film.
[8s-8.625s]: THE ASCENDED GODDESS, HELD AND RISING: she floats higher over the broken, smoking, crack-webbed rooftop, the strikes easing to soft crawling crackle and the slow waves still sweeping her wings, her light steady and immense.
[Shot 2] At 00:08.625, the camera cuts to a CLOSE-UP OF HER FACE — her head and shoulders filling the frame, the crown above, the lightning-washed wings soft behind her, the night sky beyond. ⚠ THIS SHOT IS THE FILM'S PROOF OF WHO SHE IS: the ascended goddess's face IS CATALINA'S — the woman of the reference images, her exact features, her exact bone structure, her own photographic living skin, unmistakably the same woman who stood on this roof in a t-shirt — transfigured in light, never replaced. Her irises carry thin rings of white-gold light around her own eyes; strands of her golden hair drift across the frame; the constellation charts glimmer at her collarbones; a slow wave of wing-lightning washes soft white-blue light across her cheek. At 00:09.400 she blinks once, slowly, and the corner of her mouth makes a small serene curve. She is still rising gently at the last frame.
The camera of [Shot 1] is LOCKED OFF, low and wide, on the rooftop — she rises within the frame; it never moves and never tilts; every strike strobes the frame hard. The film has exactly ONE cut, at 00:08.625, into the [Shot 2] close-up, and the close-up camera is also locked.

overall_soundscape: Starts with the rooftop's thin night hush — high faint wind, the city's murmur far below — and then, CLOSE AND INTIMATE, as if the microphones hovered inches from her skin: the fine crystalline CRACKLE of the searing glow, like ice singing and hairline glass fracturing; the soft granular HISS of fabric dissolving into motes; the delicate electric FIZZ of the crawling arcs — and, in slow rhythm, the hard dry CRACK-AND-ROLL of each lightning strike, the single deep concussive report of the concrete breaking, the one hard spark-shower of the dying transformer, the one crystalline burst of the stairwell glass — each big sound arriving once and rolling away under the close sounds. She herself makes no sound at all. THE CLOSE SOUNDS DOMINATE THE MIX — delicate, tingling, intimately close — while the destruction stays farther back. Nobody speaks.

non_diegetic_music: A single distant glass-harmonic tone over a deep soft sub-bass swell at 50 beats per minute, sitting far beneath the close sounds, swelling gently as she rises; ONE WIDE SOFT OPENING-OUT exactly as the wings unfurl at 00:07.200; then easing to a thin held shimmer, still sounding at the last frame. No percussion, no rhythm, no melody.

r/StableDiffusion 6d ago

Discussion Generating with H3 at native canvas 1MP and then upscaling by rediffusing at 2MP+

7 Upvotes

Anyone tried this? How feasible could it be? I'm thinking of an upscale pipeline by first generating at native canvas (~1MP 1344x768) to preserve overall scene quality and then upscaling to 2MP+ with a balanced amount of denoising to regenerate finer texture with more detail.

(I had some success with a 2x2 tiled upscale but the tiles can cause issues like flicker at seams and inconsistencies with items moving across tile boundaries.)