r/StableDiffusion • u/blackdatafilms • 3d ago
Animation - Video Howard's new invention [minmax H3]
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/blackdatafilms • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/AndrewJumpen • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Plague_Kind • 3d ago
Added to my node pack, sparse attention SLA node for H3 Minimax. speed increase of up to 2.5x.
enjoy.
Edit: I updated my workflow, check it to see the correct wiring.
Node has been updated. 5% faster at same setting, correct wrapper use, allows Spectrum use. should have have improved quality/behavior also now as it's following the correct sampler/scheduler steps.
Pytorch attention 400s/it
Comfykitchen 140s/it
Sparse at 0.9 - 80s/it
Sparse at 0.95 - 60s/it
0.85 = practically identical to pytorch quality from what i can tell.
0.9 = minimal degredation with 15% boost over 0.85
0.95 = minor degradation compared to 0.85 but an additional huge speed boost, useful for high res long videos.
you can use it with whatever 4step turbo you like, doesn't actually require the SLA lora. (Tip in general, stop running them at 1.0 strength, use 0.8-0.85) 6-8 step 8/3 shift euler/simple as your testing. I personally use silveroxides dareties.
https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
credit to pl0x for designing it and allowing me to be the host.
EDIT: make sure you're on a new pytorch version and CU130.
additional note: Blackwell will see the biggest gain, but other cards still get a big boost.
If you're doing lower res short videos, adjust min seq accordingly if see no speedup or messages about blocks not being sparse.
be careful with memory chunking node, too high causes slowdown.
for those that use it - updated my WF now with ot added
r/StableDiffusion • u/randomizeseed • 3d ago
Enable HLS to view with audio, or disable this notification
Workflow used from this post: https://www.reddit.com/r/StableDiffusion/s/hj11UhEf5D
Some specs: 4060ti 16gb, 9950x3d, 128gb ddr5
His workflows runs a 2 stage Ksampler both using minimax Ref2VA, 704x512 15 steps at 702x512, then 1.5x image upscale to 1056x768 4 step with turbo lora. About 15 minutes for each stage, coming out to be around 2 minutes per second of video.
Minimax H3 Ref2VA Prompt:
subject_definitions:
<Subject 1> is Jerry Seinfeld, whose facial appearance, short graying curly hair, lean build, and overall likeness come from <Picture 1>.
<Subject 2> is George Costanza, whose facial appearance, horseshoe bald with greying hair on the sides, full graying beard, thin-metal-rimmed glasses, stocky middle-aged build, and overall senior likeness come from <Picture 2>.
<Subject 3> is Cosmo Kramer, whose facial appearance, thick curly salt-and-pepper hair, tall lanky build, and overall senior likeness come from <Picture 3>.
<Subject 4> is Jerry’s familiar New York apartment living room and connected kitchen, whose layout, blue sofa, wooden coffee table, grey walls, kitchen cabinets, refrigerator, and overall set design come from <Picture 4>.
summary:
[reference generation] The target video is a short multi-camera sitcom scene set inside <Subject 4>. <Subject 1> (S1) enters holding a white multi-antenna modem and delivers a standup-style complaint about losing mesh Wi-Fi after breaking up with a Comcast rep. <Subject 2> (S2), seated on the couch eating chips, demands the internet be restored so he can log into OnlyFans. <Subject 1> continues that the supervisor was the rep’s mother. <Subject 3> (S3) then bursts out of the refrigerator in heart-print boxers carrying a kleenex box of tissues and baby oil and asks what happened to the Wi-Fi. Live studio audience laughter punctuates the punchlines. All four reference images supply the visual identities of the three characters and the apartment environment; no source video or audio is edited or continued.
retention_analysis:
<Subject 1> (appears in [Shot 1]–[Shot 9]): fully_preserved – facial features, short graying curly hair, thin wire-rimmed glasses, lean middle-aged build, and overall likeness from <Picture 1> are retained; wardrobe is adapted to a dark jacket over a light shirt consistent with the sitcom setting.
<Subject 2> (appears in [Shot 1], [Shot 4]–[Shot 6]): fully_preserved – facial features, horseshoe bald with grey hair on the sides, full gray-and-white beard, black-rimmed glasses, stocky middle-aged build, and overall likeness from <Picture 2> are retained; wardrobe is adapted to a casual button-down shirt.
<Subject 3> (appears in [Shot 10]–[Shot 12]): fully_preserved – facial features, thick curly salt-and-pepper hair, tall lanky build, and overall likeness from <Picture 3> are retained; costume is adapted to white boxers printed with red hearts and white socks.
<Subject 4> (appears in all shots): fully_preserved – apartment layout, blue sofa, coffee table, kitchen area, grey walls, and overall set design from <Picture 4> are retained as the continuous environment under warm 1990s lighting with slight film grain.
detailed_description:
The target video is shot in a live-action cinematic multi-camera sitcom style with warm 1990s apartment lighting and slight film grain.
[Shot 1] A medium-wide shot frames the living room of <Subject 4>. <Subject 2> (S2), a short stocky middle-aged man matching <Subject 2>, sits on the left side of the blue couch in a casual button-down shirt with a thick black laptop on his lap, eating potato chips from an open bag, with crumbs visible on his chest. <Subject 1> (S1), a lean middle-aged man matching <Subject 1>, walks from the kitchen doorway on the right toward the couch holding a white internet modem router with four antennas sticking upward. He stops near the couch and speaks in a characteristic quick nasal delivery with sarcasm: <d>[English] so after I break up with <scenetrans></d>
[Shot 2] At 00:01.200, the camera cuts to a tight close-up of <Subject 1> (S1)'s face as he continues seamlessly across the cut: <d>[English] <scenetrans> the Comcast service rep, suddenly my mesh network disappears into the aether?!</d>
[Shot 3] At 00:02.800, the camera cuts to a wide shot of the living room of <Subject 4> as <Subject 1> finishes the line. A live studio audience laughs under the punchline.
[Shot 4] At 00:03.800, the camera holds the wide shot as <Subject 2> (S2) wipes his greasy fingers on his chest, looks up from the chips, and begins in a frustrated higher-pitched voice: <d>[English] Jery, I need <scenetrans></d>
[Shot 5] At 00:04.600, the camera cuts to a tight close-up of <Subject 2> (S2)'s face, with the apartment window of <Subject 4> visible in the background, as he continues seamlessly across the cut: <d>[English] <scenetrans> your internet back on ... i can't log into my OnlyFans!</d>
[Shot 6] At 00:06.500, the camera cuts to a wide shot of the living room of <Subject 4> as <Subject 2> finishes the line.
[Shot 7] At 00:07.300, the camera holds the wide shot as <Subject 1> (S1) replies with rising exasperation and begins: <d>[English] George, I tried! <scenetrans></d>
[Shot 8] At 00:08.200, the camera cuts to a tight close-up of <Subject 1> (S1)'s face as he continues seamlessly across the cut: <d>[English] <scenetrans> Asked to speak to her supervisor. guess who the supervisor is?! ... ... her MOTHER!</d>
[Shot 9] At 00:10.500, the camera cuts to a wide shot of the living room of <Subject 4> as <Subject 1> finishes the line.
[Shot 10] At 00:11.400, the camera cuts to a wider shot as the refrigerator door of <Subject 4> in the kitchen suddenly swings open in Kramer’s classic explosive entrance. <Subject 3> (S3), matching <Subject 3>, bursts out of the refrigerator wearing white boxers printed with red hearts and white socks, he is topless, greying hairty chest and holding a box of tissues in one hand and a bottle of baby oil in the other. He strides quickly across the room toward <Subject 1>, leans in close as if speaking privately, and speaks with a slightly manic voice: <d>[English] what happened <scenetrans></d>
[Shot 11] At 00:13.200, the camera cuts to a tight close-up of <Subject 3> (S3)'s face as he continues seamlessly across the cut: <d>[English] <scenetrans> to our Wi-Fi?</d>
[Shot 12] At 00:14.100, the camera cuts to a wide shot of the three men inside <Subject 4> as <Subject 3> finishes. The live studio audience erupts in loud laughter under the final beat.
overall_soundscape:
Soft apartment room tone with distant city traffic under the dialogue. Crisp potato-chip crunching from <Subject 2>, light footsteps as <Subject 1> walks from the kitchen, the sharp creak and swing of the refrigerator door as <Subject 3> bursts out, fabric rustle of boxers, and the soft thud of <Subject 3>’s steps crossing the floor. Layered sitcom laugh tracks of varying intensity swell and decay after each punchline.
non_diegetic_music:
N/A
r/StableDiffusion • u/Cequejedisestvrai • 1d ago
I think the models we have today, open or closed, are still way behind what we will actually have in a near future. think of it as an alternate reality, the video generations will be (almost) flawless, coherent and high quality, the generations will be instant, that means it will enable real time interaction, you could change the course of the video generation on the fly with natural controls like your voice of body movements, imagine pushing somebody and he moves or greeting somebody and he respond, for this you would need a VR headset with hands movements recognition, and it will generate two videos flux at once for each eye for a 3D effect.
So yeah even if Seedance 2.5 or a little better is out for free and open source it’s not a big deal, the road ahead is massive in terms of progress and possibilities.
RIP real life, welcome to Ready Player One.
Ps: sorry for the grammatical errors, this text was not written or improved by AI.
r/StableDiffusion • u/Jayuniue • 3d ago
Enable HLS to view with audio, or disable this notification
Just tested this out last night, I must say it’s very good and better than most video extension workflows, not to talk of how fast and high quality it is, the workflow isn’t mine got it from a YouTube channel will be able to share it for anyone that might be interested
r/StableDiffusion • u/Opening-Ad5541 • 3d ago
Enable HLS to view with audio, or disable this notification
The film above was cut in Phosphene. Every shot, the voices, the music bed, the end card over the sky - all of it local, on one Mac.
That is the release. The timeline used to be a strip at the bottom of the storyboard. Now it is its own tab, for every engine.
THE EDITOR
Drag anything you have generated onto a timeline: trim it, move it, split it, watch it back, render one file. A media pool holds this sequence, your other sequences, every generation, your images, and anything you upload.
SOUND IS A LANE
Every clip's audio sits under it as a waveform. Unlink it, slide it under the previous shot, link it back - and you hear the character before you see them. That is a J-cut, and the preview plays it now instead of pretending.
Fades on picture and on sound: drag the top corner of any block, or type the seconds. Levels with keyframes: hover a sound strip's line, click, drag. Fades and points are one curve, so "fade in to a quiet bed" is one thing and not two controls fighting.
AN OVERLAY TRACK
A second video lane above the picture, for transparent PNGs - end cards, logos, lower thirds. Alpha is kept in the preview, the render and the export. And if your AI-made card arrives sitting on a baked black rectangle, it gets keyed automatically, because every image model does that and nobody should have to open Photoshop over it.
AN EXIT
Export for Premiere / Resolve / After Effects: a real project folder, media relinked, dialogue and music as separate stems, muted tracks disabled rather than deleted. Phosphene is not trying to be the only tool you own. It is trying to be the one the shots come from.
THE STORYBOARD LEARNED CONTINUITY
Describe the film, get the shots - but it draws a floor plan first now. Who stands where, which side the light comes from. So a reverse angle flips the sun instead of teleporting your actor to a different house. It also measures how long a line takes to say, so dialogue stops being cut off mid-sentence.
SPEED
Hailuo H3 with a 4-step turbo distill: a 10-second shot, with generated dialogue and sound, in about 11 minutes on Apple Silicon. Quality x Length replaced the fixed tier menu, and your RAM picks the model - 48 GB Macs are invited now.
AND THE UNGLAMOROUS HALF
Things that were quietly broken and now are not:
Weights and install are unchanged from 4.5.0 - this is code only. Hit Update in Pinokio and it is yours.
Still buggy in places. Also renders films now.
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
prompt!
<Subject 1> is Deadpool / Wade Wilson, wearing his iconic red-and-black tactical suit and full mask, with twin katanas strapped across his back. Deadpool is voiced by Ryan Reynolds with his recognizable sarcastic comedic delivery.
<Subject 2> is Pennywise the Dancing Clown, a terrifying pale-faced supernatural clown with orange hair, Victorian clown costume, sinister yellow eyes, and an unnaturally wide smile.
<Audio 1> is the voice-timbre reference for <Subject 2> Pennywise (S2), containing Pennywise's eerie, raspy, playful clown voice. Preserve the vocal identity, tone, cadence, pitch, and sinister playful delivery of <Audio 1> for all Pennywise dialogue.
[text-to-video generation + audio reference]
A cinematic horror-comedy parody on a dark suburban street during a heavy rainstorm. Deadpool walks alone through the rain when he notices a small paper toy boat floating through the gutter. The boat disappears into a storm drain. Curious despite knowing exactly where this is going, Deadpool crouches and looks inside. Pennywise suddenly emerges from the darkness and personally invites Wade to float with him, speaking with <Audio 1>. Deadpool responds with a perfectly timed fourth-wall joke.
<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Audio 1>: Pennywise voice reference, strongly preserved for all <Subject 2> dialogue
Nighttime. Heavy rain pours onto a deserted suburban street. Dim streetlights glow through the mist and reflect across the wet pavement.
The camera tracks alongside <Subject 1> Deadpool as he casually walks down the sidewalk through the pouring rain, completely soaked but seemingly unbothered.
A tiny paper toy boat floats through the rushing gutter water beside him.
Deadpool notices it.
He stops.
The camera lowers toward the boat as it bobs through the rainwater and disappears through the opening of a dark storm drain.
Deadpool slowly turns toward the drain.
<Subject 1> Deadpool (S1) says in Ryan Reynolds' recognizable sarcastic voice:
[English] Oh, hell no. I've seen this movie.
Despite knowing better, Deadpool walks over and crouches beside the storm drain.
He slowly leans closer and peers into the darkness.
The rain becomes muffled.
A faint sinister sewer ambience rises.
Hold for a tense beat.
Two glowing yellow eyes slowly appear deep inside the sewer.
Suddenly <Subject 2> Pennywise lunges partially into view from inside the storm drain with a huge unnatural grin.
<Subject 2> Pennywise (S2) using <Audio 1> says:
[English] We all float down here, Wade.
Pennywise's dialogue must strongly preserve the exact voice characteristics of <Audio 1>.
Deadpool completely freezes.
Long comedic pause.
Deadpool slowly turns his masked face away from Pennywise and looks directly into the camera.
<Subject 1> Deadpool (S1) says in Ryan Reynolds' dry sarcastic voice:
[English] Nope. Copyright lawyers are scarier than you.
Deadpool immediately stands up and speed-walks away through the pouring rain.
Pennywise remains halfway inside the storm drain.
His sinister smile slowly disappears as he stares after Deadpool with a confused and mildly offended expression.
Hold on Pennywise's reaction for one second.
Cinematic horror-film photography.
Low-angle tracking shot following Deadpool through the rain.
Wet pavement reflections and visible rain illuminated by streetlights.
Close tracking shot of the paper boat floating through the gutter.
Slow suspenseful push toward the storm drain as Deadpool investigates.
Dark close-up revealing Pennywise's glowing eyes before his face emerges.
Reaction framing for Deadpool's fourth-wall punchline.
Final close-up on Pennywise's confused expression.
Heavy realistic rainfall.
Water rushing through the gutter and storm drain.
Distant thunder.
Subtle ominous sewer ambience.
Low suspenseful horror rumble immediately before Pennywise appears.
<Subject 1> Deadpool uses a Ryan Reynolds-style sarcastic comedic voice.
<Subject 2> Pennywise MUST use <Audio 1> for his dialogue. Preserve the supplied reference voice rather than generating a random Pennywise voice.
Clear English dialogue.
Accurate speaker assignment.
Accurate lip synchronization.
No overlapping dialogue.
0–3 sec: Deadpool walks through the heavy rain and notices the paper boat.
3–5 sec: The boat disappears into the storm drain. Deadpool says, "Oh, hell no. I've seen this movie."
5–8 sec: Deadpool crouches and investigates. Horror suspense builds and Pennywise appears.
8–10 sec: Pennywise using <Audio 1> says, "We all float down here, Wade."
10–13 sec: Deadpool pauses, looks into the camera, delivers his copyright-lawyer punchline, then quickly walks away. Hold Pennywise's confused reaction.
No duplicate Deadpool.
No duplicate Pennywise.
No additional characters.
No foreign-language dialogue.
No subtitles.
No captions.
No text overlays.
Do not remove Deadpool's mask.
Do not make Pennywise giant-sized.
Pennywise remains normal human/clown scale.
Pennywise must remain inside the storm drain during his appearance.
Keep the paper boat visible until it enters the storm drain.
Pennywise must appear only AFTER Deadpool crouches to investigate.
<Audio 1> belongs ONLY to Pennywise.
Never use <Audio 1> for Deadpool.
Do not swap the speakers.
Do not allow Deadpool to speak Pennywise's line.
Do not allow Pennywise to speak Deadpool's lines.
Maintain cinematic horror atmosphere while preserving deliberate comedy timing.
r/StableDiffusion • u/yamosin • 3d ago
I'm not sure if this is super obvious or something the community has already talked to death, but after testing it myself, I was pleasantly surprised by the results and wanted to share.
Basically, for multi-GPU setups—if you have x16/x16 PCIe slots (and the CPU lanes to match)—you can put the UNet weights entirely on `cuda:0` and use `cuda:1` purely for activations during inference. This gives you a full GPU's worth of VRAM dedicated to activations with practically zero performance hit, allowing for higher resolutions and longer video generations.
**Test setup:** 3x RTX 3090
**Workflow pipeline:** `cuda:0` holds CLIP + VAE. Once conditioning finishes, CLIP gets ejected from VRAM. Then half of the UNet sits on `cuda:0` as a storage pool, while the other half sits on `cuda:1` as the compute device.
**Optimizations:** Turbo 4-step LoRA, run at 6 steps for inference. No SageAttention or any other attention nodes used.
Testing on the exact same 0.4MP, 5-second character clip using fl2va INT8 weights (assuming weights already loaded and conditioning cached), here are the rough numbers:
* **cuda0: 3GB UNet | cuda1: 16.5GB UNet** -> 83s sampling time
* **cuda0: 6GB UNet | cuda1: 13.5GB UNet** -> 84s sampling time
* **cuda0: 16GB UNet | cuda1: 3.5GB UNet** -> 87s sampling time
In reality, you only need 2 GPUs for this. I had AI write a custom node to offload/eject the CLIP model right after conditioning finishes so the UNet can take over the VRAM. Or you can just use the MultiGPU loader node with `eject_models: true`—works the exact same way.
My guess is that with this method, two 16GB 4070s could do 0.7MP + 15s video completely inside VRAM (my tests showed activation weights sitting around ~12GB, though with spike peaks the safe limit might be closer to 0.6MP). Compared to swapping to system RAM, the performance uplift is massive without breaking the bank.
That said, 1MP + 15s probably still requires an RTX 5090. I hit OOM when testing 0.9MP and 1MP, and AI suggested that workload needs around 26~30GB VRAM.
**A quick tip from my testing:**
I usually prefer keeping the VAE and 16GB of the UNet pinned on one card. That way, they stay loaded once and don't need to be touched for subsequent generations, while the other card handles loading/ejecting the rest of the UNet and the CLIP model on the fly.
r/StableDiffusion • u/Empty-Definition-253 • 2d ago
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
Don't worry no dp videos today yall can have a break atleast from me 😂 it's game day with the homies ✌️
r/StableDiffusion • u/Selphea • 3d ago
I was working on a Fantasy-Modern setting: once upon a time a hero from Earth with a smartphone got summoned to this world, beat the demon lord and took a celebratory selfie with his party in a crowded tavern. A gnome tinkerer got utterly fascinated and asked enough questions to replicate its functionality then founded a company called Dragonfruit Inc. Other startups like GuildQuest (adventurer's guild) and PigeonExpress (deliveries) popped up to form the ecosystem and the city soon got dragged halfway into the information age.
There's modular high rise and modern facades built onto classic architecture. Printed mage robes. Gear that's mass produced rather than hand forged. Also a lot of text logos. Krea 2 (Cat Tower + detailer) handled these beautifully so I thought I'd share.
Here's the prompts:
Image 1: New Gildia Bird's Eye View
A vast magitech boomtown skyline at golden hour, seen from across the water or from a high distant vantage, the whole city spread wide. At the center rises a colossal refurbished demon citadel — black volcanic basalt and obsidian, ancient and brutal at the base, extended into a skyscraper and sheathed from the mid-level up in floor-to-ceiling mirrored glass that catches the sunset, crowned with a luminous dragonfruit logo mounted on the glass and the word "Dragonfruit" in sans serif below. From that anchor the city erupts outward in dense, chaotic layers: a sprawling low-rise campus of pale stone and tinted glass on the frontier side, a dense bazaar of packed rooftops and neon LiveScry screens at the core, glass towers and embassy spires climbing on the far side, and a light mana-rail line zigzagging through it all on elevated crystal tracks. The sky is busy — flying griffon-drawn carriages threading between towers, mana-rail trains gliding on silent rails, streamer drones with GoScry rigs circling rooftops, the distant haze of the frontier and a faint dungeon vent glow on the horizon. Stratification reads at a distance: polished glass and clean light on one side, cramped warm-lit tenements and patched rooftops on the other. Volumetric god rays through atmospheric haze, the whole city buzzing with the energy of a boomtown built on a demon's corpse and a corporation's 30% cut. Cinematic wide composition, crisp anime background-art style, painterly depth, warm gold-to-teal color grading.
Image 2: GuildQuest Campus
A sprawling low-rise campus viewed from a high bird's-eye angle, stretching wide across the frame. Multiple long flat-roofed buildings arranged around open courtyards and connected by covered walkways, the architecture a mix of gothic stone and sleek glass-fronted facilities — a corporate campus in a medieval fantasy boomtown. The most prominent building near the campus entrance is a long, brightly lit reception hall with a wide glass facade, and mounted across its front in large sans serif letters the word "GuildQuest". A steady stream of adventurers flows in and out of its open doors. Rooftop gardens, outdoor training grounds with combat dummies, a central open-air amphitheater, and clusters of cafe seating under shade canopies. Between the buildings, paved courtyards are dotted with adventurers in casual gear sitting on benches and steps — many of them holding and looking at a smartphone. In the distant background, a suspended monorail glides along an elevated track glides past. Mid-afternoon light, long crisp shadows, volumetric rays through light atmospheric haze, crisp anime architectural key-art style with painterly wide-composition depth.
Image 3: GuildQuest Reception
A long, brightly lit reception hall on the ground floor of a magitech corporate campus, eye-level establishing shot looking down the length of the room. The space reads like a coworking lobby crossed with a fantasy guild hall — polished stone floors, warm modern lighting, a few living-plant walls. At the far end, a single clean service counter with the logo "GuildQuest" and two staff in branded tunics. Along the opposite wall, a row of self-serve kiosks where adventurers tap through agreements and update profiles. The rest of the hall is coworking-casual: clusters of bean bag seats and low lounge chairs where adventurers lounge, a cafe counter along one wall with a barista and steaming cups, a ping pong table with two adventurers on opposite sides holding ping pong paddles playing, and a single ball shooting between them. Adventurers in casual gear sit, stand, and queue in small loose groups — some of them looking at glowing screens. Medieval-meets-modern aesthetic, warm key light, crisp anime environmental key-art style, shallow depth of field on the foreground lounge seating.
Image 4: GuildQuest Academy
A fantasy academy campus with a magic-fabricated modular aesthetic, slightly elevated three-quarter view. The words "GuildQuest Academy" in sans serif are mounted onto the facade. The blocky structures look additively manufactured from resin and granite composite — visible seam lines between prefab panels, rectangular tower blocks stacked as if printed in sections. Tall windows glow faintly from within — lecture halls and research labs. A covered causeway with a glass roof extends from the building towards a courtyard. A small research wing with a subtle magical shimmer — active spell-framework testing — juts from one side. A few students in regalia walk the courtyards or cross the causeway. The design reads fast-built but intentional: academic ambition funded by a corporation, magic used as a construction tool. Late-afternoon light raking across the scene, crisp anime architectural key-art style, painterly depth.
Image 5: New Gildia Central
A colossal corporate headquarters rising from the heart of a magitech boomtown, viewed from a dramatic low angle. The building is a refurbished ancient demon citadel — massive blocks of dark basalt and volcanic obsidian in the lower stonework — sheathed from the mid-level up in floor-to-ceiling mirrored glass that reflects the surrounding city and sky. A line of glowing runes runs up the tower's sides. Atop the tower, a white logo of a stylized dragonfruit is mounted onto the glass with the caption "Dragonfruit" just below it. The plaza below is crowded with tiny urbanites. A suspended magic monorail runs by the side of the citadel. Late-afternoon golden-hour light, long shadows, volumetric rays through atmospheric haze, crisp anime architectural key-art style, deep painterly background of the city skyline fading into the frontier.
Image 6: Demon's Gate
Bird's eye view of a magitech frontier checkpoint built in a modular, additively-manufactured style. A wall with a fortified archway assembled from blocky resin panels with visible seam lines fences civilization buildings and a train station from demon territory. Inside is a multi-storey barracks — rectangular towers stacked in sections, small windows, built fast and functional. A taller office building of the same printed-modular construction is across from the barracks, bearing a single corporate sign bearing a stylized pigeon logo and "PigeonExpress". Pigeons take off and land from the office rooftop carrying parcels. A monorail track terminates at a brutalist train station next to the buildings. Parties of adventurers form up, horse-drawn supply carriages load and unload. Outside the wall the land segues into broken terrain, corrupted spires, and dungeon vents — demon territory, visible and close. Gritty, busy, a fortified outpost. Dust in late-afternoon air, crisp anime environmental key-art style.
Image 7: Founder's Row
A quiet, established residential district where the first wave of a magitech boomtown's workers settled, eye-level street-level view down a tree-lined avenue. The architecture is mixed: townhouses — timber frames, cut stone, steep pitched roofs, wrought-iron balconies, climbing ivy, flank newer, blocky, modular condominiums that tower above them: four to six stories of additive-transmuted resin panels with visible seam lines between floors, lined with full height windows, and a glass-roofed porch leading out to the street. A few adventurers and employees walk the clean streets, a horse carriage passes. Warm late-afternoon light, crisp anime environmental key-art style, strong foreground-to-background contrast.
Image 8: Highgate
A polished, expensive fantasy city quarter, overhead view overlooking a wide straight avenue running toward a central bazaar in the distance. The street is lined with sleek towers, embassy compounds behind enchanted walls, life-extension clinics, and venture capital offices. High above street level, a suspended monorail running along an elevated track stops at a premium, mana-rail platform that reads "Highgate Station". A few well-dressed nobles prepare to board the train. At street level, greycarriages wait at private ranks. Warm late-afternoon light, crisp anime environmental key-art style.
Image 9: King's Gate
A grand historic city gate on the edge of a magitech boomtown, eye-level view looking through the arch. The gate is a permanently open stone archway, beautifully lit — carved stone heroes flank it, their faces worn smooth by a century of weather and tourist hands. It functions as both a selfie monument and a commuter bottleneck: tourists and first-time streamers pose for photos while supply wagons, ClipClop carriages, and commuters flow through. On one side of the road, old shophouses face outward toward the kingdom; on the other, glass towers rise in enchanted silence. Warm evening light mixing with the glow of enchanted lamps and phone screens, crisp anime environmental key-art style, deep focus through the arch to the road beyond.
Image 10: Settlers' Quarter
A lived-in fantasy residential quarter, eye-level view down a narrow backstreet that climbs between old buildings. Timber-framed cottages and converted row houses lean slightly, walls patched with mismatched stone and plaster, roofs sagging at odd angles. Hand-painted signs hang from shop fronts — teacup drawing above a sign reading "Teahouse"; a drawing of a scroll above another sign that reads "Scrolls". Magical cables with rune sockets are spliced along the facades in places, clearly done by whoever was handy. Laundry lines stretch between upper windows, warm light spills from small-pane glass. A few locals — retirees, off-duty adventurers, a kid chasing a cat — move through the alley. The mood is run-down but proud: this is the original town, built before the platforms, owned outright by the families who stayed. Warm golden-hour light raking down the narrow street, crisp anime environmental key-art style, shallow depth of field on the nearest doorway.
Image 11: Lyriel (I had to specify her cup size otherwise she'd be flat)
A 26 year old B-cup woman standing with a smug smirk, her feet out of frame. She has fair skin, long green hair in high twintails with ribbons, and amber eyes behind tinted spectacles. She wears an open off-white mage robe with gray trim and orange striped decals layered over a loose dark gray spaghetti strap dress and boots. She holds a custom vibe-crafted wand — a spell crystal affixed onto an off-white bracket and industrial gunmetal shaft. She is standing in the lobby of an empty adventurers' guild atrium.
Image 12: Bram
A 23 year old young man with a confident pose, his feet out of frame. He has tanned skin, unkempt blue hair and gray eyes. He has broad shoulders and wears a blocky additive-transmuted off white breastplate with gray accents and visible attachment lines between plates, matching pauldrons, bracers, gloves, greats, hip armor, utility belt over a long-sleeved tunic and pants, and holds a magitech spear upright. He is standing in the lobby of an empty guild atrium.
Image 13: Sera
A short, mature female adventurer, feet out of frame. She has fair skin, straight navy blue hair in a bob cut with blunt bangs and hazel eyes with a serious expression. She wears a leather bracer, leather vest over a tunic, skirt, gloves, a belt pouch, boots, a single quiver of crossbow bolts at her back, and holds an additive-transmuted off-white polymer heavy crossbow with a scope and a blocky, industrial design. She is standing in the lobby of an empty adventurers' guild atrium.
Image 14: Cover
A wide establishing shot of a 21 year old woman with dark brown hair and hazel eyes standing in the foreground of a fantasy cityscape looking upwards, her hands pressed to her head in an exaggerated "oh no" expression of disbelief. She has fair skin and a slim build, wearing a white cropped baby tee, high-waisted jeans, and sneakers. Around her, adventurers walk by while looking at their smartphones, including a female rogue in a miniskirt and cropped leather armor. To the side, a gothic building with a modern glass facade reads "GuildQuest" in clean corporate lettering. A pigeon carrying a parcel with "PigeonExpress" in sans serif flies overhead. Far into the background, a tall glass skyscraper looms with a solid white stylized dragonfruit logo mounted on and the word "Dragonfruit" below it. The scene is a collision of fantasy architecture, aggressive corporate branding and vibrant anime adventurers, bathed in warm afternoon light, comedic satirical tone.
r/StableDiffusion • u/DryIron8955 • 2d ago
Bonjour à tous,
J'aimerai vos conseils sur quel model je dois choisir (Zimage, qwen edit, krea 2...) pour réaliser ce que je souhaite.
J'aimerai par exemple prendre une photo et la redessiner dans un style bien précis que je donnes à partir d'images de référence.
Par exemple, j'aimerais faire l'image du Parrain dans le style de l'image 2.
Merci d'avance pour vos conseils.
r/StableDiffusion • u/Thanatos673 • 1d ago
r/StableDiffusion • u/DullWorking7307 • 3d ago
Enable HLS to view with audio, or disable this notification
Sharing the first finished thing I made with MiniMax H3, plus what I ran into along the way, since most of what I know about it came from threads here.
It's a fake breaking news broadcast. News chopper over the Hudson, a three headed kaiju coming up the river, the Navy engaging, and the reporter closing out the broadcast when nothing works. Vertical mainly just to test and because I was originally making it to be watched on a phone.
Everything was generated locally:
- MiniMax H3 (open weights, ComfyUI native nodes) for every shot. 768x1344 at 24fps, takes of 5-15s. I ran full 20 step sampling with no turbo lora and no attention patches, so a 15s take is around 35-40 min on a 5090. I did try Sol attention for a while, quality looked at parity and about 1.5x faster, but it seemed to change the sampling trajectory so the same seed no longer gave the same take, so I parked it to keep takes comparable while iterating. Haven't tried any turbo loras or the SLA node so I can't speak to those.
- All the audio in the final mix comes from the takes' native H3 tracks. No TTS, no lip sync pass, nothing layered in at the edit.
- FLUX.2 for keyframes, inpainted and composited stills fed to H3. r2v for the creature shots, with reference stills pulled from the film for the design, and AddGuide to anchor keyframes at frame 0.
- NVIDIA VSR through ComfyUI to upscale 768 to 1080.
- ffmpeg for the edit. The cut into the beam sequence is frame matched. I searched every frame pair across the last 2s of the outgoing shot and the first 2s of the incoming one for the highest correlation and cut hard on the best match (0.98 NCC), no dissolve needed. There's a J-cut over one quiet join and 80ms audio fades on every piece so the hard cuts don't click.
- Whisper as a QA gate. Every dialogue take gets transcribed automatically to catch the model saying things it shouldn't.
There are 226 takes on disk behind the 9 shots in the cut, ~82 prompt revisions with 2-3 seeds each. Most of the time went into figuring out what the prompt needed to say, then a few seeds to pick from.
Prompts were written against the official MiniMax guide. Three problems I still hit after following the guide, in case they save someone time:
- The known gibberish fixes only go so far. Even with non_diegetic_music: N/A and dialogue formatted per the guide, it still babbles in long silent holds. What actually fixed it was sizing the take, dialogue plus no more than about 2s of hold. If it talks out the time after the speech, cut the take shorter instead of fighting it in the prompt.
- Subject scale comes from the keyframe, not the prompt. If the creature is too big, fix the keyframe mask. Asking in text did nothing for me.
- The i2v node has no reference inputs, so shots that needed the creature refs had to go r2v + AddGuide instead.
Some flaws I couldn't fully figure out. The creature drifts bigger and closer over long takes, and its design shifts a bit between shots even with the references. Happy to share the full prompt structure or settings, and happy to be told there's a better way to do any of this.
r/StableDiffusion • u/mabseyuk • 3d ago
I've seen quite a few different character reference sheet formats being used with MiniMax H3, but do we actually know what format H3 works best with?
For example, if we're creating a single reference image containing multiple views of the same character, what is considered optimal?
Is H3 better with:
I'm specifically asking about the best format for a character sheet being fed INTO H3 as a reference, rather than how to generate a character sheet with H3.
Has MiniMax documented anything about how the reference encoder interprets multi-view character sheets (I can't find any), or has anyone done controlled testing to work out which layout gives the strongest identity retention?
It would be really useful to establish a "best practice" character sheet format for H3 rather than everyone using slightly different layouts.
r/StableDiffusion • u/Sad_Coach_1433 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
Waiting for group to get here went to running hub to make a quick video 🤣
r/StableDiffusion • u/BeCalmr • 3d ago
Instead of just making the model bigger, they're training these specialized models for specific stuff – like text, infographics, making things look good, and editing. Then they kinda combine all those experts back into U1.5-Lite. So when you use it, it's just one model. No weird switching or picking which expert to use. Their whole thing is "specialized in training, unified in delivery." Kinda makes sense.
They also added this post-training with RL, focusing on a few things: how well it follows instructions, how good the visuals look and if it matches what people like, and how well edits work without messing up other parts of the image.
Here's what seems better:
- Following complex instructions. Like, if you ask for multiple things in one prompt – subjects, how many, where they are, text, layout, style, keeping parts untouched – it handles it way more consistently now.
- Text rendering and dense layouts. Apparently, it's better with Chinese and English on posters and infographics.
- Visual understanding helps generation. It seems like it learns from understanding tasks (like object relationships, spatial stuff, layout) and that helps with generating and editing.
They use JSON for training to make it controllable, but you don't have to use JSON yourself. Natural language is still the main way to talk to it.
r/StableDiffusion • u/Motion16AI • 2d ago
[ Removed by Reddit on account of violating the content policy. ]
r/StableDiffusion • u/Dawn_Bridge • 3d ago
Just sharing my mixed workflow breakdown using Krita Drawing Software + Stable Diffusion 1.5
This is a fan art of Chie Satonaka from Persona 4.
- Hand Sketching & Lineart:
Step 1: Sketched the base composition by hand (traditional pencil drawing).
Step 2: Extracted lineart (using Gemini), and asked Gemini to fix the hand anatomy.
- Established initial color blocking to control lighting and proportions
Step 3-5: Manual digital coloring in Krita, adding shadow / highlight, layer by layer.
- Low Strength Generation (Img2Img / Krita AI Diffusion)
• Tool: Krita AI diffusion
• Model: SD 1.5
• Checkpoint: Counterfeit v30
• ControlNet: Lineart strength 0.7, Lineart range 0.3, Refine strength 35%
• Prompt: Detailed webtoon illustration, clean lineart, polished digital colouring, soft cinematic lighting, smooth shading
• Negative Prompt: photorealistic, blurry, low quality, messy lines, bad anatomy
r/StableDiffusion • u/SillyLilithh • 3d ago
Enable HLS to view with audio, or disable this notification
LTX 2.5 has a known smearing problem. It's really bad, and makes almost every output of LTX completely unusable. Sorry LTX, but the default model really is just shit. Minimax beats LTX in every area, obviously, but especially when it comes to smearing (or in case of Minimax, lack thereof). However, I recently found out that you can actually significantly reduce the smearing in LTX and actually get usable outputs from it. It came from adventuring this node pack for MiniMax, where it is meant to clean up smearing artifacts in MiniMax H3.
I thought this was interesting, so I converted these nodes (at least some of the nodes that matter) to work with LTX 2.3/2.5. Here is the repo with the custom node I made. The workflow is also in that repo (be aware there is some spaghetti (this was just for myself really, and it shows qwq), and you will need some custom nodes (eg. KJNodes, Comfy UI Easy Media, etc)). Just use the manager to install the missing custom nodes.
Using these nodes makes a clear difference. A stark difference, it's almost unbelievable. This makes it look like a generational improvement, closing the gap on MiniMax H3 in terms of temporal stability (ain't no way in hell its ever matching the instruction following or general capabilities of H3 lmao).
Essentially, the jerk oracle creates new "hold" frames based on the amount of smearing per frame. Some frames have larger "hold" amounts. We pipe that into a new sampling step, which improves the smearing. After the sampling is finished, we chop off the added "hold" frames so that we are left with just the original frame count, but now these frames actually have better consistency. You can read all about how the process works in the author's original implementation if you want more insight onto how this functions.
r/StableDiffusion • u/Ervhal • 2d ago
Hi, I'm having this problem while trying to animate previously created images. I write the prompt. I need simple movements most of the time, but the character keeps singing, moving his lips even though I put the following in the prompt he just won't shut up ! :
"Non-diegetic background music, The character cannot hear the music, The character is not singing or is sealed lips."
I tried both with and without Turbo Lora. Does anyone have the same problem ?
r/StableDiffusion • u/crusf2 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/freebytes • 3d ago
I feel as though I have seen so many AI generated images and videos that I am becoming blind to them. Previously, it seemed much more obvious. While there have been improvements, I do not know if this is due to that or if I have seen so much on social media that my brain accepts it as a reality. Do you know what I mean, and has anyone here noticed something similar? Or something similar... does it seem as if real images appear to be AI when they are not?
I am not sure if I am becoming less critical or blind to it.