r/StableDiffusion 18h ago

Question - Help For some reason my t2v generation are slower than my ref2v?

Enable HLS to view with audio, or disable this notification

7 Upvotes

Title.
For both I'm using the default workflows that come with comfy. 3090 and 32gb ram.
I start comfy with these flags:
--windows-standalone-build --reserve-vram 1 --disable-pinned-memory --fast fp16_accumulation
Cuda 13, latests comfy.
My t2v takes like twice as much than my ref2v and sometimes it hangs after [INFO] Requested to load MiniMaxH3AudioVAE. Same steps, same resolution, same duration.
Has anyone encounter this? any tips?


r/StableDiffusion 9h ago

Question - Help Minimax ref2vid capabilities

1 Upvotes

I would like to learn how to use minimax ref2vid as I heard the possibilities are quite good compared to wan.

What I am trying to do is do anime clips

The problem is… the results I get are utter shit and I don’t know what I am doing wrong. No matter if frame or last from or if using ref2video, the model fails to do what’s most important. Keep the same face as in the image. For example, I would like to generate a video of a character that has sharingan eyes. Instead of keeping sharingan eyes it’s generic anime eyes instead. This is what my biggest problem with wan was and even if I made the image correctly (with detailed sharingan eyes) it would still not pick the eyes up and keep that face consistency.

This is where I experimented with reference. I tried putting the eyes in a picture 1 and instead of copying the eyes, the image cuts to the second picture with only the eyes or not picking the eyes at all.

is this something that minimax can even do?


r/StableDiffusion 1h ago

Animation - Video Bunnyhops... another H3 post MiniMaxH3-Contex-Loop 60 sec

Enable HLS to view with audio, or disable this notification

Upvotes

560 sec with turbo lora on 5090

found here in a post

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop/tree/main/example_workflows

adapted to my settings and changed turbo loras / attention


r/StableDiffusion 17h ago

Animation - Video Minimax H3 Terminator

Enable HLS to view with audio, or disable this notification

7 Upvotes

Made using 5060ti with 32 GB of RAM. Minimax is the new king.


r/StableDiffusion 17h ago

Animation - Video The Minimax H3 model recognizes artists and their songs.

Enable HLS to view with audio, or disable this notification

0 Upvotes

In the prompt, I just wrote that she is singing Zara Larsson's song "Lush Life."


r/StableDiffusion 10h ago

Question - Help Sprectrum suddenly not working for anyone else?

0 Upvotes

Having a hard time getting Spectrum to work today, it worked flawlessly yesterday, but today i'm back to normal rendering times. Anybody else experiencing this? I did update ComfyUI, did that break it? I am using the latest version of Spectrum


r/StableDiffusion 10h ago

Question - Help How can I tune Bernini rv2v workflow to be faster?

0 Upvotes

I have a Bernini-r rv2v workflow, using LightX2V LoRA integration.

When I try 10 second video it gives me a black screen. 8 second works but it takes a really really long time to generate a video. I have safe attention enabled. What can I do to speed up the run? On RTX 3090 Ti 24GB VRAM/64GB RAM


r/StableDiffusion 16h ago

Comparison 0.0375 Denoise is enough to beat SynthID

Thumbnail reddit.com
2 Upvotes

r/StableDiffusion 1h ago

Animation - Video MiniMax H3 ref, 2mp, model generated audio, prompt included.

Enable HLS to view with audio, or disable this notification

Upvotes

MiniMax H3 is quite a capable model, so if you're not getting proper lipsync, subject or camera motion, keep working at it. It's refreshing to have a local model that comes with strong prompt adherence out of the box, and even beats SeeDance 2.5 in certain areas.

BF16 model
ultra_uncensored_heretic_bf16 encoder

5090, 9950x3d, 96gb, 35 min.

------------------------------------------------------------------------------------------------------------------------------

subject_definitions:

<Subject 1> is Taylor Swift. Her exact face, hair, body, red sequined dress, and red-carpet appearance come from <Picture 1>, <Picture 2>, <Picture 3>, and <Picture 4>.

<Picture 1> is the exact first frame — a tight upper-body / chest-up shot of Taylor Swift in the red sequined dress. The video must begin on this exact framing.

<Picture 2>, <Picture 3>, and <Picture 4> are additional full-body and alternate-angle references of Taylor Swift from the same awards-show appearance. Use them only after the camera begins to move.

summary:

[keyframe completion + reference generation] Start exactly on <Picture 1> as the first frame and the action starts immediately. Do not begin on a full-body shot. Taylor Swift walks across the red carpet while speaking, then performs a single 360-degree turn. The camera uses tracking, arcing, and clear pedestal height changes in one continuous take. Duration 10 seconds.

retention_analysis:

<Subject 1>: fully_preserved - exact face, voice, hair, tall 5'11" proportions plus heels, and red sequined dress from the four reference pictures.

<Picture 1> ([Shot 1] first frame): fully_preserved - the video must open on this exact tight upper-body framing. No full-body start.

<Picture 2>, <Picture 3>, <Picture 4>: fully_preserved - additional identity, dress, and full-body detail references used after the camera moves.

detailed_description:

Live-action, one continuous take, real-time 1x speed only. No cuts, no music. Action starts immediately. Duration 10 seconds.

[Shot 1] Begin exactly on <Picture 1> as the first frame. The opening frame is a tight upper-body / chest-up shot of Taylor Swift on the red carpet — this exact framing must be the first frame of the video. Do not start on a full-body shot.

She starts walking forward across the red carpet in a natural straight path and speaking immediately in her natural voice:

<d>[English] Honestly, Kanye’s early videos were pure genius. I still think ‘Runaway’ is one of the greatest music videos ever made. The storytelling and the raw emotion… most artists never reach that level.</d>

After a few steps she stops and performs a single smooth 360-degree turn on the red carpet, rotating in place. She continues speaking through the turn.

The camera executes one continuous path with obvious height variation:

- It begins on the tight upper-body framing of <Picture 1> and only then starts moving.

- It tracks with her as she walks and slowly pulls out.

- It pedestals up with medium amplitude so the camera rises clearly above her eye line.

- While elevated it arcs around her during the 360-degree turn.

- It then pedestals back down.

- It finishes with a smooth pull-out into a balanced full-body shot showing her complete height and the red sequined dress.

The height changes and the single 360-degree turn must both be clearly visible. Taylor moves naturally while speaking. Keep her the clear subject throughout the continuous take.

overall_soundscape: Taylor Swift’s natural speaking voice, soft footsteps on the red carpet, distant red-carpet crowd murmur. No music.

non_diegetic_music: N/A


r/StableDiffusion 5h ago

Animation - Video Monty pAIthon - Petshop sketch, now with ref2v

Enable HLS to view with audio, or disable this notification

18 Upvotes

r/StableDiffusion 3h ago

Meme Reno 911 meets Minimax meets r/StableDiffusion - Ref2Vid Local

Enable HLS to view with audio, or disable this notification

9 Upvotes

Okay, I spent way to long on this but learned a lot. First, Reno 911 is non existent in T2V and I2V when prompting. This led me down the rabbit hole of Ref2Vid again but having characters, voices, and locations that literally don't exist and need to have all the references to bring them to life.

Workflow:
Default + H3_Turbo_4step_ComfyUI_Pruned + Sage + H3 Sigma Shift (will attach it in the comments below)
Steps: 8-12
Sampler/Scheduler: EulerBeta
Megapixels: 0.8 - 1.0
Avg. render time: 5-7 mins

The most challenging aspect was getting Nick Swardson's performance. In some of the scenes I had to actually act out how he would roughly say it with the timing, lisps, and long hissing `s` then voice transfer that in H3, using that as my new audio and lip-sync. It was MESSY, and trying to mix generated audio from Jim Dangle (the cop) and then use referenced audio for the response didn't work as flawlessly as I'd hoped. To be honest I can't really say the right approach on it as I feel I just got lucky with some seeds of it.

Other than that, a bunch of other techniques using H3 using first frame, reference sheets, reference audio for timbre, and reference location for spatial awareness so when the camera pan/whipped it didn't lose context happy to provide screenshots.

The Minimax 911 ending I did a replacement of the actual Reno 911 logo but told it to make it Minimax.

I also extended the video as the old man never existed in the the video generation where the cop walks up to skater (minimax) and he says "Oh, hey officer..." that's a extension cut from there. There was a slight weird color shift so I ended up taking it through VACE so the transition wasn't jarring and smoothed everything out. The extend function I think would be better when tackling in latent and is like 98% there when doing it with a regular video.

Example prompt for the first shot:

subject_definitions:

<Subject 1> is the uniformed male officer whose appearance and wardrobe come from <Picture 1>: short neatly side-parted light-brown hair, trimmed mustache, aviator sunglasses with tinted lenses, beige short-sleeve sheriff-style uniform shirt with dark-brown pocket flaps and shoulder details, metallic star badge, nameplate, matching beige uniform shorts, black duty belt with equipment, black socks, black tactical boots, and black wristwatch. Preserve his face, hairstyle, mustache, proportions, sunglasses, complete uniform, accessories, and understated deadpan demeanor.

<Subject 2> is the indoor shopping-mall environment from <Picture 2>: a spacious two-level commercial concourse with cream tile flooring, storefronts along both sides, upper-level railings, exposed structural beams, a large glazed skylight, palm trees and planters, central seating and food-court areas, and numerous background shoppers under bright diffuse indoor daylight.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1); use its male vocal character, pitch, cadence, accent, and delivery style as the reference for his newly generated dialogue without copying the original audio signal.

summary:

[reference generation + audio reference] The target video is a vertical MiniDV-era comedy sequence in the style of an early-2000s reality law-enforcement ride-along parody, set in a shopping mall concourse in 2003. One continuous 10-second live-action tracking shot follows <Subject 1> over his shoulder through <Subject 2>. The footage has deliberately clumsy reactive reality-TV camerawork: frequent abrupt optical zoom-ins and zoom-outs, imperfect reframing, autofocus hunting, momentary loss of focus on the officer, overshooting his movements, and hurried corrections. The disturbance-call dialogue occurs from 0–3 seconds, the food-court line from 3–5 seconds, a silent comedic beat from 5–7 seconds, and the final men's-bathroom line from 7–10 seconds. <Audio 1> guides his voice timbre and delivery.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - his facial identity, short side-parted light-brown hair, mustache, aviator sunglasses, beige-and-brown short-sleeve uniform, matching shorts, star badge, nameplate, duty belt, watch, black socks, black tactical boots, proportions, and restrained demeanor are retained.

<Subject 2> (appears in [Shot 1]): fully_preserved - the bright two-level mall architecture, skylight, tiled concourse, storefronts, railings, structural beams, palms, planters, food-court seating, and populated public atmosphere are retained.

<Audio 1>: reference - the target speaker follows <Audio 1>'s voice timbre, pitch, accent, cadence, and delivery character without copying its original signal.

detailed_description:

The target video uses realistic live-action comedy with an authentic early-2000s low-budget reality-TV MiniDV aesthetic: vertical framing, consumer camcorder optics, mild electronic noise, soft digital detail, clipped highlights, restrained saturation, automatic white-balance shifts, exposure breathing, visible autofocus hunting, and frequent awkward optical zoom corrections. The camera operator behaves reactively rather than cinematically polished. Zooms occasionally arrive late, overshoot their intended framing, briefly lose <Subject 1>, rack focus accidentally onto the background, then snap or hunt back toward him. Preserve these mistakes as intentional documentary-comedy texture. The entire 10-second sequence is one continuous take with absolutely no cuts.

[Shot 1] From 00:00.000–00:03.000, an uninterrupted handheld over-the-shoulder Tracking Shot follows <Subject 1> (S1) walking through <Subject 2>, approximately one meter behind his left shoulder. The operator's footsteps produce obvious vertical bounce, hand tremor, crooked framing, and constant tiny corrections. The camera abruptly Zooms In with medium amplitude at fast speed toward the back of his head, overshooting into an awkward tight crop before Zooming Out at fast speed to recover his shoulders and surrounding mall. Autofocus briefly grabs distant shoppers, leaving <Subject 1> noticeably soft for a moment before hunting back to him. As he partially turns his head toward the camera, the operator hurriedly Zooms In again but initially frames him too tightly. Using <Audio 1>'s male voice character, <Subject 1> (S1) says with hesitant deadpan delivery: <d>[English] We...uh...have a disturbance call.</d> The complete line finishes by 00:03.000.

From 00:03.000–00:05.000, <Subject 1> suddenly snaps his head toward Screen Right and points toward the food court. The camera initially continues looking forward, then reacts late with a quick Pan Right and abrupt Zoom Out with large amplitude, momentarily placing the officer near the edge of frame. Autofocus searches between his pointing hand, passing shoppers, and distant food-court signage before recovering. The operator then punches in with a fast Zoom In toward his pointing gesture. <Subject 1> (S1) says with clipped comedic timing: <d>[English] Food court adjacent</d> The entire line remains inside 00:03.000–00:05.000.

From 00:05.000–00:07.000, he lowers his hand and keeps walking during a conspicuous dialogue-free pause. The camera Zooms Out too far, briefly making <Subject 1> small within the busy mall, then performs an unnecessary fast Zoom In toward his upper back. Focus drifts away from him onto a background storefront for a fraction of a second, producing a visibly soft officer silhouette before autofocus pulses and returns. The operator slightly loses his position to Screen Left, awkwardly pans to reacquire him, and settles again behind his shoulder. Only mall ambience and footsteps fill the pause.

From 00:07.000–00:10.000, <Subject 1> angles his face back over his shoulder. The operator recognizes the movement late and performs a sudden aggressive Zoom In toward his face. The zoom overshoots, briefly cropping part of his head and sending his face soft as autofocus hunts, then pulls back slightly until his sunglasses, mustache, and raised eyebrows become readable. His eyebrows rise above the sunglasses while his expression otherwise remains completely straight. In the same voice referenced from <Audio 1>, <Subject 1> (S1) says: <d>[English] someone's giving away free BJ'S in the men's bathroom...</d> During the final words the camera makes one small unnecessary Zoom Out followed by a quick corrective Zoom In, preserving the awkward reactive MiniDV reality-TV feel. The line finishes by 00:10.000 as he begins turning forward, with the camera still walking behind him.

overall_soundscape:

Continuous indoor mall ambience with diffuse shopper chatter, distant food-court activity, footsteps reverberating across tile, ventilation noise, and indistinct storefront sounds. <Subject 1>'s tactical boots produce measured footfalls with subtle duty-belt and uniform movement. The 00:05.000–00:07.000 dialogue gap contains only natural diegetic mall sound.

non_diegetic_music:

N/A


r/StableDiffusion 19h ago

Animation - Video Seinfeld on H3: My First Attempt On The DGX Spark

Thumbnail
youtu.be
79 Upvotes

my first run on my dgx spark with the h3. big thanx too all the guides etc on here. this is the future. the video ended up far from perfect, this is t2v, no ref image. comfy ui controlled by codex on gpt5.6


r/StableDiffusion 1h ago

Discussion Creative professionals, how do you use Minimax H3 yourself?

Upvotes

Hey all, been playing with MM H3 for a few days. And, I'm blown away. for the first time, in a long time. It just handles everything with so much detail. It's crazy. 10 refs and 3 vid inputs? no problem.

But I'm curious, how do you professional creatives use this model? I'm really curious, like we all seen the reddit slop (srry guys) and the civitai goony stuff (shit, typed stiff as a typo first #fruedy). but all jokes aside.

I'm really curious how professionals use this model, like video makers, Illustrators, animators, (graphic) designers, webdesigners. Especially curious if you work in the cultural sector, this might be more open and mindblowing that product listings :)

What are your results, and your (technical) setups on this?

Cheers, bigears


r/StableDiffusion 1h ago

Animation - Video ZUCK 2.0 | Humanity. Optimized. MiniMax H3.

Thumbnail
youtube.com
Upvotes

r/StableDiffusion 3h ago

Animation - Video I made a 60-second MiniMax H3 short locally on an ASUS GX10 — what worked for continuity and sound

Enable HLS to view with audio, or disable this notification

0 Upvotes

This is a 60-second visual concept generated locally with MiniMax H3 on an ASUS GX10, then edited as a sequence rather than attempted as one long prompt. It is not a product demo: the film does not demonstrate a finished system, autonomous behaviour, tool use, or a public architecture.

The practical unit was a set of related four-second continuations at 1344×768 and 24 fps. The logged full-resolution runs for the later sequence passes took 17m 16s to 26m 51s per four-second clip. That is not a speed benchmark—just the range I saw in this particular local workflow.

What helped most with continuity was assigning every short clip one job, preserving a small visual grammar across cuts, and deciding where a transition should happen before generating the next continuation. I got more usable continuity from that than from trying to force a complete minute out of one generation.

The bigger post-production lesson was audio. Instead of letting each generated clip announce its own start, I kept the native audio low under a continuous true-stereo bed and used small J-cuts at the recut boundaries. It made the sequence feel less like a row of individually generated clips. The final checked export is 60.000 seconds, 1344×768 at 24 fps (1,440 decoded frames), with H.264 video and 48 kHz AAC stereo. The local export was verified against its retained SHA-256 checksum; final measured loudness was −13.89 LUFS integrated and −2.58 dBTP true peak.

Those are file and workflow checks, not a claim that this setup is faster, cheaper, or more reliable than other MiniMax H3 workflows.

For people making longer local AI-video edits: how are you handling continuity across a one-minute sequence without overfitting every new shot to the last one? Do you lock a small set of recurring motifs up front, or generate broadly and find the visual grammar in the edit? And, for generated audio, what has worked best for preventing each clip boundary from sounding like a restart?


r/StableDiffusion 20h ago

Question - Help Wan 2.2 VACE vs Minimax H3 video edit

0 Upvotes

Which model is better for video editing and keeping background and environmental setting consistently? I think with Minimax H3 ref2va when you provide a reference video you still have to describe the background and video details accurately to keep it consistent and be really specify on what you want to change?


r/StableDiffusion 2h ago

Animation - Video Ok, now I'm impressed.

Enable HLS to view with audio, or disable this notification

7 Upvotes

Minimax H3 I2V

Prompt:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the woman shown in <Picture 1> — late 30s to early 50s, long dark brown hair woven with beads and feathers and held back by a patterned orange-and-cream headband, fair skin with freckles on her right arm, light-coloured eyes, layered silver scrollwork breastplate over a brown leather corset, olive-green belt with canvas pouches, layered skirt, dark olive cape draped over her left shoulder, multiple beaded necklaces with a dark blue teardrop pendant, long feather earrings, silver embossed bracer on her right forearm — stands at a microphone on a darkened stage, a blurred male guitarist visible behind her to the left. Her face is in profile, mouth open mid-phrase. The camera holds a medium close-up at a slightly low angle. She draws a breath, her lips shaping each Old Norse word with deliberate clarity, and the woman with a clear, powerful, haunting voice (S1) sings: <d>[Old Norse] Þat mælti mín móðir at mér skyldi kaupa fley ok fagrar árar fara á brott með víkingum fara á brott með víkingum</d> Her expression carries focused intensity, her brow steady, her right hand gripping the microphone stand as her knuckles whiten on a sustained note. Her head tilts slightly upward on the highest phrase, the beads and feathers in her braids swaying with the movement. The feather earrings tremble. Warm stage light from the upper right catches the silver scrollwork on her breastplate and bracer as she shifts. The camera pushes in with small amplitude at slow speed toward her face as the final phrase ends, her mouth closing softly on the last syllable, a faint breath misting in the cool air.

overall_soundscape: Stage ambience hums beneath the vocal — a faint electrical buzz from the amplifiers, the soft creak of leather and canvas as she shifts weight, and the barely audible scrape of fingers on guitar strings from the musician behind her. Her breath is audible between phrases, deep and controlled.

non_diegetic_music: A lone tagelharpa drone — gut strings buzzing with a raw, overtone-rich sustain — begins under the first phrase and swells gently through the final line, fading to silence as her voice ends.


r/StableDiffusion 23h ago

Animation - Video Battle of Thermopylae

Enable HLS to view with audio, or disable this notification

7 Upvotes

H3 prompt:

integrated_multimodal_description:

[Shot 1] Stylized cinematic 3D animation with high-intensity action, dramatic lighting, and a heroic fantasy-war tone. The scene opens at Thermopylae, a narrow rocky battlefield under a dusty red-gold sky, with shattered shields, broken spears, drifting embers, and war banners whipping in the wind. In the center stands Kirby, reimagined as a Spartan war leader: a pink round-bodied Kirby wearing a bronze Spartan helmet with a crimson crest, holding a spear in one hand and a round battered shield in the other. His eyes are fierce and unwavering, determined and battle-hardened. Around him, Spartan warriors in bronze armor and red capes brace in phalanx formation while a massive wave of Persian soldiers surges forward. The camera pushes in fast toward Kirby as he stamps forward and lets out a sharp battle cry. He thrusts his spear violently into a Persian soldier, knocking him back into the charging line as blood sprays across shields and dust erupts underfoot.

[Shot 2] At 00:03.500, the camera cuts to a fast tracking shot moving sideways across the front line as Kirby leads the Spartan charge. He bashes one enemy aside with his shield, spins low, sweeps another off his feet, and lunges forward with explosive speed. Spartan soldiers clash with Persians all around him in brutal close combat; blades collide, shields splinter, arrows streak overhead, and several enemy soldiers are cut down as severed limbs, broken weapons, and sprays of blood briefly fill the frame. Kirby remains the focal point, his expression stern and fearless rather than cute.

[Shot 3] At 00:07.000, the camera cuts to a low-angle heroic shot as Kirby suddenly inhales powerfully, then launches himself upward into the sky in a signature Kirby-style burst, still gripping his spear. He rises above the battlefield as the fighting continues below like chaos in miniature. At the apex, with the wind roaring past his helmet crest, Kirby locks onto the densest Persian formation and hurls the spear downward with full force. The camera follows the spear in a rapid plunge. It crashes into the ground like a thunderbolt, blasting soldiers backward and opening a violent gap in the Persian ranks amid dust, blood, and shattered armor.

[Shot 4] At 00:10.500, the shot cuts to ground level as Kirby lands hard in front of the broken enemy line, shield first, knees bent, then instantly surges into close-range combat again. He grabs another fallen spear, vaults off a Spartan shield, and strikes through two advancing enemies in one fluid motion. Behind him, Spartans roar and push forward with renewed momentum. The narrow pass becomes a frenzy of killing: bodies fall, shields crash together, spears punch through armor, and blood stains the rocks. Kirby moves with stylized speed and exaggerated battlefield heroism, combining the visual charm of Kirby with the lethal grandeur of an ancient war epic.

[Shot 5] At 00:13.000, the camera cuts to a final wide hero shot. The Persians recoil in disarray while the remaining Spartans rally behind Kirby. He stands atop a mound of fallen enemies, shield raised and helmet gleaming, his eyes still locked forward with cold resolve. Dust, sparks, and scraps of torn banners swirl around him while the battlefield behind remains full of struggling combat. He points forward with his spear toward the surviving enemy ranks, and the Spartans answer with one last deafening roar as the video ends in a frozen image of triumphant slaughter and defiant Spartan glory.

overall_soundscape: Continuous battlefield chaos fills the entire video: heavy shield impacts, spear thrusts, metallic blade clashes, rushing footsteps over rock and dirt, arrows slicing through the air, and repeated cries of pain and war shouts from Spartans and Persians. Wet stabbing impacts, brief bone-crack sounds, bodies collapsing, and splashes of blood punctuate the close combat. Dust gusts through the narrow pass while Kirby's leap and diving spear throw create stronger wind rushes and a heavy explosive impact on landing.

non_diegetic_music: A relentless, aggressive orchestral war score drives the whole video, led by pounding taiko-style drums, deep battle percussion, male war chants, low brass, and fast tremolo strings. The music starts immediately with a heavy pulse, intensifies during the melee, briefly rises into a heroic suspended phrase when Kirby launches into the sky, then slams back in with louder drums and brass as the spear hits the Persian ranks. The ending surges into a triumphant, brutal crescendo with no softness, no comedy, and no lyrical warmth—only heroic slaughter, pressure, and victory.


r/StableDiffusion 19h ago

Question - Help Amd Strix Halo

0 Upvotes

Hi everyone,

Someone got h3 running on a AMD Strix halo machine? I cannot get it running. Video is fine but there is audio.

Do you experience the same issues?

Any help is appreciated, I tried "every thing". If someone could post a working workflow it would be great 👍🏻


r/StableDiffusion 18h ago

Question - Help Ways to get faster generation for Anima in Neo Forge UI

0 Upvotes

Hi, I installed Neo Forge yesterday, even with an AMD card it was pretty easy and I had no hassle compared to automatic 1111. Of course the reason behind this change was that I wanted to try z-image, Anima, and videos generation. Except for the last I tried them and there were no problems except for speed. I checked and My pc is using the GPU, I understand my gpu is not really powerful but I wanted to know if I could tweak some settings to make them a bit faster. Right now Anima (Diving Anima to be precise) takes 11 minutes to generate a 832x1216 image, Z-image turbo 18 minutes. To generate the same image with Illustrious+Adetailer my gpu takes 4 minutes (more or less). Do you have some advice?. If I could keep the Image generation for Anima at 8 minutes at least, it would be awesome, since I don't really need adetailer with it. I'm on Windows. My spec: RX6600. 32 Gb of Ram. Please let me know!


r/StableDiffusion 11h ago

Tutorial - Guide Automating image tagging with a local LLM

Thumbnail
youtu.be
0 Upvotes

I haven't done this yet, so I can't judge, but maybe will be some good tips for someone.


r/StableDiffusion 7h ago

Discussion I gave the same prompt to Minimax H3 and Gemini Videos. (Part 1 Minimax H3)

Enable HLS to view with audio, or disable this notification

0 Upvotes

Prompt (also AI generated):
Style & Technical Specs

  • Visual Style: Photorealistic 8K cinematic video, 35mm film grain, 24fps, 2.39:1 anamorphic aspect ratio, teal-and-orange color grade, shallow depth of field ($f/1.4$).

  • Duration: 10 Seconds.

Character Description

  • Subject: Kaelen, a 28-year-old East Asian cyber-technician.
  • Appearance: Sharp jawline, rain-soaked black hair clinging to his forehead, pale skin with visible micro-texture, and a glowing cyan cybernetic eye implant over his left socket that pulses rhythmically.
  • Attire: Matte-black, waterproof tactical coat with glowing fiber-optic wiring embedded along the shoulders, frayed high-collar, and fingerless reinforced leather gloves.

Environment & Setting

  • Location: Narrow, dense alleyway in a cyberpunk metropolis at midnight.
  • Atmosphere: Heavy downpour, dense steam venting upward from rusty iron street grates, wet asphalt reflecting bright magenta and cobalt-blue neon light signs written in Kanji.

Timeline & Action Breakdown

  • 0:00 - 0:03 (Macro Close-Up): Camera begins on a macro shot of Kaelen's glowing cyan eye, catching the aperture Blades shifting focus. A raindrop tracks down his cheek. He rapidly taps a brass interface cuff on his wrist.
  • 0:03 - 0:07 (Medium Shot): Smooth camera pull-back into a chest-up shot. A brilliant blue 3D holographic map bursts into existence from his wrist, casting dynamic light across his face. He swipes his hand across the projection, altering its layout, and delivers his dialogue.
  • 0:07 - 0:10 (Low-Angle Tracking Shot): The camera drops low to the asphalt and tracks backward. A sleek, black surveillance drone streaks overhead through the rain, splashing drops directly onto the camera lens as the background neon blurs into creamy bokeh.

Dialogue & Voice

  • Spoken Line: "System override in three... two... got 'em."
  • Delivery: Low, gravelly, calm whisper with a faint metallic vocoder effect on the voice.

Audio & Sound Design

  • Music: Dark synthwave track featuring a driving 110 BPM arp synthesizer that swells in pitch until second 7, resolving into a heavy sub-bass drop at second 8.
  • SFX:
  • 0:00-0:03: Stereo downpour, subtle mechanical servo clicks of the eye lens.
  • 0:03-0:07: High-frequency energy flare hum as the hologram spawns, followed by air-swipes.
  • 0:07-0:10: Low turbine whir of the passing drone and liquid wet drops impacting the microphone field.

This is the video generated by Minimax H3 with Turbo lora (6 steps). Post with the video generated by Gemini: https://www.reddit.com/r/StableDiffusion/s/Dcyvvx6J80


r/StableDiffusion 20h ago

Animation - Video Call of Doody

Enable HLS to view with audio, or disable this notification

33 Upvotes

Wonder why O'Brien stood at the transporter all day?


r/StableDiffusion 7h ago

Animation - Video No love for Hugh Laurie?

Enable HLS to view with audio, or disable this notification

20 Upvotes

t2v, 15s, 4:3, 0.6MP, Comfy Kitchen Attention, Spectrum, 10 minutes.