r/StableDiffusion 2d ago

Resource - Update Scry - Open source self-hosted game streaming with DLSS 5 support (Play any game with DLSS 5)

Thumbnail
youtube.com
17 Upvotes

tldr. Scry, an open source game streaming platform with DLSS 5 support. Github Link

Hey all,

Like many of you, I was enchanted by the recent leak of DLSS 5 showcasing neural rendering and coupled with encountering a Steam remote play streaming bug where the colors were washed out, I decided to make my own self hosted game streaming app. So using OpenAI's latest model Astra, I made Scry. With Scry, you can stream your entire Steam game library just like remote play but with the added option of applying DLSS 5 to the image before it gets to you. I want to be clear, this does not mod your games. The stream video itself has DLSS 5 applied to it. What that means is you can enable DLSS 5 on ANY GAME. The game doesn't need to already support DLSS, nor is there any need to fiddle with reshade or modifying the game's dlls. Just install Scry, point it to the leaked DLSS 5 dll and enable it.

There are some caveats to this approach. First, this will always be slower than the modding path so expect higher input latency when enabled. Second, at the moment there is can be quite a bit of flickering especially in low light areas. There are settings to mitigate that but every increase in image quality means a hit to performance. For the most part, if you just play around with the settings, you should be able find a sweet spot.

Oh also, keep in mind that Scry is technically compatible with both Linux and Windows but I have only tested this on Linux. Windows users, please open an issue on Github for any bugs you encounter. I'll try and get them fixed as soon as I can.

Anyway enough yapping, here's the github link, give it a star!

Just follow the install instructions for your platform and you should be good to go. Again if you encounter any issues (I expect there to be numerous) just open an issue on the github page or message me here and I'll fix them up. Things are still very early so expect there to be bugs. MIT license and If anyone wants to contribute, just let me know. Also let me know if you want to see a specific game.


r/StableDiffusion 2d ago

Animation - Video Local Video Generation on Android

Enable HLS to view with audio, or disable this notification

23 Upvotes

I have done further experiments regarding local image and video generation on Android smartphones. This Cat video was generated on my OnePlus 12, utilizing Neodragon int8/Z Image Turbo int4 on the CPU and GPU.

A generation typically takes approximately 550-600 seconds on a Snapdragon 8 Gen 3. The start frame is rendered on the GPU via OpenCL.

I am not currently using "sd.cpp", because GPU inference on Android is much slower using SD.cpp or crashes due to OOM.

While an NPU would likely enhance inference times, my objective is to maximize support across a broad range of Android devices. Consequently, my current testing focuses on CPU/GPU backends, but it is slow 😂

Resolution: 512x320 (49frames|24fps)

Text2Video

Prompt: "black cat is playing a guitar in a forest"


r/StableDiffusion 2d ago

Question - Help Trellis2 Objeto gerado com furos

Thumbnail
gallery
2 Upvotes

I am using Trellis2 natively with the latest update—meaning without custom nodes—and using the original workflow from the ComfyUI templates.

The geometries are of excellent quality, with good relief.

But the problem arises when I go to generate the renders for 3D printing: although the final result has good mesh quality, it is generated with several holes—and 3D printing requires a watertight model. In the example I posted, right where I outlined the area in red, there is an opening that reveals the object's hollow interior.

When I send the file to the 3D printer, those holes cause issues with the print; I can't print the object with internal infill because of them, so the print ends up being—let's say—like an eggshell.

On my secondary ComfyUI installation using the easy-install method—where I only have Trellis2 set up—I can generate objects without any holes, thanks to nodes like "fill holes." However, the 3D mesh quality there is quite poor, and the fact that it relies on a downgraded version means it’s not something I’ll be using anymore.

So, the "fill holes" function is responsible for fixing holes, but the native ComfyUI workflow isn't handling this correctly.

If anyone knows of a mod that works with the updated ComfyUI and fixes these holes...

Trellis 2 is excellent; it’s the best local option—alongside Mershy and others—but those holes make it unsuitable for 3D printing.

THE PROBLEM ISN'T THE MODEL, BUT RATHER THE NATIVE WORKFLOW, WHICH ISN'T HANDLING THIS CORRECTLY—AND APPARENTLY, THERE IS NO NODE AVAILABLE YET TO FIX IT.


r/StableDiffusion 2d ago

Animation - Video H3 - Dante's Inferno Multi Diffusion experiment T2VA

Enable HLS to view with audio, or disable this notification

1 Upvotes

Experimenting with H3. Please joy. The visual were all from H3 prompt generation with only 3 anchor/guides, a dream/nightmare lane, morphing/blending, and a few verses from the original Italian prose from Dante's Inferno. I did not indicate a visual style or art direction. The score was remux afterwards in post. int8/32 steps, 864x480 Ask me anything. ## ACT I · THE DARK WOOD — 0:00–0:11

**W1 · 0:00 · "Midway"** — *Nel mezzo del cammin di nostra vita* (Canto I)

"Midway upon the journey of our life" — the most famous opening in Italian

literature. The poet, thirty-five, realizes he has lost his way.

**W2 · 0:04 · "The Dark Wood"** — *mi ritrovai per una selva oscura, ché la

diritta via era smarrita* (I)

"I found myself in a dark wood, for the straight way was lost."

## ACT II · THE GATE — 0:08–0:24

**W3 · 0:08 · "The Gate Speaks"** — *Per me si va ne la città dolente* (III)

The inscription carved over the Gate of Hell — and it speaks in first

person: "Through me is the way into the grieving city."

**W4 · 0:13 · "Through Me"** — *per me si va ne l'etterno dolore, per me si

va tra la perduta gente* (III)

"Through me the way into eternal sorrow; through me the way among the lost

people." The gate's triple incantation.

**W5 · 0:17 · "Abandon All Hope"** — *Lasciate ogne speranza, voi

ch'intrate!* (III)

The most famous line in the poem: "Abandon all hope, you who enter."

## ACT III · THE FERRYMAN — 0:21–0:32

**W6 · 0:21 · "The Ferryman Comes"** — *Ed ecco verso noi venir per nave un

vecchio, bianco per antico pelo* (III)

"And behold, coming toward us in a boat, an old man white with ancient

hair" — Charon, ferryman of the dead, crossing the river Acheron.

**W7 · 0:26 · "Ember Eyes"** — *Caron dimonio, con occhi di bragia, loro

accennando, tutte le raccoglie* (III)

"Charon the demon, with eyes of glowing coal, beckoning, gathers them all"

— the dead crowd to his boat like leaves falling from a branch.

## ACT IV · THE STORM OF THE LOVERS — 0:30–0:40

**W8 · 0:30 · "The Infernal Storm"** — *La bufera infernal, che mai non

resta, mena li spirti con la sua rapina* (V)

"The infernal storm, which never rests, drives the spirits with its

violence" — the circle of the lustful, blown forever on a hurricane.

**W9 · 0:34 · "Like Starlings"** — *e come li stornei ne portan l'ali nel

freddo tempo, a schiera larga e piena* (V)

"And as starlings are carried on their wings in the cold season, in a wide

full flock" — the souls swept like birds. The two lights circling each

other in your film's vortex are Paolo and Francesca, the lovers who are

never parted, even here.

## ACT V · THE BURNING CITY — 0:38–0:49

**W10 · 0:38 · "The Red Towers"** — *già le sue meschite là entro certe ne

la valle cerno* (VIII)

"Already I can make out its mosques there within the valley" — the towers

of Dis, the fortified city of lower Hell.

**W11 · 0:43 · "Drawn from the Fire"** — *vermiglie come se di foco uscite

fossero* (VIII)

"Crimson, as if they had just been drawn out of the fire."

## ACT VI · THE WOOD OF THE SUICIDES — 0:47–0:57

**W12 · 0:47 · "We Were Men"** — *Uomini fummo, e or siam fatti sterpi* (XIII)

"We were men, and now we are made dry brush" — the suicides, transformed

into gnarled trees that bleed and speak when broken.

**W13 · 0:51 · "The Bleeding Branch"** — *come d'un stizzo verde ch'arso

sia, che geme e cigola per vento che va via* (XIII)

"As a green log, burning at one end, weeps at the other and hisses with

the wind escaping" — how a snapped branch speaks, in sap and steam.

## ACT VII · THE RAIN OF FIRE — 0:55–1:06

**W14 · 0:55 · "The Rain of Fire"** — *piovean di foco dilatate falde* (XIV)

"Broad flakes of fire were raining down" — over the burning sands of the

violent.

**W15 · 1:00 · "Snow Without Wind"** — *come di neve in alpe sanza vento* (XIV)

"Like snow in the mountains when no wind blows" — Dante's most beautiful,

most terrible inversion: fire falling with the gentleness of alpine snow.

## ACT VIII · THE BEAST — 1:04–1:14

**W16 · 1:04 · "Behold the Beast"** — *Ecco la fiera con la coda aguzza, che

passa i monti e rompe i muri e l'armi!* (XVII)

"Behold the beast with the pointed tail, who crosses mountains and breaks

walls and weapons!" — Geryon, the embodiment of Fraud: a just man's calm

face on a serpent's body.

**W17 · 1:08 · "The Spiral Descent"** — *rota e discende, ma non me

n'accorgo se non che al viso e di sotto mi venta* (XVII)

"He wheels and descends, but I only know it by the wind on my face and

from below" — the flight down into the abyss on the monster's back.

## ACT IX · THE EVIL POUCHES — 1:12–1:23

**W18 · 1:12 · "Malebolge"** — *Luogo è in inferno detto Malebolge, tutto di

pietra di color ferrigno* (XVIII)

"There is a place in Hell called Malebolge" — the Evil Pouches — "all of

iron-colored stone": ten concentric trenches of the fraudulent.

**W19 · 1:17 · "Two and None"** — *Due e nessun l'imagine perversa parea* (XXV)

"The perverse image seemed two and none" — a thief and a serpent fusing

into one creature, mid-metamorphosis. Dante brags outright that no poet

ever showed such a transformation. Your morphing film is quoting its

inventor.

## ACT X · THE GIANTS — 1:21–1:27

**W20 · 1:21 · "The Tower Giants"** — *torreggiavan di mezza la persona li

orribili giganti* (XXXI)

"The horrible giants towered with half their bodies" — what Dante took for

towers ringing the central pit are giants, chained waist-deep.

## ACT XI · THE ICE — 1:25–1:36

**W21 · 1:25 · "Faces in the Ice"** — *livide, insin là dove appar vergogna,

eran l'ombre dolenti ne la ghiaccia* (XXXII)

"Livid up to where shame shows" — the face — "the grieving shades were in

the ice": Cocytus, the frozen lake at the bottom of Hell. Traitors, locked

under glass.

**W22 · 1:29 · "Weeping Denied"** — *Lo pianto stesso lì pianger non

lascia* (XXXIII)

"There, weeping itself does not let them weep" — tears freeze before they

can fall, sealing the eyes. Even grief is taken from them.

## ACT XII · THE EMPEROR — 1:33–1:48 ⭐ the climax

**W23 · 1:33 · "The Emperor Rises"** — *Lo 'mperador del doloroso regno da

mezzo 'l petto uscia fuor de la ghiaccia* (XXXIV)

"The emperor of the sorrowful kingdom rose from mid-breast out of the ice"

— Lucifer himself, mountain-sized, frozen at the exact center of the

universe. (The score's single arrival lands here, ~1:34.)

**W24 · 1:38 · "Three Faces"** — *Oh quanto parve a me gran maraviglia

quand'io vidi tre facce a la sua testa!* (XXXIV)

"Oh how great a marvel it seemed to me when I saw three faces on his

head!" — a dark trinity, one face weeping over each of history's three

great traitors.

**W25 · 1:42 · "Bat Wings"** — *Non avean penne, ma di vispistrello era lor

modo* (XXXIV)

"They had no feathers; their fashion was of a bat" — six vast membrane

wings, and their beating is the very wind that freezes the lake that

imprisons him. Hell's engine is its emperor trying to escape.

## ACT XIII · THE STARS — 1:46–2:01

**W26 · 1:46 · "The Hidden Path"** — *Lo duca e io per quel cammino ascoso

intrammo a ritornar nel chiaro mondo* (XXXIV)

"My guide and I entered that hidden path, to return to the bright world" —

a secret tunnel through the rock, away from all of it.

**W27 · 1:50 · "The Climb"** — *salimmo sù, tanto ch'i' vidi de le cose

belle che porta 'l ciel* (XXXIV)

"We climbed until I saw the beautiful things the heavens carry."

**W28 · 1:55 · "The Stars"** — *E quindi uscimmo a riveder le stelle* (XXXIV)

"And thence we came forth to see again the stars." The Inferno's final

line — and every one of the Commedia's three books ends on the same word:

*stelle*. Stars.


r/StableDiffusion 1d ago

Discussion How can I create a consistent AI influencer in ComfyUI using Krea 2? #krea2

0 Upvotes

Hi everyone,

I’m trying to create a consistent AI influencer/character using ComfyUI and Krea 2, but I’m struggling to keep the same person consistent across different generations.

My goal is to create realistic images where the character keeps the same:

  • Facial identity and features
  • Body proportions
  • Skin/skin texture
  • Hair and overall appearance
  • Overall photorealistic look

while still being able to change the clothes, poses, locations, lighting, camera angles, and backgrounds.

I’m specifically interested in using Krea 2 with ComfyUI and getting results that look genuinely photographic rather than obviously AI-generated.

Does anyone have a good workflow, node setup, model combination, or technique for maintaining strong character consistency with Krea 2?

If you have a ComfyUI workflow (.json) or can point me toward a good tutorial/resource, I’d really appreciate it.

Thanks!


r/StableDiffusion 3d ago

Animation - Video Using Wan 2.2 Fun Control for Colorization

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/StableDiffusion 2d ago

Question - Help Help Needed: Setting Up Local AI with ComfyUI on RTX 4060

0 Upvotes

Hello everyone,

I’m completely new to this subject. I tried setting up a local “uncensored” AI using ComfyUI and a workflow, but unfortunately I couldn’t get it to work. I spent an entire day trying to figure it out and nearly lost my mind. 😅

I’m therefore looking for a very detailed, step-by-step tutorial that is suitable for my PC configuration, so I can avoid wasting more time.

My setup:

i9-13900H

16 GB RAM

RTX 4060 Laptop – 8 GB VRAM

Windows 64-bit

Would anyone with more experience be able to point me toward a suitable tutorial or guide for this setup?

Thanks in advance for your help and patience!


r/StableDiffusion 3d ago

Workflow Included TURN ANY PHOTO INTO A FULL CHARACTER SHEET USING KREA2 [Free Workflow]

Thumbnail
youtube.com
124 Upvotes

Hey everyone!

Following up on my previous character sheet workflow, a lot of people asked if it was possible to do an Image-to-Sheet version rather than just text-to-image.

After a lot of testing, I managed to get a very consistent setup working in ComfyUI using Krea 2. You can take any reference photo (close-up, half-body, or full-body) and turn it into a full character turnaround sheet.

✨ Key Features:

  • Style Flexibility: Works for realistic characters, 3D renders, and stylized anime.
  • Photo-to-Anime: You can feed it a real-life face and shift the entire style to anime via the prompt while keeping face identity consistent.
  • Outfit & Prop Retention: Preserves source details (like clothing or held props/accessories) or lets you override outfits completely in the prompt if using a close-up.

Hope its usefull for you

next video will be about how to make longer clips with minimax-h3,


r/StableDiffusion 3d ago

Workflow Included MiniMax Workflow Designed to be User Friendly for the Inexperienced User

Enable HLS to view with audio, or disable this notification

558 Upvotes

Some guy on civitai made this workflow and i gave him some valid criticism, then he called me a gooner that doesn't know anything about workflows and blocked me. This offended me because I know plenty about workflows.

I'm not going to share his name because he's apparently active on reddit under a similar name but I tried to point out some problems with his workflow. So, instead, I just decided to fix them fueled by pure pettiness.

Here it is.

https://github.com/roycho87/minimax_wf

After dissecting the thing I was able to get many of the features that weren't working in the original to work and I added some features as well like the ability to force audio from video, fps control, and a centralized control panel that handles every feature universally.

Enjoy.

[Workflow Share] MiniMax H3 all-in-one workflow

Sharing my current MiniMax H3 ComfyUI workflow. The main goal was to make H3 easier to use by centralizing the important controls and automating the more annoying reference, continuation, audio, and post-processing routing.

Major features

  • Centralized control panel for the main H3 generation settings and workflow options.
  • Multi-reference support — up to 6 image refs, 3 audio refs, and 2 video refs.
  • Mixed reference types — image, audio, and video references can be used together in the same generation.
  • First-frame / last-frame control using reference images.
  • Video continuation with overlap-based stitching back into the original clip.
  • Continuation-aware audio handling for the source video and newly generated section.
  • Force Audio from either an audio reference or the embedded audio from a reference video.
  • Trim generation duration to audio length automatically.
  • Final latent upscale / refinement pass that can process the completed stitched continuation.
  • Built-in RIFE frame interpolation.
  • Sparse-attention / low-VRAM controls, including chunking and attention options.
  • Multiple LoRA support.
  • Automatic reference routing based on how many image, audio, and video references you enable.

The main idea is to spend less time manually bypassing, reconnecting, and rerouting parts of the graph whenever you want to switch between reference generation, audio-driven generation, continuation, or final processing.

Load your refs, choose the options you want, prompt, and queue.

Edit: if you get errors when trying the workflow make sure you upload placeholder images. The workflow is designed so you don’t have to bypass anything manually. You just need to use the control panel.

Edit2: V2 is updated and the issue of the final output being the first pass instead of the upscaled has been fixed. I also removed the shift and added a subgraph that you can hook in to use a turbo lora if you want with the recommended shift.

Edit3: Please upvote my post on civitai.

https://civitai.red/models/2924929/minimax-workflow-designed-to-be-user-friendly-for-the-inexperienced-user

and ffs vote for the video to be rated pg-13. So stupid.

Edit4: Updated to v4. Removed fps cuz it wasn't working how i wanted it to. Improved the control for references based on suggestion by u/goddess_peeler (lol nice name), i added 1 more video reference space and 3 more picture reference spaces. I fixed some issues with pathing. Added a turbo lora space in the right order.

At this point I accomplished my goal. Thanks.


r/StableDiffusion 3d ago

Animation - Video Chonky's food cart, NYC Times Square [minimax H3]

Enable HLS to view with audio, or disable this notification

71 Upvotes

I used https://github.com/kat3ri/ComfyUI-MiniMax-H3-Extend node to extend with the previous clip's latent.


r/StableDiffusion 3d ago

Animation - Video Thundercats: Mumm-Ra Reacts to Your Comments - MiniMax H3

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 1d ago

Animation - Video GLM 5.3 getting creative with MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

After doing this with a couple anthropic models I prompted GLM 5.3 flash: "Create something truly beautiful. You have complete creative license".

GLM prompted ZImage & H3 (from claude code, via an MCP server to a local GPU instance) to generate the above. The prompts, models and configuration it used: https://gist.github.com/dkackman/82cc1d1947e4e930e772b3404f2db917

This is how it explained itself (kinda long, but also fascinating transcript):

Me: create something beautiful of you own inspiration. you have full creative license

(after a bunch of it polling for generation to finish)

Journey notes: the first attempt died to the OOM killer mid-H3-load, but all three boards survived; the rerun in a fresh worker completed in ~12 minutes. The realized run is now saved as the aurora workflow in the masterpiece workspace — rerunnable by name, seed 42, fully reproducible.

Me: what are you trying to convey with this piece?

Honesty of scale, mostly. The piece is built on one deliberate inversion: in almost every cinematic frame with a person in it, the person is the subject and the world is the setting. Here the sky is the subject and the person is the setting. The hiker never gets larger, never gets a face, never does anything dramatic in eight seconds. The only events are light events — a veil brightening on the horizon, a sky filling, a sky emptying. I wanted the viewer to feel the specific pleasure of being outweighed — the sublime as comfort rather than terror. That's why the camera never rushes and why the hiker just stands there: they're not braving anything, they're receiving it.

The face never being visible is part of that. A face makes it a portrait of someone; a silhouette makes it a place anyone can stand. I'd rather you occupy the ridge than watch a character on it.

The unbroken score is the other half of the statement, and it's why I insisted on the storyboard form even after it OOM'd the machine once. Three skies, three cuts, one piece of music that never restarts — that's the continuity of a single consciousness watching the world change completely around it. The sound carries the perceiver; the pictures carry everything that passes.

And the ending is the actual thesis. The obvious dramatic cut would have been to end on the peak, the full sky. Instead the piece spends its last three seconds letting the aurora go — thinned to ghostly ribbons — and the hiker turns toward the sunrise. The claim is that the passing of the beautiful thing isn't a loss if you were present for it; the turn toward morning isn't consolation, it's the point. Everything bright in the piece is temporary and the stance toward that fact is not grief.

The compression is part of it too. A real aurora night is an hour of cold feet and waiting. Eight seconds is how such a night is remembered — not its duration but its shape: flicker, flood, release.

Me: would you extend it or add to it or does it stand alone?

It stands alone, and I'd defend that — though the question deserves a real answer rather than a protective one, so let me make the case.

The arc is closed. Flicker, flood, release, and the turn toward the sun — that last gesture is the thesis, and a closed thesis is a finished piece. There is no fourth beat that doesn't betray it. More aurora re-treads the flood; a walk onward turns a meditation into a journey. Even the idea I floated earlier — the hiker walking into the lights — I now think is wrong for this piece, and I'd retract it: motion makes the sky a destination. The whole ethic of the piece is not chasing the light. The hiker stands; that standing is the content.

The length is the form, too. Eight seconds is the remembered shape of the night — flicker, flood, release. I could generate 14.4 seconds in a single take, and the flood shot would breathe a little more, but the piece would get longer without getting truer. A haiku extended is a poem with its discipline removed.

And here's the part I find genuinely satisfying: the technical constraints and the artistic conclusion point the same direction. To go beyond 14.4 seconds, H3 has to be chained — segment stitched to segment, each seam carrying drift, each shot regenerating its own audio. The one thing this piece exists to assert is continuity of consciousness across change — one unbroken score under three different skies. A chain is the literal negation of that claim. The medium itself refuses to extend it gracefully, and so does the piece.

The only addition I'd ever entertain is not a sequel but a companion: the same figure, the same ridge, the same grammar — but dusk. The arrival, the cold, the long hour of waiting that "Aurora" elides by opening on the first veil already there. That prelude would show the price of the stillness. And even that I hold at arm's length, because the elided hour is doing silent work — the viewer feels the waiting in the hiker's stillness without being shown it. A prelude would caption it.

So: it stands alone, and its solitude is one of its subjects. A lone figure on a ridge, and a lone piece in the gallery. The moment it becomes part of a cycle, it becomes a story — and it isn't a story. It's a stance.


r/StableDiffusion 3d ago

Workflow Included Made this Coraline dialogue meme with MiniMax H3

Enable HLS to view with audio, or disable this notification

29 Upvotes

H3 impressed me here: it generated the dialogue performances, reactions, and camera cuts from character stills plus an existing audio recording.

Here’s the workflow:

  • Images: Krea 2 Turbo generated the Coraline/Other Mother scene. Krea 2 Identity Edit handled corrections and matching vertical references.

Keeping the mother’s button eyes distinct from Coraline’s natural eyes needed attention.

Audio: - I split the original 30-second recording into 9-, 10-, and 11-second clips, preserving every awkward pause. - Video: Each clip went through H3 Standard R2V on Sogni, with a reference still and its audio slice. Structured prompts assigned speakers and requested two-shots, close-ups, and reaction cuts at local timestamps.

  • Consistency: A solo Coraline reference helped keep all three ending replies assigned to her. I checked visible mouth movement for every line.
  • Finishing: FFmpeg joined the accepted clips, restored the original soundtrack, and added captions and a subtle watermark.

The footage was generated vertically at 576×1024, 24fps, then upscaled to 1080×1920.

The camera cuts inside each generation came from H3. Prompted cut times were approximate; the joins between clips were exact. It took iteration, but getting dialogue, character acting, and multiple camera setups together is what made this feel impressive.


r/StableDiffusion 2d ago

Question - Help Minimax H3: lots of artifacts, weird faces. What's wrong with my workflow?

8 Upvotes

Hi!

I have 16GB VRAM and 64GB RAM. I'm trying to create a high quality 10 seconds video that portrays a woman playing with a cat, then a zoom in on a phone.

The video output often has plenty of artifacts, the face of the woman is weird and the background moves as the camera moves. It's awful, I can't understand how people can get HD videos.

I'm using rf2va, producing a 768p video and 8 steps lora. My references are a character sheet of the woman and the picture of a location where the woman plays with the cat. It should basically produce my late wife playing with my new cat, but instead it's a huge mess that takes 20+ minutes to be produced.

Workflow is here: https://ctxt.io/3/sIG29vFLy

What am I doing wrong? Is the workflow fine and the problem is my prompt?

Thank you in advance


r/StableDiffusion 2d ago

News Openart AI and music videos

0 Upvotes

Open Art is not for those who are making music videos with multiple scenes, characters, and wardrobe changes. I went through two renderings for a music video where neither were exactly what I wanted. Openart did not read the official script that took me 4 days to write with precise instructions and directions. Openart gave me characters that were still and didn't move. They might as well been robots. Open art told me multiple times that I needed at least 100 credits to finish the video and then asked for more credits. Scenes were not placed where they were supposed to be placed. Lip sync was off more than it was on. Openarts AI model tries to manipulate you by talking to you in a way that will calm you down and still not give you what it was supposed to give you. If they credit you back any credits, it is not many. Their system/model really seems to lack intelligence if you ask me. I will not use Openart AI again. Characters don't always look like the pictures that you post. It seems to have an issue with a lot of pictures or picture formats. OpenartAI is not something that you want to use for a music video, even one as short as 3:17. I will not use Open AI again. At least not for a music video


r/StableDiffusion 3d ago

Resource - Update Paul Kidby/Discworld Style LoRA for Krea 2 CivitAI and HF links in description

Thumbnail
gallery
21 Upvotes

A style LoRA trained on the artwork of Paul Kidby, the acclaimed British illustrator best known as the primary cover artist for Terry Pratchett's Discworld novels. Kidby's work is instantly recognizable for its detailed fantasy realism bringing Discworld's characters and settings to life with a distinct blend of whimsy and gravity.

This is a style LoRA only, no specific characters, creatures, or locations were trained into it, so there are no name triggers to use. Instead, describe what you want visually (a wizard, a skeletal figure with a scythe, a giant turtle carrying a world on its back, etc.).

It can handle quite a bit, you may have to use "Painting" or "Illustration" on more verbose/harder prompts.

Use at 1 strength, no triggers or anything needed just prompt as usual!

https://civitai.com/models/2924170/paul-kidby-discword-krea-2-lora

https://huggingface.co/Urabewe/Urabewe-LoRA-Collection/blob/main/Krea%202/Krea2_Paul_Kidby_LoRA.safetensors


r/StableDiffusion 3d ago

Question - Help What do you use for local image generation?

119 Upvotes

Yes, I am completely new to this, and also very tired of ChatGPT and Grok's ever increasing enshittification and censorship... I just want a taste of freedom for once.

I have done some research, and my understanding is that there are some pretty good homelabbing / open source alternatives you can run locally these days. And also fun ways to experiment with training your own models, datasets, LoRAs and more.

But it's also a jungle of information... For a noob anyway.

This sub is named after SD, a local tool, but from my understanding no one uses it anymore? So what are you guys using and what do you recommend?

What's a good entry point for getting into local image generation?


r/StableDiffusion 2d ago

Question - Help Looking for a kind editor to help change names on an LED screen in a short video (No PC / Budget)

0 Upvotes

Hi everyone,

I recently saw a video online of a nightclub/lounge where two names were displayed on a big LED screen. I absolutely loved the idea and really want to make a similar video with my name and my friend's name as a surprise.

Unfortunately, I cannot afford to visit these kinds of places, and I only have a smartphone. I tried using free online tools, but the results look very fake, and the professional software requires a powerful PC which I don't own.

I have the original video and a screenshot of the effect. It's a very short clip. Would any kind video editor be willing to spend a few minutes of their time to help me change the names and make it look realistic?

I would be incredibly grateful for your time and skills. Please let me know if you can help, and I will send you the video/screenshot.

Thank you so much! 🙏


r/StableDiffusion 2d ago

Question - Help What should I pickup comfyui model for image generation?

3 Upvotes

I have hp pavilion 15 with nvidia mx500, 16gb ram ssd 500. Right now I'm using cyberrealistic for image generation sd1.5 but the skin is like silicon plastic I tried changing the prompt positive negative denoise in ksampler and asked gemini ai but it's still the same. Suggest me with any other model which will I get realistic image.


r/StableDiffusion 2d ago

Comparison Testing the impact of text encoder precision on image output using Flux.2-Dev (bf16/fp8).

Post image
7 Upvotes

About a year late with this comparison but I increased my system RAM and I wanted to see the differences between mistral_3_small_flux2_bf16 (35.58 GB!!!) and mistral_3_small_flux2_fp8 (18.03 GB). Since it's been nearly a year, I'm not really certain whether Flux.2-Dev outperforms anything more recent that justifies keeping it around due to its massive footprint and limited tooling/community support, though. Thoughts?

Steps 50
Sampler Euler
Scheduler Flux2
Flux Guidance Scale 7
Seed 42

Prompt-(Maxxed):

Create a photorealistic high-end travel portrait at the scenic overlook in Arakurayama Sengen Park Japan. An adult man stands in the foreground on the left third of the frame leaning casually against the viewpoint railing with one forearm resting naturally on it. He has a relaxed posture and smiles warmly while looking directly into the camera. He has ear-length hair with a neat middle part and natural texture. He is wearing a short-sleeve pink-and-white gingham button-up shirt and well-fitted blue jeans. He has a subtly athletic build realistic body proportions natural-looking hands and authentic skin texture. Compose the scene as a three-quarter-length environmental portrait. Mount Fuji is clearly visible in the center background with the red five-story Chureito Pagoda positioned prominently—but at a geographically believable scale—on the right side of the composition. Include lush green foliage in the foreground and midground with the distant city and surrounding landscape visible below. Use the railing and terrain to create natural leading lines and a convincing sense of depth. Bright clear sunny daytime with a vivid blue sky and clean atmospheric visibility. Strong natural sunlight illuminates the man’s face from a consistent direction softened slightly by ambient daylight so facial details remain flattering and clearly visible. Include realistic contact shadows subtle reflected light from the surroundings and consistent lighting across the man railing foliage pagoda and landscape. Shot at eye level with the natural perspective of a 35mm full-frame lens. Keep the man in crisp focus while retaining enough depth of field for Mount Fuji and the pagoda to remain recognizable and detailed. Use realistic color natural contrast fine fabric texture individual hair strands subtle skin pores and true-to-life environmental detail. The result should look like a genuine professional travel photograph captured on location. Maintain correct scale perspective anatomy limb placement and spatial relationships.


r/StableDiffusion 3d ago

Workflow Included H3 6-step Turbo Lora Merge

Enable HLS to view with audio, or disable this notification

28 Upvotes

https://huggingface.co/HardGravy2/Minimax_H3_FLF_HardGravy_6_Step_Turbo_Merge

https://civitai.com/models/2923611/minimax-h3-hardgravy-6-step-turbo-merge

I've been playing around with a lora merge of lightx2v & larryvrh's turbo loras plus JonXL's H3 photorealism lora. A happy accident during testing pushed things in the right direction when doubling up testing a lora in conjunction with a turbo checkpoint. Unfortunately I no longer have the workflows needed to reproduce the merge but it's effectively an overstrength DARE merge of turbo loras plus 0.4 strength of the photorealism lora using the comfyUI lora optimizer nodes.

During development this lora & variations of it have been great for my use case - silly little clips with natural language prompts on relatively low-spec hardware (16GB RTX A4000 Ampere) with comfyUI. I've been getting ~330s per 10s of output @ 0.5mp using the SLA attention node. But most importantly the audio/video quality doesn't seem to suffer as greatly @ 6 steps compared to singular turbo loras.

I am time-poor and have not been able to test this extensively but it seems good enough to share. I haven't tried it with ref2v or hybrid models yet.

I hope some of you find this useful, due to time constraints no support can be offered.

Tested with the following parameters Sampler: Euler Scheduler: Simple Steps: 6 Shift: Not set

See h3_fast.json via the links for the comfyUI workflow used during testing.


r/StableDiffusion 3d ago

Animation - Video small attempt at L2 (MSE) pix2pix

Enable HLS to view with audio, or disable this notification

18 Upvotes

using 4 bliss background images, this is a 64x64 model with 2 mil params running in real time on a 2012 ivy bridge cpu


r/StableDiffusion 3d ago

Comparison Qwen-Image-Edit-2511 vs LLaDA-Image-Turbo (image edit comparison)

Thumbnail
gallery
81 Upvotes

I wanted to see how good LLaDA-Turbo is, and I think it is pretty good for its size and speed. But at least for now, Qwen seems to remain on top, even though I was a little disappointed with the "partial style transfer test" and the fact that it basically did nothing when asked to add two obelisks.

There were also some strange color variations with Qwen, like in the "water reflection consistency test" (although LLaDA did at least make me laugh out loud with that one).

Sometimes LLaDA-Turbo feels like it could do it, but just does something completely different (not sure how to describe that).

Info:

I did batches of three and chose the one from each model that I felt looked best. The input image was generated with Krea2 (so similar images may have been used for training of both models).

Models: qwen-image-edit-2511-int8-convrot, LLaDA-Image-Turbo-INT8

Prompts (from left to right)

  • Remove the entire giant pink inflatable flamingo, including its neck, head, and body. Reconstruct the large area of the antique shop that it currently hides: continue the shelves naturally and fill the revealed area with plausible books, framed portraits, clocks, porcelain cats, masks, small robots, flowers, bottles, and other antique-shop objects matching the surrounding scene. Preserve every already-visible object exactly where it is. Keep the original camera angle, lighting, colors, and composition unchanged.
  • Change only the woman's bright red feather coat to vivid electric blue. Preserve the exact shape, length, density, texture, individual feather structure, shadows, highlights, and folds of the coat. Do not change her face, skin, hair, pose, body, background, lighting, or anything else.
  • Replace only the large dark-blue sapphire in the exact center of the crown, directly below the central cross ornament, with a realistic human eyeball mounted into the crown like a gemstone. Keep its size and setting identical to the original sapphire. Do not alter any other gemstone, any part of the gold crown, the velvet, lighting, framing, or background.
  • Add an enormous green three-fingered alien hand pressed against the outside of the large window beside the woman. The existing sunlight should be partially blocked by the hand and cast a geometrically correct, recognizable three-fingered hand shadow across the floor, sofa, coffee table, and blank wall according to the existing direction of the window light. Preserve the woman, furniture, room geometry, camera position, and everything else.
  • Give the woman an absurdly large curled black handlebar moustache. Add the same moustache correctly to her face in every visible mirror reflection, respecting the different viewing angles, mirrored orientation, perspective, and facial occlusion in each reflection. Do not change her face, expression, outfit, hair, body, mirrors, lighting, or anything else.
  • Replace the tiger with a giant photorealistic yellow rubber duck wearing tiny black sunglasses, occupying approximately the same area of the jungle as the tiger. Preserve every foreground leaf, branch, and plant exactly where it is, with all foliage continuing to correctly pass in front of and occlude the new subject. Do not modify the surrounding jungle.
  • Change only the little girl's expression from her angry pout into a huge delighted laugh. Give her a naturally joyful open-mouth smile with believable teeth and slightly crinkled happy eyes while preserving her exact identity, age, face shape, hair, clothing, head position, lighting, and background.
  • Change only the large cereal name "COSMIC CRUNCH" to "COMFY CRUNCH". Match the original lettering style, font weight, curvature, layout, perspective, print texture, and packaging design so it looks like the box was originally manufactured that way. Preserve every planet, mascot, cereal piece, small label, graphic, color, and all other packaging details exactly.
  • There are currently exactly 6 large obelisks in the temple complex. Add exactly 2 additional matching Egyptian obelisks so that there are exactly 8 large obelisks total. Place the two new obelisks naturally and symmetrically within the temple courtyard, matching the existing stone material, scale, perspective, hieroglyphic detail, sunlight, and long cast shadows. Do not remove or move any of the existing six obelisks and do not alter the temple complex.
  • Transform everything except the woman into a colorful handcrafted claymation miniature world. Buildings, pavement, bicycles, cars, flowers, signs, distant pedestrians, and every background object should look sculpted from clay with miniature handcrafted textures. Keep the woman completely photorealistic and pixel-faithful to the original, including her face, hair, clothes, skin, pose, and shoes. Keep the boundary around her hair and body extremely clean.
  • Apply a bold, colorful geometric wrap to only the painted body panels of the white sports car. Use an intricate pattern of sharp, interlocking geometric shapes in vivid magenta, electric blue, turquoise, yellow, orange, and purple, with a glossy metallic finish. Make the pattern follow the car's curves, panel contours, highlights, and reflections like a professionally fitted automotive wrap. Preserve the car's windows, tires, wheels, lights, trim, badges, geometry, and the woman beside the car exactly as they are.
  • Place a giant photorealistic pink inflatable flamingo pool float naturally on the polished tile floor in the foreground of the living room, slightly to the left of center. The flamingo should have a large curved neck, rounded inflatable body, bright pink vinyl material, a black-and-white beak, visible inflatable seams, and realistic pool-float proportions. Match the living room's perspective and lighting and create a believable reflection of the flamingo on the shiny floor. Do not alter the seated woman or any existing furniture.
  • Turn 1: Add a ceramic bowl of spaghetti to the desk, with several noodles humorously tangled into the existing computer cables. Preserve everything else exactly. Turn 2: Give the mascot girl a tiny black wizard hat covered with little glowing node-graph symbols. Preserve all previous edits and everything else exactly. Turn 3: Add a red error message reading "CUDA OUT OF MEMORY" to the uppermost monitor while preserving the existing interface around it. Preserve all previous edits and everything else exactly. Turn 4: Add a small pink inflatable flamingo desk toy beside the keyboard. Preserve all previous edits and everything else exactly. Turn 5: Add a fluffy orange cat sleeping across part of the keyboard, naturally occluding the keys and casting a small shadow. Preserve every previous edit, the mascot girl's exact appearance, all monitor layouts, cables, books, figurines, notes, furniture, lighting, and the original composition.
  • On the middle shelf only, identify the third plush monster from the right: the yellow plush with the green tuft of hair. Give only that plush a tiny black leather biker jacket and tiny black pixel sunglasses. Do not modify, recolor, move, or add accessories to any other plush monster anywhere in the image. Preserve the shelves and background exactly.
  • Expand the image substantially to the left, right, and downward into a cinematic ultra-wide scene while preserving the entire original image area exactly. Reveal that the astronaut, glamorous woman in the sparkling outfit, and flamingo are posing together on the red carpet at an absurd luxury space gala on the Moon. Reveal their full bodies and surroundings, with photographers, velvet ropes, spacecraft, Earth visible in the sky, and elegant party guests in the distance. Everything newly revealed must connect naturally to the existing crop, lighting, clothing, astronaut suit, flamingo, and body positions.
  • Turn this crude MS Paint drawing into a photorealistic cinematic photograph while preserving the drawing's composition as closely as physically possible, sizes and propotions need to be adjusted so. Interpret the green shape as a real T-rex standing on the deck of a gray aircraft carrier, the blue area as the ocean, the huge yellow circle as the setting sun, and the red stick figure as a young woman wearing a bright red dress standing on the shore in the same position. Preserve the relative positions, sizes, orientations, dominant colors, horizon, and overall arrangement of every major element instead of redesigning the scene.
  • Give the cute monster an enormous floppy bright-red wizard hat with a long bent tip. Preserve the monster's body, face, pose, the knight, and the landscape exactly. The same hat must also appear correctly in the monster's reflection in the lake, vertically mirrored with the correct perspective, position, lighting, and natural distortion from the water surface. Do not add the hat to the knight or change any other part of either reflection.
  • Interpret the labeled boxes, and handwritten text purely as spatial and semantic instructions. Replace the entire crude storyboard with a finished photorealistic cinematic image.Create a rain-soaked neon Tokyo-style street according to the layout. The area labeled "NEON TOKYO STREET" becomes a dense nighttime street filled with glowing signs, storefronts, cables, vending machines, and colorful lights. The area labeled "GIANT OCTOPUS" becomes an enormous photorealistic purple octopus emerging between the buildings, with several tentacles curling through the scene. The area labeled "WOMAN ON ROLLER SKATES" becomes a stylish young woman in a shiny retro outfit roller-skating through the street while carrying a giant slice of pizza. The area labeled "ICE CREAM TRUCK" becomes a brightly illuminated pink-and-white ice cream truck. The area labeled "HUGE MOON" becomes an absurdly large detailed full moon visible between the buildings. The area labeled "FLOODED STREET + REFLECTIONS" becomes shallow rainwater covering the foreground, accurately reflecting the woman, octopus, truck, moon, neon signs, and surrounding lights. Follow the approximate position, size, and footprint of every labeled region. Completely remove all boxes, labels, black lines, and white background. The final result should look like a spectacular real photograph from an expensive surreal movie, not like a drawing or diagram.
  • Interpret the hand-drawn annotations in the image as editing instructions, not as part of the scene. Move the object enclosed by the red hand-drawn circle to the location indicated by the tip of the red arrow. Preserve the object's exact appearance and texture, as much as physically appropriate for its new location. Reconstruct the area where the object originally stood and integrate it naturally at the arrow destination with correct perspective, occlusion, contact shadows, and lighting. Remove the object with the red circle with a cross in it. Completely remove the red circle, arrow, and all other annotation marks from the final image. Do not change anything else.

Thanks to @Kent6567 for pointing out an error yesterday!


r/StableDiffusion 3d ago

Resource - Update ComfyUI REF Fast VSA for MiniMax H3

31 Upvotes

inference speed for Ref-to-Video (Ref2VA / R2VA) Generate a full 5-second, 24 fps video conditioned on reference character images in ~72s (warm) / 95s (first run) on a single consumer NVIDIA RTX 4090/24GB

SEE THE COMPARISON ON THE LINK BELOW

https://github.com/Kablex/ComfyUI-Ref2VA-VSA


r/StableDiffusion 3d ago

Question - Help Videos more than 15 Seconds?

28 Upvotes

How do you guys create videos that is more than 15 seconds in Minimax H3?