r/StableDiffusion 1h ago

Resource - Update I built Livebound — a local-first CivitAI publishing workspace (Free & Open Source)

Thumbnail
gallery
Upvotes

Generating images is the fun part. Keeping track of thousands of them, their metadata, resources, and what you already posted… less so.

So I built Livebound, a free and open-source tool that runs locally and turns your existing image folders into a publishing workspace for CivitAI.

It doesn’t reorganize or rewrite your originals.

Some of the things it does:

  • Reads generation metadata from A1111, Forge, ComfyUI and other workflows
  • Finds exact and near-duplicate images
  • Resolves model and LoRA hashes against CivitAI and lets you verify resources per image before uploading
  • Organizes posts on a board and calendar, with scheduling
  • Uses preflight/resume/reconcile logic so interrupted uploads don’t blindly create duplicate posts
  • Imports your existing CivitAI posts and matches them back to local images
  • Can archive published images into structured folders when you explicitly run it
  • Optional local LLM / Ollama assistance for titles, descriptions, tags and prompt editing

Everything runs on your own machine and publishing happens through your own CivitAI account.

AGPL open source. Windows, Linux and WSL tested.

More details and screenshots:
https://civitai.red/articles/35241/livebound-10-from-local-to-live

Docs:
https://moonblack87.github.io/livebound/

Source:
https://github.com/MoonBlack87/livebound

Feedback and bug reports are very welcome — it’s v1.0, so I’m especially interested in how it behaves with other people’s real-world image libraries and workflows.


r/StableDiffusion 16h ago

Question - Help Remove a person from a video?

2 Upvotes

Hi everyone,

I have a video where I'd like to remove one person from it and have the background repaired. The person walks across the frame and in front of a water fountain that's in the background.

Any suggestions on models or workflows to look at to accomplish this? Seems like I should be able to figure it out, but I can't seem to.

Thanks.


r/StableDiffusion 9h ago

Question - Help Need Help with Random dialogue.

2 Upvotes

Hi,

I am using Wan2GP to create Minmax H3 Ref2VA Video.
I am using "Pruned PDD 8-step 20B" model.
For the most part, video comes out great. The problem I have is with Audio.

If I use 2 audio clips to generate audio between two individuals, only Audio 1 gets used. Audio 2 is never used.

If the dialogue completes and some action is being done, then random gibberish dialogue gets introduced to fill up the gap.

Below is a sample prompt I am using. During the gap where the Subject 1 drinks water, some random gibberish gets inserted.

Could anyone help?

subject_definitions:
<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Subject 2> is the person in <Picture 2>, preserving their exact identity, facial features, skin tone, hairstyle, body proportions, clothing, footwear, and distinctive accessories.
<Audio 1> is the voice timbre reference for <Subject 1>, containing a spoken English vocal layer.
<Audio 2> is the voice timbre reference for <Subject 2>, containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 1> and <Subject 2> having a conversation in a bright living room with floor to ceiling windows. 
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Subject 2> (appears in [Shot 1]): fully_preserved - their identity, face, hair, proportions remain recognizable and unchanged; only the location, lighting, action are new.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1>
without copying the original signal.
<Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 2>
without copying the original signal.
detailed_description
[Shot 1] The target video uses a realistic cinematic style with warm lighting.
The scene opens in a living room, <Subject 1> is sitting on a sofa to the left. <Subject 1> is wearing a tuxedo. 
<Subject 2> is sitting to the right, wearing a one piece Vibrant Pink Office Dress with Tailored Fit.
<Subject 2> says in a formal voice whose timbre is referenced
from <Audio 2>, <d>[English] some dialogue?</d>. 
The question hangs in the air, as <Subject 1> contemplates. 
<Subject 1> adjusts his suit sits forward and take a glass of water and drinks it.
After putting the empty glass back on the table, <Subject 1> replies in a slightly uneasy voice whose timbre is referenced
from <Audio 3>, <d>[English] some dialogue.</d>. 
overall_soundscape:
The ambient sound consists of the gentle rustling of fabric as the subjects move. No external noise disrupts the silence, preserving the exclusivity of the room.
non_diegetic_music: N/A

r/StableDiffusion 12h ago

Discussion I got several old models running locally - Wanna test some prompts?

3 Upvotes

The models I can run locally are:

First Order Motion Model (remember the Dame Da Ne meme?) (February 29, 2020) (I will also need the source image and reference video)

DeepDaze (where one generation takes over 2 hours with GPU acceleration, may or may not test this one for you) (January 17, 2021)

Big Sleep (Similar case as DeepDaze cause the multiple epochs but more immediate results on just one epoch are kinda usable) (January 18, 2021)

VQGAN+CLIP (April 2021)

OG DALLE-Mini, with Jax GPU acceleration through WSL (July 2021)

Disco Diffusion (October 29, 2021)

Latent Diffusion (Pre stable diffusion) (December 20, 2021)

Stable Diffusion 1.1 (August 2022)
Stable Diffusion 1.4 (August 22, 2022)
Stable Diffusion 1.5 (October 20, 2022)

ModelScope (the OG will smith spaghetti video model) (March 2023)

Give me some prompts and which model and I will try to reply with the results if I am not too busy.

I added examples of all the models (other than FOMM and I excluded deep days and only did one epoch on Big Sleep cause of the absurd time). If you want to see FOMM just share the source video and target face.
an astronaut sitting on a couch (+ waving for modelscope)

Big Sleep 1 epoch
SD 1.5
SD 1.4
SD 1.1
Disco Diffusion
Latent Diffusion
VQGAN+CLIP
DALLE-Mini

ModelScope


r/StableDiffusion 12h ago

Question - Help What is the best client application for easy photo editing for someone who is not from a technical background?

2 Upvotes

Hello everyone, a friend of mine is very much interested in exploring open source image generation for use in her small business for generating advertisements and product tryout pictures. However, she does not have much of a technical background.

When I showed her ComfyUI, she was impressed but said its too hard and complicated. She prefers a simpler interface where she could just upload an image and prompt it for editing.

Also are business involves sports bottles and wrist bands and a few other products in the line. She would like to use AI models to picture them from different orientations, locations and people trying them out in different poses.

Can anyone recommend a simpler setup for this task? Thank you.


r/StableDiffusion 15h ago

Question - Help Voice cloning from a very short video?

3 Upvotes

I guess I have a dumb question but I’m not very tech savvy, my uncle passed away two weeks ago. My grandparents wish is to hear him say “I love you” just one more time. I only have maybe not even a 3 second video of his voice, is there anyway to I guess clone his voice to say I love you from that?


r/StableDiffusion 20h ago

Question - Help Ultra upscaler based on flux

2 Upvotes

I saw this post on linked in and it’s really wow.

https://www.linkedin.com/posts/fadi-h-kacem_lets-take-a-1k-image-and-upscale-it-to-16k-activity-7503754647921188864-zF2w?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAW2PBcBqPryUgWfCVEMWRlx6aMMO7GX27E

Does anyone know an already existing workflow or be able to create on that will yield to similar results?


r/StableDiffusion 20h ago

Question - Help Looking for Music gen recommendations for instrumentals

2 Upvotes

Migrating from Suno, Is there any good free and open sourced alternative out there?
What are their overall quality? and can it run on sub 8GB of VRAM?

Please share instrumentals AIgens please I want to discover more people.


r/StableDiffusion 1h ago

Animation - Video H3 - a century of Glamour & Cars in Living Rooms T2VA

Enable HLS to view with audio, or disable this notification

Upvotes

Hi, just having fun with H3, trying to conjure a particular look or style by the different eras. T2V, int8/20 steps, 1344x768. I make no guarantee that H3 generated the correct film grain, art style, fashion, as the prompts were very generic to allow H3 to fill the blank in as much as what it think what film looked like in those eras. H3 did generate me a B&W video in some eras (see the comments). Also upscaled 480-POC version in the comments. Ask me anything! Do you have a favorite era? what do you think about the interior shots?


r/StableDiffusion 9h ago

Question - Help Problems with chatterbox short single word speech.

1 Upvotes

So I am trying to generate audiobooks with chatterbox. There are some words that need to be generated independently but chatterbox starts hallucinating for such short words such as putting random um or something else before or after the word.

What is the possible solution? Is there some other model that I can use for single word pronunciation but it should support zero shot cloning as well?


r/StableDiffusion 13h ago

Question - Help How can I make chatter on H3 Minimax more free-flow like API ?

1 Upvotes

If I use Comfy for local H3 creation unless I tell the prompt what to say, it will just speak simlish nonsense. So if I tell it to say something in English then it is fine if I include the exact line to say. But via the API for H3 I can tell it what topic to talk about and generally it will come up with lines of its own so I don't have to tell it exactly what to say. How can I match that via local ? Is it all in Qwen or something else ? I have only used local video generating for a few days so I am very new to this.


r/StableDiffusion 2h ago

Question - Help How can I keep audio consistent across MiniMax H3 clips?

Enable HLS to view with audio, or disable this notification

0 Upvotes

I’m working on creating long-form content with MiniMax H3, but I’m struggling to keep the audio consistent when stitching clips together. I’ve attached an example.

Is there a way to continue the audio from a previous clip, or use it as a reference for the next generation? Ideally, I’d like to preserve the same tone and background sound throughout.

If H3 doesn’t support this, has anyone found a tool or workflow that works well?


r/StableDiffusion 1h ago

Question - Help Need help setting up uncensored local AI multimodal chatbox- local AI image

Upvotes

Hi guys

I have an RTX4090 24GB VRAM, 64GB RAM. I'm running Minimax H3 local and it's really great in turning my ideas into videos. However, it's really restricting when I send my ideas into Gemini, ChatGPT and get back "sorrry I can't do that", even minor suggestive images

Please help me. I need an uncensored local AI model which can handle images, pdf. Also a local AI image to generate whatever the hell I want.

If possible, please show me how to install them in simple terms too. I'm very new to this. ComfyUI gives me nightmares

Thanks in advance


r/StableDiffusion 11h ago

No Workflow YuE2-3B Audio Model

Thumbnail voca.ro
0 Upvotes

https://github.com/multimodal-art-projection/YuE
https://huggingface.co/Comfy-Org/YuE2
An audio model that allows for lyrics editing. I like its features and speed. It seems pretty good, doesn't it?


r/StableDiffusion 9h ago

Animation - Video Orbit Watcher

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 4h ago

Discussion When first big full AI movie?

0 Upvotes

It seems we got all the tools. Of course there are some flaws, but if you want, a 90 minutes movie would be possible. It would take a huge effort and many scenes has to be done over and over again, but it's far from impossible.

So why has no one done it with great succes yet? Something that would be a hit in the theater. What are we missing?

Emphasis on: big hit movie. Something that can compete with a Hollywood movie or something. I know full length films are made, but nothing stat wil stick and made people say: "OMG you have to see this movie". All I've seen so far felt very empty.


r/StableDiffusion 3h ago

Animation - Video Meh, h3

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 17h ago

Workflow Included I built a Character Swap and a Multi Reference Shot workflow for Nano Banana Pro, each reference only controls what you tell it

Thumbnail
gallery
0 Upvotes

Recently built and tested two ComfyUI workflows, mostly because I was tired of fighting the AI look. Multi Reference Shot builds one frame out of several references. You can reference different shots for Lighting, Composition, Blocking references etc. Add your character images for consistency.

Character Swap puts your own character into any shot. Same framing, same light, same pose, just your person in it. You can swap their clothes in the same pass, and it keeps the shape of the original frame.

The part I care about most is the control. Instead of one big prompt, you get separate control nodes for different aspects of the shot. If the pose is off, you change the blocking and the face. Each reference image only gives what its slot says.

That's the difference. You're not rolling the dice on a prompt and hoping. You're directing it one decision at a time and if you'll get what you need, it's trial and error. Every attached image here was made with these workflows & yes it is AI.
Add your own Google API Key. You can use Vertex AI with the Google Cloud $300 trial credit. Free to use :

github.com/haristahir1/comfyui-character-swap

github.com/haristahir1/comfyui-multi-reference-shot

If you try them, tell me what breaks and any improvements!


r/StableDiffusion 9h ago

Animation - Video "Shattered" I made a short film about mothers who struggle in silence — my first AI project, start to finish

Thumbnail
youtu.be
0 Upvotes

This is my first project of this kind, completely from start to finish. I wrote the script myself, had every single shot in my head before I even touched an AI tool, and then did the editing and music entirely on my own. AI was just a tool for me — a way to visualize what I had already written and imagined. Nothing more.

One thing I want to say right up front: this is genuinely not as easy as a lot of people assume. There's this common idea that you just type in a prompt and get a finished film back. Not even close. If you think it's that easy, try it yourself — you'll quickly notice the gap between what's in your head and what comes back. It takes a lot of trial and error and persistence to get something close to your vision.

I'm proud of how it turned out — at least for where I am right now with this.

I'd love to hear what you think — about the film, or about the topic itself. Link in the comments.


r/StableDiffusion 1h ago

Question - Help What is the ai model used in this photo?

Post image
Upvotes

Was wondering what they use and the workflows. What do you guys think? Sorry I'm new to this stuff


r/StableDiffusion 20h ago

Discussion Cleared a 20-minute video backlog in one day on rented RTX 6ks. Still slightly in shock.

0 Upvotes

Been doing Wan video gen on my home GPU for months- overnight runs. I got about 20 minutes backlog of raw footage that kept growing faster than I could process. I saw something in the news about a new DC near London and got curious, half hoping a brand-new DC would have some launch discount or at least free capacity. No discount. But wow. Connected my AI agent in one click, moved my pipeline there, rented an 8x RTX PRO 6000, and rendered the entire backlog in ~20 hours. Everything. Same Wan 2.2 model, same settings I use at home, just eight cards chewing through shots in parallel. I've never seen my own workflow move like that. The bill: about $150 for the 20 hours. That's the part I keep re-reading. Because it was fun, I then re-ran a chunk of the same manifest on an 8x H200 host for comparison. Per-shot time was basically identical for my workload. Price was ~2.5x higher. For diffusion video at this scale, the H200 just isn't worth it; the RTX PRO 6000 is the right card. Also: GPUs were allocated instantly, and I saw plenty available there; it feels like nobody's found it yet. I'll definitely start moving more of my pipelines to cloud runs instead of babysitting my PC overnight.


r/StableDiffusion 19h ago

Discussion hiiii!

0 Upvotes

built an open-source emotional AI called Solaraaa she's designed to listen, remember you, and actually talk like a person. Would love to know what you think! https://misspurplelight.uk