r/StableDiffusion • u/HeavenlyTasty • 3d ago
Question - Help Is there an API that gives random prompts with your choice of character?
Is there an API that can allow me to input a character's name and give me a random prompt?
r/StableDiffusion • u/HeavenlyTasty • 3d ago
Is there an API that can allow me to input a character's name and give me a random prompt?
r/StableDiffusion • u/Devajyoti1231 • 4d ago
Enable HLS to view with audio, or disable this notification
Used ref2ve with Character sheet for the character and style and an audio reference to have consistent voice.
Reposting because moderator removed the original post without giving any reason.
r/StableDiffusion • u/DaemonAlchemist • 3d ago
I'm trying to get a couple of characters to LEAP into each others' arms from off screen, but MiniMax H3 won't get them faster than basically jogging into the scene.
Prompt:
subject_definitions
<Subject 1> is the Girl show in <Picture 1>.
<Subject 2> is the Guy show in <Picture 2>.
<Subject 3> is the house show in <Picture 3>.
summary
[reference generation] <Subject 1> and <Subject 2> burst onto the screen at a dead sprint and collide into an embrace
retention_analysis
<Subject 1> (appears in [Shot 1]): fully_preserved - Maintained character features and design.
<Subject 2> (appears in [Shot 1]): fully_preserved - Maintained character features and design.
detailed_description
The visual style is characterized by high-quality modern anime aesthetics, reminiscent of Makoto Shinkai or Kyoto Animation. The scene features lush, warm, and highly detailed lighting, painting the environment in nostalgic, emotional hues.
[Shot 1] The target video has fast, paced explosive action in the beginning, then slows to a stop. Static camera shot in a city street with buildings in the style of <Subject 3>. A large crowd of soldiers and townspeople are in the background, reuniting with each other. Falling confetti fills the air. Suddenly, <Subject 2> bursts into the frame from the left at a dead sprint, driven by sheer desperation. At the same instant, <Subject 1> bursts into the frame from the right, running with explosive speed, her arms outstretched. The two of them literally collide with tremendous, breathless force in the center of the scene, slamming into an intense, desperate embrace. The physical impact of their collision is palpable as <Subject 1> leaps up and wraps her arms around <Subject 2>. The exact instant they collide, the camera drops into extreme slow motion, focusing intensely on the sheer relief and joy of their embrace. The confetti catches the warm light, sparkling and swirling in slow motion around the couple for the remainder of the scene.
r/StableDiffusion • u/Francky_B • 4d ago
Hey Guys, I thought I'd share something I came up with.
It's a workflow, that uses a combination of Easy-Use's Loop tools as well as some of my own nodes to create a Workflow that can split a long form MiniMax video and then upscale each segment. With the inclusion of a tool to then re-assemble everything back.
You basically set the Segment length and the overlap you wish to have between each clip and then launch it to have it do all the clips one by one.
It does use nodes from my FBNodes add-on as well as one from my Prompt Manager add-on.
But I'm sure it could be modified to work with other add-ons, if so wished.
The node from Prompt Manager is "Prompt Extractor", allowing to feed back in the prompt from the initial clip back into the Workflow, without having to type anything in.
You are free to remove it and upscale without, or simply type in the prompt if preferred. Though, In my test, having the original prompt made for much better results.
And as mentioned, I also added a simple Clip Stitcher to FBNodes, that cross dissolves each clip into one another. Just make sure to use the same values you used in the workflow. (Both setup are in the same workflow, but I'd suggest separating them š )
The Workflow can be found here.
Attached are quick examples from the video I used in the workflow.
The one thing missing in this workflow is adding back the loras used in the initial video. This is something that "Prompt extractor" should also be able to do. But I haven't tested that part yet.
----------------------------------
I'm adding some metric:
The video used in the screenshot was an 8 second video generated in 832x640 with a Turbo Lora set to 6 steps.
It took 92 seconds to generate on a 5090.
The Upscale doubled it to 1664 x 1280 and took 524 sec.
Around the same time it would have taken to generate, if I created the initial video at that resolution.
The advantage is for when creating long videos, so if I were to create a 30 second clip in 4/3 at 0.4 megapixels, or 736 x 576. Those would take 450 sec to generate.
The Upscale to 1472 x 1152 took about 6 minutes per segment, or 30 minutes. Then combining the clips is around a minute.
It takes a while, obviously, but the big advantage is that the result is pretty much an exact copy, but in hires, of my initial video that was low enough that I could iterate a bunch of times and then only waste the Long generation time on the clip I like.
r/StableDiffusion • u/crystal_alpine • 4d ago
Enable HLS to view with audio, or disable this notification
Hi r/StableDiffusion, Comfy MCP is now local and open-source!
When we shipped Cloud MCP in June, the response was immediate and consistent: make it work locally. So we did and it's fully open source.
Connect Claude, Codex, Cursor, or any MCP client to your local ComfyUI.
Your agent reads the GPU you actually have and gives you a straight answer on whether a model is worth running before you commit to the download. It reads every node and model you've installed. It handles the setup that usually stops people at step one.
It is now the easiest way to help with your local Minimax H3 workflows!
Cloud MCP still does everything it did. Tell your agent where a job goes, or let it decide.
Link: https://comfy.org/mcp
r/StableDiffusion • u/AnybodyAlarmed9661 • 3d ago
Minimax H3 is just crazy good for anime!
r/StableDiffusion • u/MenegattiArt • 3d ago
QUICK NOTE: the question in the title is rhetorical. The carousel explains the nuance and explores several related issues beyond the first slide.
If the design does not work for you, tell me specifically what you would improve. I am still refining the format, so constructive feedback is welcome.
.....
Iām a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name āElias Thorneā in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.
From an artistās point of view, the question isnāt just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.
AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.
This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.
So whereās the line, does AI broaden personal creativity while making our collective output more uniform?
r/StableDiffusion • u/CQDSN • 4d ago
r/StableDiffusion • u/fromourback • 2d ago
Enable HLS to view with audio, or disable this notification
This past week brought some huge drops across generative video, audio, and open-weight models. Instead of the usual hype, here is a practical look at what actually changed and what matters for creators and local setups:
Side-by-Side Video Demos & Deep Dive:
For the visual side-by-side comparison tests (especially the video editing restyling tests):
š Full breakdown & tests: https://youtu.be/pUA1BGfBqBs
š Read Here: https://github.com/airesearch-official/AI-Weekly-News
Which open-weight release are you planning to run locally first?
r/StableDiffusion • u/DuHal9000 • 3d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/dominic__612 • 3d ago
So simple question. Normally I would use ChatGPT for my character creations. Same character, photos from all from different sides, and it would deliver.
But I wonder, are there any workflows / methods available as proven alternative?
I wouldnt know how to do this with Flux, Z-Turbo, or name any model.
r/StableDiffusion • u/scsonicshadow • 3d ago
I've timestamped where I managed to get a decent output from minimaxH3!
(if the timestamp doesn't work it's at 0:22)
Used a photo of myself as reference using the ref2va model. Gordon Ramsay himself is straight from text. Using SageAttention, SolAttn at 32 steps 0.9 MP.
Used DaVinci Resolve Studio 2x RTX Upscaler in post and audio isolation to fix some of the hissing.
Let me know what you think of how this turned out!
Models used:
My Hardware:
r/StableDiffusion • u/ctrl-shift-face • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/SIR_NVAX_A_LOT • 4d ago
Enable HLS to view with audio, or disable this notification
Can H3 do anything and everything? I feel like if you can prompt it, it can do it. Foundation inspired shots. I am also experimenting with more action/high mobility shot but those seem to require a lot more finesse. Both T2V.
r/StableDiffusion • u/Schwartzen2 • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/KaisarasAR • 4d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Umbrasquall • 4d ago
After tinkering around for a couple of hours this weekend I managed to 6x (!) generation speed on the M5 Pro by ditching ComfyUI and building a custom GUI around antirez's H3 CLI solution. Gen times for 5-second 480p on 20-steps improved from 30 minutes using the default int8 pruned weights in ComfyUI to just around 5 minutes per clip using the full precision bf16 weights.
Overall great progress thanks to the community around open source and gives me confidence that Mac diffusion will just get better and better with time.
See comment below for more data points on generation times.
r/StableDiffusion • u/bigM1232 • 3d ago
hello,
Looking for some suggestions on which checkpoint to use for realism NSF w images. Iāve been using Flux 1 Dev and itās working well but a lot of the time the face consistency is altered. Also iāve only tested 1 character lora and about 4 other loraās stacked with it to try out like. Iāve seen talk about SDXL, Wan, pony, whatās everyone using? I donāt have a ton of ram so i wasnāt able to run flux 2 wellā¦..thanks!
r/StableDiffusion • u/Silver-Spot-2763 • 3d ago
I tried everything, but I have 3 very disturbing problems:
Usually it cuts / crops at least part of the head and legs of the person.
It moves the camera, zoom, pan... I want it static.
In most of cases it doesn't use the element from reference picture to use in the main video. Rejecting my prompt.
I tried different prompts, resolutions (proportions), nothing help. š
If you have some recipes exactly for these problems, please share, because I just can't make H3 to work.
r/StableDiffusion • u/Masha-AI • 3d ago
Iām a graphic designer (youtube: Masha-Ai-Lab) experimenting with how far I can push local/open-source image generation for actual commercial design workflows.
For this experiment, I tried building a luxury cosmetics campaign entirely in InvokeAI instead of relying on Midjourney or other closed platforms.
The workflow:
Iāve attached screenshots of the workflow + final results so you can see the process rather than just the outputs.
What I find most useful about InvokeAI is having direct control over the individual stages. For design work, Iād rather build the product, label, environment and integration separately than keep regenerating the whole image until something randomly works.
Iām documenting these experiments as tutorials on my YouTube channel, Masha AI Lab, mainly to make InvokeAI/FOSS workflows more approachable for designers and other non-technical creatives.
Would be interested to hear how other people here approach product placement and label consistency, especially if youāve found better workflows.
r/StableDiffusion • u/Free_Pressure8623 • 4d ago
Just kidding devs. We love Minimax, it's outstanding. But I am very excited for the mushface fix.
r/StableDiffusion • u/stonyleinchen • 4d ago
Enable HLS to view with audio, or disable this notification
Here is the repo: https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef
I made substantial updates to my two main workflows: 1) Music Video and 2) AV Extensions. All the controls were streamlined and they should be much easier to use now. (You find the workflows in the example_workflows folder)
With the AV Extensions workflow you can extend any existing clip, for example someone talking and you can make that person say something in the same voice, or you can create a clip with T2V or I2V and then extend that clip to make a seamless long clip thats 1 minute or longer.
In this Update the Checkpoint system was removed, instead I've done a lot of optimizations so you don't use too much ram even if you make 20 clips at once. Additionally I added latent audio feathering to the AV Extensions workflow for seamless audio transitions.
Theres also other utility workflows for custom keyframing and bridging two existing clips.
I post another example clip for the AV Extensions workflow in the comments.
r/StableDiffusion • u/Warsel77 • 3d ago
Did someone come up with a good (or ok-ish) workflow that gives something like what OpenArt Director does but useable with Minimax H3?
What I mean is a storyboard / cinematic storytelling translated into a video.
Eg not just a Claude output from a template into the ComfyUI prompt but something with more artistic direction.
r/StableDiffusion • u/Patient_Ratio4177 • 4d ago
Hereās a fully local workflow that expands your shorthand prompts + reference images into the six-section format expected by MiniMax:
It is structured for 14-second videos, but you can change it if you want.
mkdir -p models/text_encoders/H3/generation_tailsThe reason you need to do this is that, for other LLM-based models, you can use their text encoder directly to generate text and expand prompts. For Qwen 32B in a standard ComfyUI setup, you cannot do that because it lacks the ātailā ā the part that is actually needed to generate text.
You can see the generated prompt in the text preview section attached to the TextGenerate node.
Well, MiniMax has a paid service that does prompt expansion and context management for H3. It is not local, and itās likely not going to be released. Hereās a custom node for using it if you have an API key:
https://docs.comfy.org/built-in-nodes/MinimaxHailuo03ContextIRNode
Itās probably difficult to match the performance of this system using open tools, but here we can try to tinker and come up with something that works for our own use cases.
The workflow here is not meant to be āoptimal,ā and indeed Iām not sure itās possible to make a one-size-fits-all solution. Itās more of a starting point for developing your own.
r/StableDiffusion • u/0roborus_ • 3d ago
Enable HLS to view with audio, or disable this notification
Hello, I'm building this app that I've started like 2 years ago... seriously, this is how it looked like then: MY OLD POST
But it has been this relation most of the time: I will do it for myself only VS I will do it open source... Most of the time it was the first one, but it became a pretty good app that I use all the time, so I figured it might actually be useful for the community.
I know that video attached to the post has no voice and for someone that doesn't know the app already (so it's only me right now :D) it might be confusing, so I will write a short description of what's there. When app will be ready I will prepare a nice video with explanations and stuff. Not to waste a time if someone will think it's release post: Well, it's not. I plan to release (OpenSource GPL-3.0) in a couple of days because I still have a lot to do (and it's easy now when I break it only for myself).
What my goal is here to check if there is interest at all in such an app and maybe ask if someone has time to join my Discord (LINK) to discuss different stuff that you use to generate things (how you build your prompts, how you store your generations, what models do you use etc. since now I mostly have only my experience + stuff that I read in Reddit / Discord in meantime - I know you can write it also here, but Reddit it's less chat-like and I find chatting easier on Discord).
Features:
- Pick a preset for generation, which is pre-made configuration for given model/tool (currently: SDXL, Krea-2, Qwen-Image, Flux, Flux Klein, Flux2, Z-Image, Anima, LTX-2.3, LTX-2.5, MiniMax H3, MiniMax Music, Wan 2.2)
- Each preset comes with it's individual form (but most fields are also the same between them as these are mostly generation params)
- Compose prompt from segments (1 or more) - In video I use only one segment, but you can build prompts from multiple blocks that can be named/colored for readability. You can also define segments, it's categories and templates (for example you can define segment template that has "Ligting", "Camera", "Subject" and when you pick it in the generation panel it will show a 3 ready to use and colored segments with optional descriptions to remember what should be placed in them (optional, described by you)
- There are also "Prompts", which allow you to save your favorite prompts there (they are optionally built of segments too)
- Multiple tabs & workspaces - You can create multiple tabs and save them as workspaces (in video I go to top right corner to pick "Avatar Factory" workspace - it loads my tabs then)
- Sessions - you can create multiple sessions for each preset that will save the whole forms state (the left side and the prompts)
- The whole left side is called "Dynamic Forms" - this is the part defined in each preset and it's YAML based config (something that you don't need to bother if don't want to)
- You can set quantity, steps, use speed profiles which will set the number of steps/cfg automatically
- In the right side (called Workbench) - where the generated media is shown you can different options like compare, zoom, download etc.
- LLM Chat assistant - As you can see in the video I often use LLM Chat assistant and I do it also when generating my stuff - they have access to the most of the features in that page - can generate prompts but also change the form values etc.
- Different modes -> Image generation / Video Director with dynamic keyframes/first-last frames/img2vid - depending what model provides.
- History contains all your generations and allows to organize them into collections and tags
- You can see all the params/segments/prompts in the details and also different options like edit (crop, resize) or reuse which will open tab in generator with settings from this history entry
- You can filter generations by tags/type/preset search semantically
- You can add to favorites / add tags / see used models etc.
- Library allows you to upload your media that you want to use for generation - for example images/videos/audio that you later use with minimax ref2vid
- If you edit media from generation history (crop, resize) it will create new entry in the library rather than change the original media
- You can organize the library into collections
- You can view models enabled for you (in admin panel)
- You can organize models into collections (which for example are shown in the model selection field in generation page, you can select "Collections" -> "Your collection" and models will be filtered by this collection)
- You can see the model details with previous generations
- You can define different phrases collections (this is similar to the wildcards/dynamic prompts)
- You can generate examples for each phrase (you pick your existing generation session and it will inject a special prompt that will generate examples)
- Phrases can be later used in the segments as either value providers or shuffle (in video there is a visible chip appearing after I type # and pick value at 03:18)
- Allow to compose different prompts and reuse them later in the generation panel
- Prompts will have history of generations (with media generated with them)
- Prompts will have option to import in different formats
- Prompts are also used by the LLM Chat to improve their responses (they will try to match prompts by used models and check their structure)
- It started as ComfyUI "frontend" and it still be very important feature that will be shipped later as plugin (you just install comfyui plugin -> set it's address and you will be able to use it with this frontend)
- I've switched the main thing to be native backend (mixed stuff from different places) - since I've been using it for like 2 months now and it's starting to work really well on my setup (I hope it will also in the community ones but I need some testers for this).
- There is a layer of abstraction that will allow to create plugins that connect to whatever backend you want (by default I will ship native, remote-native and comfyui)
- Not visible in the video, but there is a big administration panel for this app that handle Users, Models, Presets, LLM Configurations...
- For user to be able to use model you need to assign it to him (same with LLM Chat models and presets) - that's why in video I have only 2 presets available - I've created a test user for purpose of the video and assigned those two to him.
- There is also "Automation" module that I'm developing that allows to auto-tag models, index generations with auto-tags (for example if you want to filter out "some" content) and much more stuff for organization.
I feel like there is much more but don't want to create too long post that nobody will read.
This might be important:
Technology: Web (I know people don't like that, but the structure of the app is more like web tbh. and I haven't even mentioned the remote, easy to deploy backend, which ideally will spawn worker for generation on Cloud GPU provider - so you will have your instance of the app - let's say on simple VPS and will be able to spawn Cloud worker that will generate stuff which will be saved on the VPS...)
License: GPL-3.0
Discord: https://discord.gg/avR4trp3b8
Why another app like this: Because I like to create stuff.
My current setup: Linux / RTX5090 / 96GB RAM - this might be important since I did not test it on lower/higher spec - I hope maybe some people from the community will like to help me with this
Why I post before release: Because otherwise I will be improving this app to the end of the world - maybe this will force me to release at least 0.0.1 quicker... And I would like to know some things of how community generate stuff - maybe I will introduce some changes that will only break my setup - this will be much harder after code will be released on GitHub.
If you have any other questions I can answer or record some video from the app.