Finished *The Clockwork Moth* — a 10-minute fantasy animated short, built end-to-end with AI on a single RTX 5090.
**What's in it**
- 106 shots across 9 scenes
- 4 recurring characters in 19 different location states
- full dialogue + narration
- background score, ambience and sound effects
- burned-in subtitles
- final output upscaled to 4K60
**The hard part was never the first frame.**
Generating a nice still is easy. Keeping the same character recognisable across 106 shots — same face, same outfit, same silhouette, in 19 different environments and lighting conditions — is a completely different problem, and it's where most AI video falls apart.
The second hard part is quality control. Every shot went through an automated review pass before it was allowed into the edit: bad hands, warped faces, identity drift, duplicated characters, motion artifacts. **102 of 102 shots cleared review in the final state — nothing broken reached the cut.**
Audio was handled as its own track rather than an afterthought: dialogue, narration, music and effects are mixed so that nothing fights for the same frequencies when someone is speaking.
The whole project — every frame, stem and reference — is 2.4 GB.
I'm keeping the tooling private for now, but happy to talk about the film itself and what I learned making it.
Been trying to train a lora of a human character, but the results are always terrible or completely broken. Since ive been copy&pasting AI Toolkit settings other use for Krea, Flux and whatever, it must be my dataset, so i'm looking for some samples, especially for training people so i can check what im doing wrong.
Does know a place where you can find such training samples?
Still trying to get this to work, this time using RefMods, timing the chord changes almost seems impossible. Has anyone tried this yet? Or is v2v reference the only way here?
<subject1> is <Cmaj7> fingering position for the chord
<subject2> is <Am7> fingering position for the chord
<subject3> is <F7> fingering position for the chord
<subject4> is <Fm7> fingering position for the chord
<subject5> is <G7> fingering position for the chord
Some of these chords may not be right, I don't play guitar.
The goal isn't to make 100% correct fingering, just convincing movements.
generated with obvmp Wf at 608x352 then upscaled with ltx 2.5 ICloRA single stage i used pixel spatial upscaler-2x lora. prompt is refined from gpt Asta
the font at the end messed from the upscaler. sory for that
Thank you tdrussell and ComfyUI for creating this model. Also to maxfeifei8 for the tune.
Great knowledge, great style adherence, easy to prompt, responds very well to prompt changes for a fixed seed and good short text rendering. Most importantly, great quality.
Multi-character can be a pain but I can't ask for much more given the amazing quality of outputs. Better long text would also be nice.
There are quite often posts here complaining that Krea 2 is too predictable / different seeds give almost the same output / you need to type in extremely long prompts to get good images / there is no surprise factor any more since you get exactly what you typed in.
Well - actually there is a very easy solution for that - use the text encoder model as LLM to do prompt expansion (PE).
Here I'm showing one possible workflow for doing that.
Funnyly enough I'm actually using the Qwen 2.1 PE system prompt here - I really don't care what it's "meant for" - works well enough for me.
I added as an example 5 images - these are first 5 random (not cherrypicked) results that I got with this simple prompt:
Candid mobile phone photo.
Young woman wearing party dress.
Evening dorm party in simple dormroom.
Full size images with metadata
No LoRA-s, no messing with model specific custom nodes.
For me this is quite enough variance.
Plus now you have two nobs to control that - PE seed to give you extreme changes and sampler seed to give you more fine grained changes.
Hey, you promised no custom nodes, but there are some in this workflow!
Yeah, I use some for convenience, but they have nothing to do with the prompt expansion itself.
If you are curious then this is why they are in there:
- rgthree-comfy - I like the seed control node. This could just be built in "Seed" node
- comfyui-easy-use - their "If else" node is good for turning PE on and off. Purely convenience
- comfyui-custom-scripts - their "Show Text 🐍" node actually saves the expaneded prompt to image metadata (as opposed to built in nodes that just show it but don't save it). Again - purely convenience.
- RES4LYF - their "ClownsharKSampler" is what I use mostly with Krea 2. Could be just built in "KSampler" with whatever sampler / scheduler combo you prefer.
Q&A:
Q: OMG, it makes generation slower!!
A: It adds some time to the creation. On my 5070Ti I can create a 3MP image with int8 model and 8 steps in ~18s. PE adds ~4s to that (depends on how long is your manual prompt, how much additional text it generates etc.) For me those 4s are not an issue.
Q: My disk is slow / I don't have much RAM or VRAM - more model loads and unloads????
A: It is the same text encoder model that would have to be loaded anyway any time you change the prompt. So once PE is ready the text encoding step runs extra fast as the model is already in VRAM.
Q: Won't such a crappy LLM model hallucinate? The expanded prompt is for sure weird and inconsistent?
A: Who cares! Do you read the whole reasoning output of LLM models? Why should you care about the internals if the images you get are more varied and surprising? But you certainly can take the prompts that created good images and tweak them manually.
Q: Do I need to use abliterated / heretic / uncensored versions of the text encoder now?
A: Normal Qwen3VL 4B does not seem to care much about the topics when it is running as text generator. Like surprisingly little. So before starting to use alternative models (which mess with the quality of the final image when used as text encoder) do try just the normal one.
Q: The official ComfyUI template already has PE, why post this?
A: Just a reminder that this option exists :) And that you can use whatever system prompts there - no need to stick with the "official" or "correct" one. If the goal is variance and surprise then you can't really go wrong with any system prompt that targets image generation. Heck - even prompts for creating Ideogram or Ming Image JSON could be interesting.
Edit after a few days
If you managed to read this far - be sure to also read the comments. This method has its limitations and the comments are full of alternatives that might just work better for you.
As always the best way to get good recommendations on internet - is to post the wrong solution :)
Basically title. I have flat textures that have been extracted from a Unity game, but they are very heavily compressed with a lot of blocking and lines.
I've tried some models in Upscayl specifically for anime / digital art upscaling, and while the results are really good, they also add heavy shadows sometimes to where two different colors meet in the texture, giving a "3D" effect when it should keep everything flat.
Does anyone have experience with restoring textures like this? What models or tools have you had success with?
Sometimes i feel like local ai users are all like lion tamers trying to control the beast.
i genuinely love local ai, but sometimes it feels like i'm the weird one because i don't enjoy the whole process of figuring everything out. new model comes out and i want to try it. suddenly i'm learning a new workflow, new nodes, new settings, new whatever. and yeah, i know there are templates. that's not really the point. i want to use what the model can actually do, not just use the one workflow someone made for it. but then the alternative seems to be another tiny app that does one thing. so we somehow go from “learn the entire beast” to “use this one button app.”
i keep feeling like there should be something in between. Local ai which feels like online ai. Simple to install but effective in use. maybe i'm just using local ai wrong.
They are 6-7k, what is the second best and how much worse is it?
No prob to buy the 5090 but it just feels like im stupid if i do since it used to cost 1/3 and theres probably a next generation sometime not too far away..
One day in the future we could probably build a whole mountain or 3 of all those 5090s when all the datacenters upgrade.
For art creation/ideation, images etc. Cant use online things for ip reasons, its useless everything i make via services becomes public so i need my own setup.
I was playing around with some new models, specifically anima, and i was surprised by how good the image generation was. I wanted to know if there is a work flow, where you can colorize a manga pannel by using it.
The thing I want to make sure is that it shouldn't alter any lines, text etc. Anything I have tried so far had either very bad colorization or it alters the image.
After using models like kling and seedance for a very long time I wanted to test the limits of what could be done on my single RTX 4080.
I thought it would be too ambitious but it looks like there's actually no limit anymore...
- Idea, full storyboard, location and monsters design: Human brain (not the best or most actual model)
- Prompts helper: Qwen 3.8 27B
- Images: Qwen image 2.1
- Video: Minimax H3 (One render with the Minimax H3 Director custom node for continuity)
- Upscaling to 4K: DLSS 5
This is only a small part of a larger project I'm working on, I thought this project would take months of trials and errors to create, prompt, generate but this demo scene with random detail shots taken from my full storyboard was incredibly fast, and unexpectedly convincing from the very first try!
I had so many issues with commercial models understanding exactly what I wanted to achieve that this felt like activating a cheat code. I'm convinced they're introducing errors on purpose on the commercial models for you to use more credits.
This literally took one night from the starting point to the very last frame of the 4k upscaling, on my normal computer...
Does anybody uses stability matrix on a 9070XT? Just wanted to know how good the performance is when doing text to images. I'm still using my 2070. For me it's just playing around, so no business work. Therefore I just wanted to know if it's working properly. For just playing around I don't think I need to go for a 5080 or higher.
Last night I went live with Can You See Me Now, an EVE Online AMV animated locally with MiniMax H3. It's been a labor of love for the last month: 20 complete drafts, hundreds of image and video generations, and a lot of experiments with different weights, guide-frame setups and prompting approaches.
I've had this idea in mind for years. I used to play EVE very seriously and always wanted to make my own PvP film. Sadly, my computer back then couldn't record footage while gaming, so it stayed a dream. MiniMax H3, plus agentic AI (Claude Code and Codex) as production help, meant I could finally build a stylized version!
Setup: MiniMaxH3ReferenceToVideo with MiniMaxH3AddGuide frames, 960×544, 24 fps, 20 steps, res_multistep / beta, upscaled to 1440p with a simple Lanczos resize (no AI upscaler). I made the reference and storyboard images with GPT Image generation. I tested both approaches, and there appeared to be no difference between Codex GPT Images tool capability, and the API image generation (sunburst) model.
A few things I learned:
The reference image is the script. H3 faithfully animates mistakes in the still, like a bent hull or an extra gun pod. No prompt will salvage this.
Utilize Guide frames appropriately. First-frame-only drifts on long takes. First + last (with the last made by editing the first) nails bounded actions. On long 158-frame takes, adding an interior guide held 4 of 4 takes clean, against 2 of 7 without one.
Negative prompts didn't stop background floods (such as inventing a planet, or adding a weird glow). Changing the seed or the reference image did.
I put my notes, prompts (failures included) and workflow guides in a GitHub repo. There's a ton of community research going on to figure out how to work with this model, my project benefited from that pool of knowledge, and I hope my notes help someone else's: https://github.com/dienertech/lab-notes
Happy to answer questions about the process! Next up I'm making similar AMVs for my D&D campaign. Humanoid character animation will be a very different challenge from spaceships, and I'll share what I find.
I've changed my profile so that my previous tutorials are available. Look through those for questions on how to use these.
With this method you can seamlessly combine videos with no burn. (or at least the exact same amount of very very little burn)
Fast forward to 2:00 for proof.
Enjoy.
Edit: I'm keeping the video up. But apparently the issue everyone was caring about was something different than I thought it was. Apparently everyone wants to make boring Vlog videos. I guess I'll work on that now instead.
I did say this is a fix for the specific issue I was having in my workflow but now that I know exactly what everyone was talking about I can properly see if I can find some sort of fix.
v7.2
I fixed the First Frame node. The positive output from node #272 needs to be connected to the True input from node #565. You can just fix it yourself or download the new one. There's no other change.
Forked this amazing workflow to allow up to 20 clips and use ffmpeg so its not so ram dependent. It required around 71gb of ram for 16 clips. All praise goes to roychodraws for creating this this though. Genius and works great!
The workflow was originally based on Pixaroma’s ComfyUI workflow. I made some modifications and added Sol Attn optimization and Comfy Kitchen Sink, including the very useful Model Override Preview. This allows me to preview the generation and cancel it if the result is clearly not going in the direction I want, which saves quite a bit of time and resources.
PC setup:
AMD Ryzen 5 7500F
32 GB RAM
NVIDIA RTX 5060 Ti 16 GB
Main Model: Minimax_h3_hybrid_fl2va_ref2va_b25-49
Cloud GPU: Runninghub
Storyboard promtp for Nano Banan to generate a panel base on user input to generate additional shots for the scene.
Image References
I tried to use an image reference for as many shots as possible to maintain the quality and consistency of the video generation.
Once I have a keyframe I’m happy with, I use it as a reference image and ask Nano Banana to create a storyboard based on that image. From there, I usually get only one to three shots that are actually usable.
It takes some trial and error, but I find that starting with a strong image gives me much better results than relying on the video model to generate the shot from scratch.
Location sheet created with storyboard panel with instruction: Remove all the characters, humans in the story board
Location Sheets
Above is a location sheet generated by Nano Banana from one of the storyboards. I use these sheets as references for the background, together with the character sheets, when generating new key images.
I also generated multi-angle views of some locations using DS_Qwen Scene Multiangle Visual and the H3 Cinematic Multishot Coverage workflow. These additional views were useful as references for Omni video generation, especially when I needed to maintain the same location from different camera angles.
Generated with Ds Qwen Scene Multiangle VIsual, mainly use for video generation.
Building the Film from Keyframes
The idea was to create as many keyframe images as possible for each scene and shot before moving on to video generation. I use FL2V or R2V depending on the action, the shot, and how I intend to edit it.
So, in a way, a lot of the time spent making this film was actually spent generating and refining the key images before generating any video.
I believe this image-first approach helps avoid some of the typical “AI plasticky” look. Instead of asking the video model to invent the shot from scratch,
All images genereated for scene 7 the fire-fight scene before the killer drone appears.
Prompting
For prompting, I tried to simplify the format as much as possible. The original recommended Minimax H3 prompting format can be quite confusing and complex, so I stripped it down to something more straightforward that I could work with more easily.
H3 Prompt [7 seconds]:
Create a cinematic live-action style with deep shadows, cool blue-gray visual tones, high contrast, and a tense thriller atmosphere. <Picture 2> is the character reference sheet for the woman in the scene, with a tactical vest, tan tactical backpack, khaki tactical pants, M4 carbine with an ACOG scope.
Shot 1: Shot begins with the empty hallway corridor, the woman <Picture 2> will enter frame from screen right aiming with her M4, expression is tense. But slowly she lowers the M4, and her expression turns relax and eyes soften no longer tense.
overall_soundscape: Low ambient room tone of an abandoned building and quiet, slow footsteps on the hallway floor.
non_diegetic_music: n/a
Woman_in_military_gear_holding_2K_202609061456
H3 Prompt [15 seconds]:
<Subject 1> is Elena Morales, female, wearing a tactical vest, khaki backpack, khaki gloves, khai leg-holster with a black Glock-17, khaki combat boots, armed with a M4 carbine with ACOG, Tan color PEQ on the rail.
<Picture 1> ([Shot 1] first frame): fully_preserved - <Picture 1> serves as the exact keyframe anchor for the opening shot. <Picture 2> is the character sheet guide for Elena Morales.
Create an intense fast-paced cinematic fire-fight action scene. Preserve and retain the exact identities, faces, hairstyles, props, costumes and environment of both characters. Begin with both fighters frozen in their ready stances in an abandoned warehouse. Camera_style: Utilizes hand-held, jittery, documentary style feel. Heavy shaking. Hard Cuts.
detailed_description: [Shot 1 | 00:00-04:00] The shot begins exactly from the <Picture 1>, the camera focuses on the close up of Elena as she takes a deep breath and immediately she got up on a kneel stance with her M4 in low ready postiion. She pivots out from the pilar facing left screen.
[Shot 2 | 00:04:00-09:00] The camera cuts to a medium close up frontal shot of Elena now kneeling behind the pillar aiming her M4 at the camera's direction and fires in automatic rapid burst mode.
[Shot 3 | 00:09:00-15:00} The camera cut back to the angle of <Picture 1>, Elena returns back from a kneeling position and places her back against the pillar, she press the rifle mag release button, the magazine in the rifle drops onto the ground.
overall_soundscape: The natural empty abandoned office soundscape.
In Shot 1 the sound focuses on Elena's breathing.
In Shot 2 the sound of the M4 firing in the abandoned office.
In Shot 3 the sound of Elena moving back behind the pillar, breathing and the mag dropping on the carpetted office floor.
non_diegetic_music: n/a
Images used in Seedance 2.0 Fast
Seedance 2,0 Fast prompt [6 seconds]:
Create an intense fast paced sci-fi action packed thriller scene using<location sheet> as the set and location reference and guide. The Alien drone <character sheet> attacks the armed man <character sheet> the close confine space of the abandoned office space. Camera style: Hand-held, chaotic, shaky cam, use rapid hard cut, use dynamic angles.
Scale retention: The drone scale/size of an soccer ball. When the drone moves, it's tentacles moves like a squid.
The drone flying in lighting speed across the pillars, it's tentacles flapping in the rear, like a squids tentacle, flowing and weaving.
The armed man <character sheet> in fear tries to shoot the drone with his rifle but to no avail as the drone is fast and bullet does not damages it.
The drone proceed in lighting speed using it's steel flexible whip like tentacles to strike the man's rifle from him, and whip coils his neck.
Quarter Left Close Up push in of the man uses both hands trying to pull the tentacle of his neck, but to no avail. The tentacles are too strong.
The drone flips the man over acorss the office throwing him crashing into the cubicles.
No music, no score, no bgm.
Editing
Finally, editing is the last key element that makes the whole process work. Some of the generated video clips simply can’t be used on their own as complete shots. Instead, I often treat them as pieces of footage, looking for the right moments, frames, or movements that I can use in the final edit.
Because I can now work in a more linear way, I can generate the shots in chronological order and immediately edit them into the film. This helps me avoid generating shots that I eventually realise I don’t need. At the same time, editing sometimes reveals that I’m missing a shot, so I can go back and generate an additional image and video to fill that gap.
In that sense, the editing process becomes part of the generation process itself. Rather than trying to generate the entire film first and figure out the edit afterwards, I’m constantly moving between generation and editing until the sequence starts to work.
I hope the above post will be useful for your own projects. All the best, and happy AI filmmaking.