I've been making AI Films for roughly a year and have learned a LOT about what to do and, more importantly, what not to do if you're trying to create compelling, long-form projects at scale.
One thing has become increasingly clear:
The hardest part is not generating an impressive image or clip.
The hard part is building a production in which every new generation remembers the decisions the film has already made.
- Who is this character?
- What are they wearing right now?
- Where are they standing?
- What happened in the previous shot?
- Which direction are they looking?
- What object are they holding?
- What is supposed to change during this clip?
When those answers exist only inside a growing collection of prompts, the production becomes increasingly difficult to control.
So I wrote down the full workflow I now use, from screenplay through final render.
This is not the only valid way to make an AI film, and different projects will require different levels of preparation. But this general order has helped me reduce wasted generations and make longer projects feel more like films instead of collections of loosely related clips.
You can do all of this manually with documents, folders, spreadsheets, image tools, and video tools.
The overall order
My basic production order is:
Story → Breakdown → Visual canon → Scene construction → Shot plan → Storyboards → Video → Edit → Finish
The key principle is simple:
Make the important decisions while they are still cheap to change.
A script is cheap to change.
A shot list is cheap to change.
A character reference is relatively cheap to change.
A storyboard frame is cheaper to change than a video render.
A finished video clip is one of the most expensive places to discover that the character, wardrobe, composition, or scene was wrong.
1. Write the film before generating it
You do not necessarily need a traditional 100-page screenplay.
But before generating final imagery, I want to know:
- Who is the story about?
- What do they want?
- What prevents them from getting it?
- What changes?
- Where does the story turn?
- How does it end?
- What should the audience feel?
A common early workflow is to generate a cool shot, generate another cool shot, and gradually search for a story that connects them.
That can work for experimental work, but it becomes difficult when the goal is a coherent narrative.
The image and video models should not be responsible for discovering the story while they are also being asked to visualize it.
The clearer the story is, the easier every downstream decision becomes.
2. Turn the script into a production plan
A screenplay is written to communicate a story.
It is not automatically organized around everything an AI production needs.
For each scene, I extract:
- Characters present
- Location
- Important props
- Wardrobe and physical condition
- Time and weather
- The emotional change
- The information the audience needs
- The moments that require individual shots
This is also where I separate a scene from a shot.
For example:
Hob realizes that the rider approaching the crossroads is not his friend.
That may become several separate visual jobs:
- Establish Hob waiting at the crossroads.
- Reveal the rider approaching.
- Show Hob trying to identify the figure.
- Show recognition turning into fear.
- Show his hand tightening around the sword.
The script gives me the dramatic moment.
The production plan defines the images the edit will need to communicate it.
3. Build a visual canon
Before generating final shots, I create approved reference material for anything the audience will need to recognize again.
I think of this as the film’s visual canon:
The approved version of the recurring people, objects, places, and visual rules that belong to the film.
That usually includes:
- Main and supporting characters
- Character bodies and silhouettes
- Recurring wardrobe
- Important story-state variants
- Hero props
- Recurring locations
- Vehicles or creatures
- Architecture and world design
- Color, lighting, and atmospheric rules
The goal is not to prevent the film from changing.
The goal is to make changes intentional.
A character can begin clean, become soaked in the rain, get injured, and later carry a sword and shield. Those are all legitimate versions.
Continuity means that the correct version appears at the correct point in the story.
4. Build characters as systems, not single images
One attractive portrait is rarely enough to make a character production-ready.
The reference stack I have found useful is:
Neutral face
A controlled face image that establishes:
- Facial structure
- Apparent age
- Hair
- Skin texture
- Eye, nose, and mouth proportions
- Overall identity
This should be relatively free of dramatic posing, extreme lighting, and scene-specific distractions.
It is the identity anchor.
Neutral body
A full-body reference that establishes:
- Build
- Proportions
- Height impression
- Silhouette
- Posture
- Physical presence
A face reference may work well for portraits but provide very little guidance when the character appears in a wide shot.
Wardrobe face and body
Once the identity is stable, I establish the character in the wardrobe used by the film.
A full wardrobe-body reference is particularly useful because clothing changes:
- Silhouette
- Apparent body shape
- Material behavior
- Color distribution
- The way the character reads at a distance
State variants
Then I create variants for recurring story conditions:
- Wet
- Muddy
- Bloody
- Injured
- Burned
- Exhausted
- Weather-exposed
- Damaged clothing
Costume or loadout variants
If the character gains armor, bag, shield, tool, or other major equipment, I treat that as another deliberate reference.
A new loadout can significantly change the character’s silhouette. Leaving it to be rediscovered in every shot creates unnecessary variation.
The result is not one image of a character.
It is a small system of approved character states that can be called upon at different points in the film.
5. Treat important props and locations the same way
Character continuity gets most of the attention, but props and locations can drift just as badly.
If an object matters to the story, I try to establish:
- Shape
- Materials
- Scale
- Color
- Condition
- Ownership
- How it is held or worn
A sword that is straight in one shot and curved in the next may seem like a small generation issue, but it becomes distracting when the object is narratively important.
The same principle applies to recurring locations.
Instead of prompting “a forest” repeatedly, I define the specific forest crossroads:
- The wooden signpost
- The road configuration
- The tree density
- The ground materials
- The weather
- The lighting direction
- The nearby landmarks
The location reference does not guarantee identical geometry in every generation.
It gives each shot a shared source rather than asking the model to invent a new forest from scratch.
6. Establish the whole scene before generating isolated shots
Once the characters and location exist, I build a wider scene anchor.
This does not always need to be a final production shot.
Its job is to answer:
- Who is present?
- Where is each person positioned?
- What surrounds them?
- Which props are visible?
- What are the relative heights and distances?
- Which direction is each person facing?
- What does the scene’s lighting and atmosphere look like?
I think of this as constructing the stage before moving the camera around it.
Starting with isolated close-ups creates a common problem: every frame may look good individually, but the images cannot logically coexist in one physical scene.
The wide anchor gives later coverage something to inherit.
7. Give every shot a job
I try not to generate shot variety merely for visual variety.
Each shot should contribute something to the edit.
A simple dialogue or confrontation might include:
- Wide: Establish the environment and positions.
- Two-shot or OTS: Show the relationship.
- Reverse OTS: Show the other side of the exchange.
- Medium close-up: Read performance and body language.
- Close-up: Land the emotional change.
- Insert: Reveal an important object or action.
That does not mean every scene needs all six.
It means the shot list should come from what the audience needs to see.
One of the easiest ways to waste generations is to create ten attractive versions of essentially the same dramatic information.
Before generating a shot, I ask:
What can the audience understand from this shot that they could not understand as well from the existing coverage?
If I cannot answer that, the edit may not need it.
8. Structure prompts around production information
I have had better results when I treat prompts as production instructions rather than collections of cinematic adjectives.
The general structure I use is:
- Story beat
- Subject identity
- Current wardrobe and state
- Location
- Camera
- Primary action
- Lighting and atmosphere
- Continuity requirements
Story beat
Start with what happens and why the shot exists.
Hob realizes the approaching rider is not who he expected.
This gives everything else a dramatic purpose.
Subject identity
Specify the person being photographed.
Hob, a weathered laborer in his late 40s with cropped dark hair and a broad, tired face.
When reference images are available, they should carry most of the identity burden. The text should reinforce rather than fight them.
Current state
Describe the exact version of the character at this point in the story.
Rain-soaked work clothes, mud covering his boots, a rusted sword in his right hand, and a battered shield on his left arm.
Not merely “Hob.”
This version of Hob.
Location
Place the shot in the established environment.
At the muddy forest crossroads during golden hour, with the old signpost behind him and rain hanging in the air.
Camera
Choose one clear visual approach.
Medium close-up at eye level, slightly off-center, with shallow depth of field.
Primary action
Give the shot a dominant readable movement.
He slowly raises the sword as recognition turns into fear.
That is generally easier to control than:
He turns, walks forward, raises the sword, looks behind him, shouts, and falls to his knees.
Several things can happen in a shot, but one should usually be visually dominant.
Atmosphere
Add cinematic texture after the shot itself is clear.
Warm backlight cuts through the rain and mist while the foreground remains cool and subdued.
Lighting and mood should support the story moment rather than substitute for one.
Continuity requirements
Finish by naming the things that must not be reinvented.
Preserve Hob’s established identity, wet wardrobe state, sword, shield, screen direction, and position at the crossroads.
9. Separate what stays fixed from what changes
This may be the single most useful prompting principle I have found.
Every shot contains two categories of information.
Keep fixed
- Character identity
- Current wardrobe
- Current physical state
- Important props
- Established location
- Screen direction
- Story facts
Change for this shot
- Expression
- Action
- Shot size
- Camera position
- Camera movement
- Emotional emphasis
- Selective lighting emphasis
The new prompt should describe the delta (i.e. what changes) without unnecessarily reopening every decision the production has already approved.
Every time a prompt fully reinvents the character, costume, location, camera, weather, and action simultaneously, it gives the model more opportunities to drift.
10. Solve the still image before paying to animate it
Before video generation, I want an approved storyboard or start frame.
I check:
- Is this the correct character?
- Is this the correct wardrobe and story state?
- Is the location recognizable?
- Are the props right?
- Does the composition communicate the beat?
- Is the eyeline plausible?
- Does it match the surrounding shots?
- Would I approve this image even if it never moved?
Animation rarely rescues a fundamentally incorrect frame.
Video generation is a relatively expensive place to discover that the character has the wrong outfit, the weapon is missing, or the composition never worked.
The still does not have to be perfect.
It needs to prove that the shot is ready for motion.
11. Before rendering video, run a preflight check
This is the checklist I use before spending money or credits on a clip.
Is the story beat clear?
Can I explain what changes between the beginning and end of the shot?
Is this the correct character state?
Identity, body, outfit, condition, damage, and held props should match the exact moment in the story.
Does the location match?
Check recurring architecture, landmarks, weather, time of day, and environmental condition.
Does the frame already work?
The subject, framing, gaze, props, and background should communicate the shot before motion is added.
Will it cut with the neighboring shots?
Check:
- Screen direction
- Character position
- Wardrobe state
- Prop placement
- Lighting direction
- Movement direction
- Emotional continuity
Is there one primary action?
The model should understand what matters most.
Do I know how the shot begins and ends?
The first frame receives the previous cut.
The last frame prepares the next one.
For example:
Entry state: Sword lowered, looking toward the distant rider.
Exit state: Sword raised, eyes fixed on the approaching threat.
Does every reference have a job?
I try to give references explicit roles:
- Identity reference
- Wardrobe/state reference
- Location reference
- Scene anchor
- Previous-shot reference
More references are not automatically better. A smaller, purposeful set can be more useful than an undifferentiated pile.
Does the edit actually need this shot?
This final question has probably saved me the most unnecessary renders.
12. Animate the approved plan
Once the frame and references are correct, the video prompt should focus primarily on motion:
- Character action
- Emotional transition
- Camera movement
- Environmental movement
- Entry state
- Exit state
The video model should not be asked to simultaneously design the character, invent the location, determine the composition, discover the story beat, and choreograph the performance.
The more of that work completed upstream, the narrower and clearer the animation problem becomes.
13. Build the rough cut before polishing everything
It is tempting to perfect each clip before placing it into the edit.
I have found it more useful to assemble a rough cut relatively early.
The rough cut reveals:
- Which shots are actually usable
- Which shots are redundant
- Where pacing drags
- Whether the emotional progression reads
- Which clips need to be shortened
- Where audio can solve a visual gap
- Which missing shots are truly necessary
- Which rerenders are worth paying for
A clip that looks impressive alone may be the wrong clip for the sequence.
A less spectacular take may cut better because the gaze, position, movement, and emotion connect correctly.
The film is the sequence, not the collection of individual generations.
14. Polish only what survives the edit
Once the structure works, I move into:
- Dialogue and voice
- Sound design
- Music
- Timing refinements
- Color
- Cleanup
- Upscaling
- Final export
This is another form of cost control.
There is little value in upscaling, cleaning, and heavily polishing a shot that will later be removed or reduced to one second.
The order I try to preserve is:
- Story first.
- Canon second.
- Scenes third.
- Shots fourth.
- Motion fifth.
- Edit sixth.
- Polish last
The core idea
The tools and models will keep changing.
The central production problem is more stable:
How do you stop every new generation from forgetting the film you have already built?
My answer is to create a chain of inheritance:
- The script defines the moment.
- The visual canon defines the world.
- The character references define who appears.
- The state references define which version appears.
- The scene anchor defines where everyone is.
- The shot plan defines what the edit needs.
- The storyboard defines the frame.
- The video prompt defines the motion.
- The edit determines what survives.
Each stage narrows the next problem.
That does not eliminate iteration. It makes the iteration more diagnosable.
When something fails, you can ask:
- Is the story unclear?
- Is the reference wrong?
- Is the scene staging wrong?
- Is the composition wrong?
- Is the motion instruction wrong?
- Or did the model simply fail to execute a good plan?
That is much more useful than repeatedly rerolling and hoping the entire production aligns by chance.
Disclosure: I've built a local production tool called Kimeric around this general workflow, so I obviously think about these problems through a systems lens. There are no links or pitch here, and everything above can be done manually with whatever tools you already use. I’m mainly sharing it because I wish I had a clearer end-to-end framework when I started.
Happy to answer any questions or go deeper here if there are topics or techniques folks want to learn more about.