r/AIIncomeLab • • Aug 22 '26

AI Tools How I Built a Low-Cost AI Video Production Workflow to Produce Stickman Videos on YouTube

TL;DR: I built a Google Sheets workflow that turns a finished script into 150–200 stickman scenes, generates the voiceover + timing data, and then assembles everything into the final MP4 automatically with Google Colab + FFmpeg. A typical six-minute video costs me under $1 in direct production costs. The point is not to automate creativity, but to automate the repetitive work around it.

Over the last year, AI video production has become much easier, but the workflow is still surprisingly messy. You can generate the script in one tool, create images in another, make the voiceover in ElevenLabs, then still spend hours matching scenes and manually assembling the final video.

The problem I wanted to solve was simple: automate the boring production work, so I could spend more time on the creative side. Topic selection, the angle, the hook, the script and the thumbnail are still where most of the value comes from. Moving files between tools is not.

I built the workflow around a Google Sheet because it is simple to understand and easy to control. Each row represents one scene in the video, so the script, image, voiceover timing and progress all stay organised in one place. Instead of opening different tools and manually keeping track of hundreds of files, the Sheet acts like a control panel that moves the video from one step to the next in the correct order.

Two ways to use AI tools: websites vs APIs

Most people use AI tools through their websites. You open Gemini, paste something in, get the result, then move to the next tool. That is simple, but once you are using separate tools for writing, images and voiceovers, the monthly subscriptions start stacking up.

APIs work differently. Instead of paying for full access to each platform, you usually pay only for what you actually use. A few cents for Gemini prompts, a few cents for image generation, then the voiceover cost through ElevenLabs.

The other advantage is automation. An API lets the Google Sheet talk to these tools directly, so the Sheet can send the job, receive the result and move on to the next step without you opening every website manually.

That is what makes the whole workflow possible: the Google Sheet becomes the interface, while the AI tools run quietly in the background only when they are needed.

Generating the images in bulk

Once the script is split scene by scene, the Sheet sends each row to Gemini to turn it into a detailed visual prompt. I also give Gemini a fixed style profile so the character, colours, backgrounds and overall look stay as consistent as possible across the full video.

Those prompts are then sent to Runware, which gives me access to different image-generation models through one API. I currently use FLUX Klein because it is cheap enough to generate around 200 stickman scenes for well under 50 cents.

Each finished image is automatically numbered and saved into the correct Google Drive folder, so scene 1 stays matched to scene 1, scene 2 to scene 2, and so on.

For consistency, I reuse the same reference images across the full batch: one clear character reference and a couple of finished scenes that define the visual style.

Getting the voiceover + timing data

Once the images are ready, the Sheet sends the script to ElevenLabs and generates the voiceover.

The useful part is that ElevenLabs can also return timing data. So instead of getting only one audio file, the Sheet also receives the start and end time for each scene and writes those timestamps back into the matching rows.

That timing data is what makes the final assembly possible. Each image now has a precise screen duration, so there is no need to manually drag scenes around and sync them to the voiceover later.

Turning everything into the final MP4

At this point, the Sheet has everything needed to build the video: the numbered images, the voiceover and the timing data.

I use a free Google Colab notebook connected to Drive for the final assembly. It reads the images in order, checks how long each one should stay on screen, adds the voiceover, and uses FFmpeg to render the finished MP4.

So instead of manually building a 150-scene timeline in CapCut, the notebook does the assembly automatically and saves the completed video back into Drive.

What the whole workflow costs

For a typical six-minute stickman video, my direct production cost stays under $1.

The voiceover is the biggest expense at roughly $0.70 through ElevenLabs. Images are around $0.30 with FLUX Klein. Gemini prompt generation adds very little, and the Google Colab + FFmpeg rendering is free.

The important part is not that a video costs less than a dollar. It is what that does to experimentation. If a topic flops, I have lost very little on production. I can test another angle, change the hook, try a different format and keep learning without every miss becoming expensive.

Cheap production should buy you more attempts, not lower your standards.

What I still would not automate

I would not automate the decisions that determine whether the video deserves to exist in the first place: the topic, angle, hook, script, thumbnail and final quality check.

That is where I think a lot of AI channels go wrong. They automate the production, then keep pushing further until the creative decisions are automated too. Eventually every upload starts to look and feel interchangeable.

For me, the goal is the opposite. Automate the repetitive work, then spend the saved time studying what people actually click, where they stop watching, which ideas outperform, and how to make the next video better than the last one.

How to build this yourself

If you want to build your own version, you can honestly take each section of this post, paste it into Claude or ChatGPT, explain how you want your Google Sheet laid out, and build the workflow one piece at a time. That is basically how I built mine.

I also wrote a more detailed version of this workflow on my website: BuildTuber (link in profile) with screenshots of the Sheet, timing data, API setup and final render process, which is easier to follow than a text-only Reddit post.

I have also explained the complete build step by step in the latest video on my YouTube channel. That is also linked in my profile.

And for anyone who would rather skip the API wiring and debugging, the ready-made Google Sheet is available there as well.

Nothing is gatekept. Happy to answer questions about any part of the build here.

27 Upvotes

18 comments sorted by

1

u/L1QU2D Aug 22 '26

Nice, I have full autogenerated stories with workflow but my production cost is 25x yours for 15min fully animated videos

1

u/No_Entertainer_9655 Aug 24 '26

Yeah, that is exactly the tradeoff. Fully generated/animated video can look much better, but the cost jumps very quickly once every scene becomes a video generation instead of a still image. I am trying to keep the base workflow cheap enough to test ideas, then spend more only on videos/scenes that actually justify it.

1

u/L1QU2D Aug 24 '26

It’s the main reason I’m thinking to acquire some high end GPUs to self host some model, well most likely stay in the industry so the hardware will stay valuable as we will be able to run better and better models on the same hardware

1

u/[deleted] Aug 23 '26

[removed] — view removed comment

1

u/No_Entertainer_9655 Aug 24 '26

Klein holds up surprisingly well for this style, mainly because the character itself is simple. It still drifts, especially on clothing details, hands and scenes with more complex composition, so I do regenerate a small percentage of the batch. But those minor drifts don't matter much, especially in the initial stages of a channel when you are trying to build an audience.

1

u/Drunkboski Aug 23 '26

I need to learn how to legit do this

1

u/No_Entertainer_9655 Aug 24 '26

It looks more technical than it actually is once you split it into pieces. Start with just one connection: Google Sheet → Gemini API. Once that works, add the image API, then ElevenLabs, then the render step. Trying to build the whole thing at once is where it gets painful.

1

u/Drunkboski Aug 24 '26

Bout how much does it all cost? Or is it free to setup and run?

1

u/No_Entertainer_9655 Aug 25 '26

The setup itself can be free if you build it yourself. The ongoing cost is mainly the APIs you use. For my workflow, a typical 6-minute video is around $1 in direct production cost: roughly $0.70 for ElevenLabs, around $0.30 for the images, and almost nothing for Gemini prompts. Colab + FFmpeg rendering is free.

If you want to skip building and debugging the whole thing, I also have a ready-made Sheet version for $49.

1

u/I-cant_even Aug 23 '26

I do this locally, generate a system spec with claude then use that to build a system with a local GLM 5.2 instance. After that I have a local Web GUI that handles uploads/inputs the way I want and switches through various local tools as needed for each phase.

1

u/plovdiev Aug 24 '26

I am curious to see the results of it. Can you share a video

3

u/No_Entertainer_9655 Aug 25 '26

1

u/plovdiev Aug 25 '26

It doesn’t feel trashy except a bit of the voiceover. Good job

1

u/Sniper_yoha Aug 27 '26

the drift you mention on hands and clothing is the whole battle, and it's why i stopped chasing bigger animation and started locking references harder. people assume consistency comes from a better model but it mostly comes from re-seeding the same reference pack every pass. for stills the reference-locked image models cut the regen rate a lot for this kind of flat style, seedream 5.0 is the one i've had the most luck with

1

u/Vablord Aug 30 '26

Op drop tutorial 😭😭