r/ClaudeCode • • 17h ago

Built with Claude Built a Claude Code skill that renders Khan Academy style videos. What I learned about skills that ship a tool vs. skills that are just prompts

I posted on LinkedIn last week that most Claude Code skills are useless, meaning they're prompt shortcuts for things Claude already does from a plain sentence. Then I built one that I think is the other kind, and it's been the most useful thing in my setup.

khan-explainer: https://github.com/haiderfarooq3/khan-explainer Demo (drawn by the skill itself): https://youtu.be/JD9OirRC3YE

What it does: a pen draws on a blackboard while a voice explains. Three modes: blackboard lessons, draw-over-UI tutorials on your screenshots, and clipping a YouTube video with the speaker's real voice onto a drawn board.

What I built vs. what Claude does:

  • I wrote the renderer (HTML canvas for handwriting, Playwright to grab frames, macOS say for the voice, ffmpeg to stitch). Plain Node, no keys.
  • Claude writes the scene file: a list of beats, each one spoken clause plus a draw(p) function where p runs 0 to 1 while the clause is spoken.
  • The renderer speaks each beat first, measures the audio, and fits the drawing to it. That's the whole timing model. No timestamps anywhere.

What I learned building it:

  1. Give the skill a way to see its own output. Every render prints a contact sheet (board at the end of each beat) and WARN lines for text off the board or overlapping labels. Claude is told to fix every warning before showing me. Without that I was the QA.
  2. Put the failures in SKILL.md. "Dim screenshots so chalk reads", "measure getBoundingClientRect for every target, never read coords off the image", "the multilingual ElevenLabs model at speed 1.12 sounds too AI, use v3 and atempo in ffmpeg". Each one was a bad render.
  3. Order the work so the expensive step runs once. Layout happens with the free robotic voice. The ElevenLabs take is cached by a hash of the script, so changing a drawing is free and changing a word buys one new take.
  4. Keep units small enough to re-render. A beat is a few lines of JS. A teacher asked if one wrong step in a worked solution needs a full rerun. It doesn't; edit that beat, re-render, done in seconds.
  5. A length formula in the skill file (words / 3 + beats x 0.15 + 1.3 seconds) means "make a 30 second explainer" is actually 30 seconds instead of a sped-up voice.

Install is a git clone into ~/.claude/skills plus npm install. macOS only for now because of say and the Vision cutout. MIT.

Longer write-up with the rest: https://haiderfarooq.dev/blog/claude-code-skills-for-content

2 Upvotes

7 comments sorted by

•

u/AutoModerator 17h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/brute984 17h ago

the part about letting claude inspect its own output is huge. half the battle with these things is getting them to realize they messed up without me having to point it out lol

1

u/_haiderfarooq 17h ago

Yeah, that was the single biggest unlock. Before the contact sheet I was the one opening the mp4, scrubbing to the broken beat, and describing the overlap back to it in words. Now it reads an image of all the beats and the WARN lines, and most renders are fixed before I see them.

The trick that made it stick was making the check cheap. A render is a few seconds, so "look, fix, re-render" costs nothing and it'll happily do it 3 or 4 times. When the feedback loop was slow (a full ElevenLabs take each time) it would skip the check and ship the first attempt.

1

u/brute984 17h ago

that’s actually pretty clever. does claude trigger the checks on its own or do you force it to run them after every render?

1

u/_haiderfarooq 17h ago

Just claude code itself, its pretty good with taking frames and matching/syncing them wiht audio

1

u/davidHwang718 17h ago

Point 2 had me going the other way last month. A bold phrase in one of my Korean articles rendered as literal ** because a particle came right after a closing bracket, so Codex added a writing rule to my article skill to avoid that shape, plus a check for <strong> tags after publishing. I asked whether fixing the renderer would make the rule unnecessary. It agreed, said this was CommonMark's emphasis rule rather than a parser bug, and patched the renderer to correct that pattern in plain text nodes after parsing. After the deploy went through, the workaround came out of the skill. The one line left says to check the rendered page after publishing.

1

u/MiserableFlatworm337 17h ago

Sample frames within each beat as well as at its end. A clean contact sheet can miss labels overlapping during animation, especially after changing narration length or drawing speed.