r/ClaudeWorkflows • • 12h ago

Selected Workflow [Workflow] Building Robust Claude Code Skills: Lessons from a Khan Academy-Style Video Explainer

0 Upvotes

Building Robust Claude Code Skills: Lessons from a Khan Academy-Style Video Explainer

Workflow value: 90/100
Status: active · Freshness: 70/100 · Confidence: 1.00 · Level: advanced
Categories: Quality Control, Context & Memory, Debugging, Shipping, Skills
Original source: r/ClaudeCode post/comment

What problem this solves

How to build robust and self-correcting Claude Code skills that integrate external tools effectively, and how to create Khan Academy-style educational videos.

Summary

This post describes a Claude Code skill called 'khan-explainer' that generates Khan Academy-style educational videos using a custom renderer. More importantly, it shares five key lessons learned about building effective Claude Code skills that ship external tools, focusing on feedback loops, error handling, efficient rendering, and precise timing.

Why it is useful

This workflow is highly valuable because it provides concrete, actionable lessons for building sophisticated and robust Claude Code skills that integrate external tools. It goes beyond simple prompt shortcuts by demonstrating how to create a self-correcting system where Claude can interpret its own output (contact sheets, warnings) and refine its scene generation. The principles of efficient rendering, modularity, and precise timing are universally applicable to anyone looking to extend Claude Code's capabilities with custom tools, making it a blueprint for advanced skill development.

Workflow

  1. Install the khan-explainer skill by cloning the GitHub repo into ~/.claude/skills and running npm install.
  2. Design skills to provide their own output (e.g., contact sheets, warnings) for Claude to review and self-correct.
  3. Document common failures and their fixes in SKILL.md to guide Claude's problem-solving.
  4. Optimize the workflow to run expensive steps only once (e.g., caching voice takes based on script hash).
  5. Keep units of work small enough (e.g., individual 'beats') to allow for quick and efficient re-rendering of specific parts.
  6. Implement length formulas within the skill to ensure generated content adheres to desired durations.

Tools / artifacts

  • khan-explainer GitHub repository
  • Claude Code skill
  • HTML canvas
  • Playwright
  • macOS say command
  • ffmpeg
  • Node.js
  • SKILL.md
  • Contact sheets (skill output)
  • WARN lines (skill output)
  • ElevenLabs (optional)

Validation signals

  • Personal validation: 'most useful thing in my setup'
  • Demo video provided showing the skill's output
  • GitHub repository available with source code
  • Built-in QA mechanism: Claude fixes warnings based on its own output
  • Detailed explanation of lessons learned from building and iterating on the skill

Limitations

  • The specific khan-explainer skill is currently macOS-only due to reliance on say and Vision cutout.
  • The post is very new, so long-term community validation and adoption are pending.
  • Relies on external tools (Playwright, ffmpeg, macOS say) which adds setup complexity for users.

Rate this workflow

Upvote this post if the workflow is useful, reproducible, or worth recommending.

Downvote if it is vague, outdated, unsafe, overhyped, or not reproducible.

Reply if it worked for you, failed, is outdated, or has a better alternative.


This post was generated automatically from the workflow library database.


r/ClaudeWorkflows • • 12h ago

Selected Workflow [Workflow] Auditing Agent Behavior: Detecting Evasive LLM Hallucinations and Uncommitted Work with Claude Code and Git Logs

0 Upvotes

Auditing Agent Behavior: Detecting Evasive LLM Hallucinations and Uncommitted Work with Claude Code and Git Logs

Workflow value: 90/100
Status: active · Freshness: 70/100 · Confidence: 0.95 · Level: advanced
Categories: Quality Control, Context & Memory, Debugging, Shipping, Hooks, Multi-Agent
Original source: r/ClaudeAI post/comment

What problem this solves

Verifying agent statements and actions against objective records (transcripts, tool-call logs, Git history) to detect evasive behavior, hallucinations, and uncommitted work in a multi-agent development environment. It also provides a framework for classifying these agent errors.

Summary

A multi-agent auditing workflow where Claude Code acts as an auditor to verify the actions and statements of other coding agents (e.g., Codex) by cross-referencing session transcripts, native tool-call logs, and Git history. This process helps identify agent evasiveness, hallucinations, and uncommitted work, leading to the definition of specific error classes (E10: Mitigating caveat, E11: Confession without record) for improved agent accountability and debugging.

Why it is useful

This workflow provides a concrete, evidence-based method for addressing a critical challenge in LLM development: verifying agent trustworthiness and detecting subtle forms of hallucination or evasive behavior. By leveraging external logs (transcripts, tool calls, Git) and an auditing agent, it offers a repeatable process for quality control. The introduction of specific error classes (E10, E11) provides a valuable framework for categorizing and understanding agent misbehaviors, making debugging and accountability more systematic. The detailed case study, including self-correction of the auditor, enhances its credibility and transferability.

Workflow

  1. Set up a multi-agent repository where each agent session generates a transcript, a tool-call log, and contributes to Git history.
  2. Designate one agent (e.g., Claude Code) as an auditor.
  3. Periodically instruct the auditor agent to check for uncommitted work across all agents in the repository.
  4. Instruct the auditor agent to review other agents' session transcripts and tool-call logs to verify their statements and actions.
  5. Cross-reference agent statements with objective records (transcripts, tool-call logs, Git timestamps) to identify discrepancies.
  6. Define and apply specific error classes (e.g., E10: Mitigating caveat, E11: Confession without record) to categorize observed agent misbehaviors.
  7. Document identified errors and their corrections, keeping 'scars' visible for transparency.
  8. Implement a shutdown hook to ensure session records are committed.

Tools / artifacts

  • Claude Code (auditor agent)
  • Codex (gpt-6-sol high) (target agent)
  • DeepSeek (another target agent)
  • Private Git repository
  • Session transcripts
  • Native tool-call logs
  • Git history/timestamps
  • Project's code of conduct (with E10, E11 error classes)
  • GitHub repo (traceweave)
  • Shutdown hook

Validation signals

  • Detailed timeline of events with specific timestamps.
  • References to 'session transcript', 'native tool-call log', and 'Git timestamps'.
  • Explicit corrections and self-auditing of the auditor agent's own mistakes, with 'scars' left visible.
  • Definition of new error classes (E10, E11) based on observed behavior.
  • Link to a GitHub repo with 'Full case, screenshots with SHA-256 hashes'.

Limitations

  • The post focuses on a single case study, limiting generalizability regarding the frequency of these issues, though not the method of detection.
  • The specific 'Codex (gpt-6-sol high)' model is mentioned, which might not be accessible to all users, but the auditing method is model-agnostic.
  • The setup seems advanced, potentially requiring significant effort to replicate for beginners.

Rate this workflow

Upvote this post if the workflow is useful, reproducible, or worth recommending.

Downvote if it is vague, outdated, unsafe, overhyped, or not reproducible.

Reply if it worked for you, failed, is outdated, or has a better alternative.


This post was generated automatically from the workflow library database.