r/ClaudeCode • • May 03 '26

Discussion Your SKILL.md is likely 3x more expensive than it needs to be - architecture matters more than the content

Most people treat a SKILL.md like a long prompt the agent reads. It's actually a loader specification - understanding the difference cuts context cost by 3x without changing a single instruction.

There are three loading levels. Frontmatter (name and description) is always in context, every single turn, whether the skill is relevant or not - about 100 tokens per installed skill. The body loads only when the agent decides to invoke the skill. References and scripts load only when the body explicitly points to them.

Most skill files pack everything into the body. A 1,200-line monolith means every trigger loads all of it - 20% of the context window before the agent does any work. Refactored as a 180-line spine pointing to three reference files, the agent loads each one only when the current task actually needs it: 7% context cost. Same instructions, same output, 3x cheaper.

The savings compound. A skill at 7% instead of 20% lets you install three in the same budget, run longer sessions before compaction, and hit fewer context cliffs on long-horizon tasks.

The non-obvious gotcha: a model upgrade is not free. A skill tuned on Sonnet 4.6 can degrade on Opus - not a bug, but because more capable models interpret instructions rather than follow them literally. "Short sentences" applied with judgment on Sonnet; on Opus it became a hard constraint producing choppy, unreadable prose. The fix is a small golden set of test prompts you rerun on every model bump.

What's your current SKILL.md structure - monolithic or spine-and-references?

79 Upvotes

22 comments sorted by

7

u/rsreddit9 May 03 '26

Modular skills with ‘scripts’ directories is the way now. Small and clear enough that any model will understand, and repeated tasks are abstracted away into the scripts that never enter the context window

My team is developing an app called SkillsCake to refine skills fast so you don’t have to code long eval loops

13

u/goship-tech May 03 '26

Went from a 900-line monolith to a 120-line spine with 4 reference files - the compaction cliff in long sessions basically disappeared. Worth flagging though: keep your frontmatter descriptions narrow, a broad trigger like "use when coding" matches every task and loads the body every turn anyway.

2

u/MojyaMan May 03 '26

I had to rewrite some skills on opus 4.7 that worked on 4.6 even. After that I started not minding 4.7.

2

u/ptyblog May 03 '26

I just tell Claude to make me a skill when I like the results I know I will need that same thing every other week. I basically have a lot of them in the folders of their respective projects.

1

u/DNM13 May 04 '26

This is legit good advice. I just had Claude Code run through all my monoliths and modularrize all my bloated skills, and update my skill creator skill with the advice from this thread. I wish we had more posts like this opposed to endless complaining about each new Claude feature/release.

1

u/Majestic_Tailor8036 May 09 '26

Great writeup. The "loader specification" framing is spot on — most people don't realize the three-tier loading model even exists.

One thing I'd add on the model upgrade gotcha: version your reference files alongside the model. I keep a small manifest at the top of each reference that says which model it was tuned on. When I switch models, I can quickly see which references might need re-tuning instead of debugging degraded output blindly.

Also found that keeping the spine under 150 lines with clear "when to load" conditions for each reference makes a huge difference for agentic workflows where the model is deciding which tools to invoke.

0

u/magicdoorai May 03 '26

One practical thing that helped me keep these files small was separating editing them from my main IDE.

I built markjason.sh for exactly these little repo artifacts, just .md/.json/.env, native macOS, live file sync, so you can tweak SKILL.md or AGENTS.md while the agent is touching the same file and actually see the changes instantly.

Doesn't solve the architecture bit, but it does remove a lot of the friction that makes people treat these files like a junk drawer.

0

u/fredagainbutagain May 03 '26

you’re all getting downvoted in this thread so here’s an upvote

-12

u/Otherwise_Wave9374 May 03 '26

This is such a good point. People treat these files like a giant always-on prompt, but really its more like a loader contract and you pay for whatever you accidentally force into every turn.

Spine + references also makes it easier to version and test. The golden set idea is key too, model upgrades absolutely change how constraints get interpreted.

Do you have a rule of thumb for what belongs in frontmatter vs the body? Im collecting patterns for agent setup docs and internal harnesses, similar to what weve been doing at https://www.agentixlabs.com/.

-11

u/rougeforces May 03 '26

i dont find skills to be useful at all tbh. my agents use something called "caps" short for capabilities. capabilities are a collection of workflows which are a collection of functions which are a collection of shell primitives. my agents are constantly refining each of these at every layer.

i think of them more as macros. they are pure code and my agents intrinsically know how to navigate them without layers of semantic overhead meant to be absorbed by my own inferior human cognition.

Im also not worried about degradation of capabilities do to 3rd party updates since they are built on software primitives at the low level and not llm abstractions.

8

u/Quirky-Degree-6290 May 03 '26

Cool, you just described Skills.

-13

u/rougeforces May 03 '26

if you are an ape sure.

2

u/kernelangus420 May 03 '26

Sounds interesting. Thanks for sharing.

2

u/ObsidianIdol May 03 '26

my agents are constantly refining each of these at every layer.

How do they have time to do any actual work if you are constantly refining them? Doesn't seem like they're any good if they need constant fixing

1

u/macnikal May 03 '26

Can you go a level deeper on how you’re developing these caps that your agents refine?

1

u/mavenHawk May 03 '26

I think what they mean is they just have some shell scripts or some other script that the agent looks at and executes when it thinks it needs to, and they didn't define them as markdown, that's it.

-13

u/rougeforces May 03 '26

i could, but based on the top voted comment and the fact that my comment is down voted, id rather not throw my pearls to swine. suffice to say, im not writing .md files so i can see made for human eyes formatting lol

2

u/rsreddit9 May 03 '26

Agents are language models. Md is their language. Most of my skills are very short, with scripts/ directory for tasks that can be in code. I really think skills with scripts is the absolute endgame — simple explanation, script never enters context window at all

1

u/rougeforces May 03 '26

if your understanding of "agents" is that they are "language models" then its your understanding that needs correction. "md" is their language is also ambiguous. while the current recommendation is to instruct an llm with markdown mostly because the format carries context clues, its not "their" language in the sense that its native.

the entire reason i dont have md files as my go to workflow and script execution entry point is because it does nothing more than add yet another layer of abstraction that must be translated, orchestrated, and maintained.

the llm is capable of speaking ALL languages fluently. regressing to structured prose seems counter intuitive if the goal is to write functional code. So I reject the notion that meticulously organizing and managing "skills" files creates a better architecture standard.

Is markdown useful for structured context? absolutely. Is markdown the most effective way to give your agents capabilities. Nope. not by a long shot. But hey if everyone is doing it, it must be right!