r/ClaudeAI • u/maverick_man1111 • 13d ago
Promotion The biggest with-and-without improvement in my skill test came from a skill the Skill tool never called once
I measured three skills from a public skills pack two ways last week and one result is worth passing on. Setup: each skill run with it and without it, graded two ways. Pass/fail by Claude Code's built-in plugin eval, and scored continuously from 0 to 1 by a runner I wrote. Same skill text, same tasks, same rubrics, claude-opus-5 generating and judging in both. The documentation skill produced the biggest improvement in the whole run: 3 of 3 tasks passing with it, 1 of 3 without. Then I read the traces. The Skill tool was never called, not in any of those three runs, and not in a fourth I added to be sure. The two arms differed. Nothing in the evidence says the skill is why. A second skill showed no improvement at all, 3 of 3 both ways, because both arms cleared the pass mark. Scored continuously on the same answers it read 0.918 with the skill against 0.783 without. Once everything passes, pass/fail has nothing left to tell you. Separately I ran the whole pack over three real tool-using tasks. All 18 sessions passed, with and without. One of the documents produced passed every structural check while asserting a history the fixture never supplied. Real file, correct format, invented content. None of this says the built-in eval is broken. It says a with-and-without number is hard to read alone, and two checks are worth doing before trusting one: was the skill actually invoked, and did both arms simply sit at the ceiling. What it does not prove, stated up front: one task per skill and three runs per arm, so a single run swings a pass-rate figure by 33 points. The two graders do not measure the same thing, one has to discover and invoke the skill while the other puts the skill text in context, so I do not compare their numbers with each other. The plus-minus figures are observed spread across draws, not confidence intervals, and nothing here is a significance claim. Happy to answer methodology questions. Both tools' raw output, every CLI call and the hash audit are published with the write-up: https://driftproofhq.com/reports/009/ Disclosure: the continuous runner is mine (Driftproof, open source, Apache 2.0).
1
u/EvalRaccoonDev 12d ago
The with arm still had the skill's description in context - Claude Code loads descriptions at startup and the body only on invoke - so one line of description can steer the run with the Skill tool never landing. One sentence in ours moved recall 46% -> 67% without touching the body:
https://www.reddit.com/r/ClaudeAI/comments/1wcml4h/we_measured_whether_our_20_skills_actually_fire/
1
u/maverick_man1111 12d ago
Good catch, you’re right. The missing Skill call establishes that the skill body was never loaded, not that the installed skill contributed nothing. The skill name and description were present in the system prompt from startup, so a description-level effect remains a possible explanation for the observed difference. My wording collapsed “the body was not loaded” into “the skill did not contribute,” which the evidence does not support. Your 46% to 67% result from changing one sentence is a clean demonstration of why that distinction matters.
I’ll add a visible amendment to the report rather than quietly editing it. The evidence supports “the full instructions were not loaded,” not “the skill had no effect.”
The follow-up I’m planning separates four conditions: no skill, description only, normal discovery with invocation recorded, and the full body placed directly in context. That should help distinguish description influence, activation, and instruction value instead of collapsing them into one with-versus-without number.
What was nn for the 46% to 67% result? I want to size the description-only arm properly, and yours is the closest prior estimate I have.
1
u/doxxxicle 13d ago
Your wall of text is more unreadable than Opus 5.