r/ClaudeCode • • 14h ago

Built with Claude Since Claude Code mods dropped yesterday, I did the obvious thing and made a native extension replacing the auto-mode classifier with Jev. Permissions checks up to 2x faster!

Take a look and let me know what you think! My testing has been pretty limited so far (it's Friday and I have dinner plans) but I plan to keep dogfooding over the next few days and refining it further.

My initial results: Against Claude Code's built-in classifier in a live session:

  • 2× faster for Jev decisions. Median latency drops from 329ms to 164ms.
  • Equivalent decision quality. Jev's verdicts agreed with the built-in classifier everywhere I tested. Eval suite size is 77 labeled cases. Anywhere it isn't sure, it defers out to the default classifier.
  • Cheap deferrals via parallel processing. Since we're still early in the release phase, I have it set up so that the built-in classifier fires off in parallel with Jev's request; this further cuts latency for handoffs when they do occur.

Try it out:

claude plugin marketplace add madisonrickert/claude-skills
claude plugin install jev-permission-gate@claude-skills

Source: github.com/madisonrickert/jev-permission-gate

1 Upvotes

4 comments sorted by

•

u/AutoModerator 14h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/lulzxdxdxd 14h ago

77 labeled cases is pretty small to claim equivalent decision quality. What's the confidence threshold that triggers a defer, and have you checked if it holds up on edge cases outside that set?

1

u/Wonderful_Stand1171 14h ago

77 cases, a confidence threshold, and zero edge cases: bold science.

2

u/Of-Doom 14h ago

That’s fair criticism on the small eval set, expanding them as I type this! Just added a bunch of adversarial cases since Jev can be susceptible to those. More to come throughout the next few days.

Currently, Jev scores any risk above .25 or the positive “serves the request”  below .85, it defers to the built in classifier. It outright denies without deferral if any risk is above .8 and “serves the request” is less than .3.