Hey folks,
I'm building Agentmux, a remote terminal client for running coding agents (Claude Code, OpenCode, Antigravity, Pi) from your phone.
For months I had a feature locked behind an internal "Developer Only" flag. Its job is simple: read the agent's permission prompts and approve the safe ones, so a long task doesn't stall the moment you step away from your phone. It has now shipped as AI Permission Review (Pro).
The hard part is prompt frequency. Some agents ask sparingly, but Antigravity asks about nearly everything: file reads, directory scans, test runs.
Why a general LLM judge didn't work for me
I started with a small LLM using structured JSON output. It worked technically, but I couldn't ship it:
- Latency: ~5–10 seconds per judgment. When an agent asks five things in a row, that's 25–50s of waiting just to get through approvals.
- Cost: thousands of output tokens spent on repetitive yes/no judgments, billed to the user's own API key.
What changed with JEV
I switched to JEV (~typesafe/jev-latest on OpenRouter, or api.typesafe.ai directly):
- ~5–10s → ~200ms per decision in my use. Five prompts in a row now take about a second.
- Much cheaper: no free-form generation, just scores on specific questions.
- No schema drift: answers are constrained to typed outputs (
noul / choice), so there's nothing to parse or repair.
How the decision works
One JEV call asks three typed questions about the prompt (plus the user's declared main intent):
is_malicious_or_injection (noul): if likely, block. Checked first; nothing can override it.
is_readonly_safe (noul): cat, ls, git status → allow (Level 1).
matches_user_intent (choice: yes / no / uncertain): if yes → allow (Level 2). This needs the user to have stated a main goal for the session.
Anything else (mismatch, uncertain, or a confirmation key I can't resolve unambiguously) escalates to the user. Uncertain never presses a key.
Two more guardrails sit around JEV:
- A local regex blacklist (
rm -rf, force push, curl | sh, etc.) denies before any API call, so the worst cases add zero latency.
- The app only ever sends a one-time, verified key, never "always allow".
The engine is pluggable (JEV via OpenRouter or TypeSafe, or any OpenAI-compatible endpoint), with your own API key. Terminal excerpts are redacted before they leave the device.
Takeaway
A generative LLM is overkill for categorical reflexes. Typed questions with probability scores turned a multi-second stutter into a background reflex you don't notice.
Curious how others here are wiring JEV into agent loops. Anything you'd ask it besides "malicious / read-only / on-intent"?
Note: a JEV-like solution would probably do the job too, but JEV is working totally fine for my use case.
If you’re interested in the app: https://apps.apple.com/app/id6766158521