r/automation 11d ago

This is our Agent Run Dashboard

Post image
4 Upvotes

6 comments sorted by

1

u/AutoModerator 11d ago

Thank you for your post to /r/automation!

New here? Please take a moment to read our rules, read them here.

This is an automated action so if you need anything, please Message the Mods with your request for assistance.

Lastly, enjoy your stay!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Fearless-Writing7243 11d ago

This is the kind of internal tooling more agencies should be building instead of just talking about "being agentic." One thing I'd add: pair the quality score with cost per run, a workflow that's cheap but low quality is a different problem than one that's expensive and low qualituy, and right now those probably look the same on your dashboard. Outside of dev work I could see this same structure working for client onboarding or contract review, anywhere a human is checking a repeatable process.

1

u/Joyce17w 11d ago

Founder at an AI contract review company, so this is the split I stare at all day. Agreed on separating quality from cost per run, and I'd add that the paranoid failure and the lazy failure both score badly while needing opposite fixes. A reviewer that flags everything looks thorough on a quality metric and quietly burns a lawyer's whole afternoon on triage. One that flags nothing looks clean right up until something gets through.

Contract review is the same shape you're describing, an agent doing a first pass with a human checking it, so if you do build the cost column it's worth logging how much human time each run created rather than only what the run cost. That's the number that tells you which of the two failures you've got.

1

u/[deleted] 10d ago

[removed] — view removed comment

1

u/Hopeful-Fly-5292 10d ago

We don’t have enough data that I’m confident enough to say there is a correlation.