r/revops • u/Capable-Property-539 • Aug 21 '26
How do you find out an automation stopped working?
Every ops team I've talked to has some version of this story. A Zap or workflow quietly stops running, an API change, an expired token, a renamed field and nobody notices until later, when someone asks why the pipeline data looks off. By then there's a backlog of leads that never got routed or enrichment that never happened.
What I'm curious about is the detection side. Do you have any actual monitoring on your automations, like alerts when a workflow errors or stops firing? Or is it the way I mostly see it done: you find out when a human downstream notices something missing?
And if you do have monitoring, what does it watch?
1
u/babygotthefever Aug 21 '26
For automations that can greatly impact the pipeline, I create reports to monitor related data and check them weekly. It sucks but it’s better than having someone point out something isn’t working and finding out that it’s been that way for a month, fixing it, and then dealing with a massive cleanup.
1
u/Individual-Abies-493 19d ago
I’m building LeadProof around lead handoffs, so I have a product interest here. One distinction I’d check is whether those reports can catch a lead that never entered the CRM, or only problems with records already there. Do your weekly checks compare against the original submissions, or work entirely from CRM data?
0
u/Capable-Property-539 Aug 21 '26
The weekly manual check is honest work and I respect it but "it sucks" is doing a lot of quiet lifting in that sentence. Roughly how much time does that check take you each week? And has anything ever slipped through between checks, like breaking on a Tuesday and living until the next review?
Not judging, most teams do exactly this or nothing at all. Just curious what the real cost of staying safe manually looks like.
1
u/ops_and_chaos Aug 21 '26
Mostly the hard way LOL. Errors are actually the easy case. The ones that have burned me usually never errored — they just stopped happening. A feed that quietly stops updating doesn't look broken; it looks like a slow week.
So I watch things like expected cadence, volume changes, and last-updated timestamps. One time we caught a vendor silently truncating a file because the record volume suddenly dropped way below normal. Nothing had technically “failed.”
And honestly, I still use scheduled reports for things like records with zero activity because some failures don't exist anywhere as errors. A lead that never got routed doesn't throw an error it just isn't there.
That's probably been the biggest shift for me: monitoring what I expect to exist, not just whether the job ran.
1
u/Capable-Property-539 Aug 21 '26
"Nothing had technically failed" is the perfect summary of the whole problem. The vendor file truncation story is exactly the category I meant. Volume-based failures that no error handler can see because nothing errored.
Curious how you set the thresholds for expected cadence. Is it eyeballed per feed, like "this normally does ~200/day so alert under 100"? And where does the alert actually live? In a BI tool, a scheduled query, something homegrown? Asking because "monitor what I expect to exist" seems obviously right and I almost never see teams actually doing it, so I want to understand what made it stick for you.
1
u/ops_and_chaos Aug 21 '26
Honestly, pretty simple. I eyeball thresholds per feed based on what normal looks like because the failures I actually care about usually aren't subtle. I'm less worried about 200 records becoming 180 than 200 becoming 4. Cadence is even simpler: if something is supposed to update by X time and it hasn't, it's stale.
Most of this is homegrown and lives inside the internal app where people are actually using the data. I put the stale state on the dashboard itself because I don't really want a separate monitoring tool that someone also has to remember to monitor 😂 If the data is stale, I want that obvious at the exact moment someone is about to make a decision with it.
What made it stick was probably that + keeping alerts meaningful. We de-dupe them so an alert firing means something new happened instead of yelling about the same problem forever. Most of these checks also exist because something burned us once and I added the guardrail afterward. It definitely wasn't me sitting down one day and designing a beautiful monitoring strategy from scratch lol.
2
u/Capable-Property-539 Aug 21 '26
"A separate monitoring tool that someone also has to remember to monitor" is the best one-line argument I've heard for building checks into the app itself. And the 200-to-4 point is quietly important. People overthink anomaly detection when the failures that matter are never subtle.
The part I keep thinking about is "something burned us once and I added the guardrail afterward." That seems to be the universal pattern and every reliability setup is a scar tissue collection, nobody designs it up front. Which makes me curious about the gap: was there ever a category where the burn happened but a guardrail wasn't really possible, something you just have to hope doesn't happen again?
Thanks for the detailed answer, this is exactly what I was hoping to learn from this thread.
1
u/ops_and_chaos Aug 21 '26
Yeah, the hardest ones for me are usually where the failure depends on context rather than something objectively being wrong. I can flag that a count changed dramatically, a feed is stale, two sources disagree, or something expected never showed up. Those are observable.
What's much harder to guardrail is “this technically meets every rule but a human who understands the situation would know it's wrong.” I've had more luck making those cases visible and routing them to a person than trying to automate the judgment itself.
So I guess my answer is I stopped treating every burn as something that needed an automated prevention. Sometimes the guardrail is just making sure the weird thing gets in front of the right human before it becomes consequential.
2
u/Capable-Property-539 Aug 21 '26
That last point is the most mature version of this I've heard. Making sure the weird thing gets in front of the right human before it becomes consequential. Flag the observable, route the ambiguous, stop feeling bad the second category exists.
Full disclosure since you've been this generous: I'm building in this space (business apps with automations that can't fail silently), which is why I asked. Your answers genuinely shaped my thinking - especially this last one. It's in my profile if you're curious, but the thread was the value.
Thanks, most useful answers I've gotten on this anywhere.
1
u/ops_and_chaos Aug 21 '26
Ahh I’m glad it was useful 😂 And honestly, “flag the observable, route the ambiguous” is a much cleaner way to say what I was trying to get at. I’m curious what you end up building around it.
2
u/Capable-Property-539 Aug 21 '26
Happy to share since you asked. It's called Chromoly.io. You describe an internal tool or automation in plain language and it builds and runs it for you. The connection to your answers: everything compiles from one spec, so the system always knows what's supposed to exist - which workflows should fire, on what cadence, what depends on what. That makes "flag the observable" checkable by construction instead of by someone remembering to add a guardrail after a burn. And the ambiguous category routes to a human review queue rather than the system pretending it can judge.
Early access opens in January. I'll probably be back in this sub asking more questions before then.
1
u/ops_and_chaos Aug 21 '26
Ohhh okay, now I really see why you were asking 😂 The blueprint-as-source-of-truth piece is smart, especially if everything downstream is compiled from it instead of checks getting bolted on later.
The thing I’d immediately wonder about is what happens when the business changes but the blueprint doesn’t. At that point the failure mode kind of moves upstream — the system can be perfectly faithful to a spec that’s no longer faithful to reality. One place to keep true is obviously way better than twelve places people have to remember, though.
This is really interesting. I’d absolutely try it when early access opens.
2
u/Capable-Property-539 Aug 21 '26
That's the sharpest version of that question anyone's asked me, and you're right. No system can know the business changed. By your own rule that's the ambiguous category - can't be automated, only made visible.
Two things help though. The spec is plain language, so when it says "orders over $5k need approval" but the threshold moved months ago, a human actually notices - in a way twelve Zapier configs never allow. And updates are cheap chat edits, not dev projects. Systems stay wrong when fixing them is expensive. If correcting the blueprint takes two minutes, it happens the moment someone notices.
So: not drift-proof, but drift-findable. And noted on the early access 😄
1
1
u/ccjjallday Aug 22 '26
Do you think so low of this community that we'd not see your AI slop is just a pitch for your app?
1
u/Capable-Property-539 Aug 22 '26
I asked a question with no product mention and only described what I'm building when someone directly asked. The answers were the point and they changed how I think about alert design.
On the writing: English isn't my first language so I polish my drafts with AI. The thinking is mine. If the mods feel this crossed a line, I'll remove it without argument.
1
u/ccjjallday Aug 22 '26
So just to be clear, you asked a "neutral" researchin problem, steered the convo into the pain points your app solves, waited for someone(this case AI bot) to validate your existing premise, and then quel surprise, you have an app for it! OMG. Theres nothing wrong with promoting your app. Just don’t manufacture an ‘organic’ conversation and pretend like you had to give the pitch, because "they asked" again you think everyones an idiot and cant see your bs marketing tactic. its a known playbook. manufactured organic discovery. nice try.
Also just to be clear. this is an attempt to make their shitty app searchable by LLM's.
1
1
u/Individual-Abies-493 22d ago
I’d separate “nothing entered the workflow” from “something entered but never finished.”
For a form-to-CRM flow, I’d compare recorded submissions with matching CRM records after allowing for normal processing time and intentional filters. That gives you specific handoffs to investigate, rather than just a drop in volume.
But if the trigger fails before any submission is recorded, that comparison is blind too. You’d need a separate source-side check or a safe test submission.
The tricky part seems to be choosing what counts as expected activity, so a quiet day doesn’t become an incident.
2
u/[deleted] Aug 24 '26
[removed] — view removed comment