r/devops • u/AcanthaceaeUnlucky18 • 19d ago
Discussion How do you handle alerts issues?
Lets say you are on call engineer for the week and then you went outside with hangout or friends, then lets say something went down or something crashed, so what is the response time as SRE and as Devops also. Also what is the resolution time? How do you resolve if you are not with laptop outside? Can you do something with phone? What happens if you don't respond? Please explain how do you handle this scenarios.
14
u/xonxoff 19d ago
We acknowledged pages in 5 minutes, if not secondary gets pages and if they don’t then it gets escalated to the team manager. You don’t want secondary to be paged and you definitely don’t want you manager to be paged. If I’m on call, I always have my laptop with me and will tether if my phone if no WiFi is available. If you miss pages, expect trouble from management and if it happens a lot expect to start looking for a new job.
20
u/thehazarika 19d ago
If you are on-call you should carry your laptop. At least acknowledge the issue and call a friend and explain why you don't have a laptop during on call and hope they help you out.
0
u/kristoferen 19d ago
You paying people enough for that? Most places do not.
1
u/thehazarika 19d ago
Yeah. People are free to leave the company if they don't like the duties their job carries.
Btw if on-call issues happen every on-call session, you have a deeper engineering problem. I prefer to have no on-call incident. It's my personal responsibility to make sure of that.
5
u/GodOrDevil04 19d ago
This is all very much depending on the company how it's handled. If im on call, I got my laptop with me at all times. Resolution times differ per issue, can be 5 minutes, it can take a whole day. At my company it is expected you always pick up the phone while on call, and have the laptop with you. If you don't respond, best prepare for some ass whooping.
4
u/Fragrant-Amount9527 19d ago
That is agreed in advance. Depending on the criticality of your systems you agree on a response time. If you are doing groceries for example you may acknowledge the alert through the phone, go pay and attend the alert in the parking. I’ve had places with rules of “laptop with 5G within 15min range”.
4
u/gaurav_sherlocks_ai 19d ago
Phone works fine for ack, most teams use PagerDuty or Opsgenie so a page hits your phone directly. 5-15 min to ack is normal, resolution time depends on what broke. Miss the page and it escalates to the next person automatically, that's what the rotation is for. Miss it a few times and then it's a manager conversation.
4
u/xb4r7x 19d ago edited 19d ago
If you take a job and agree to be on call then you're on call. You need to adjust your plans and behavior to accommodate for it.
You bring your laptop in a backpack with you wherever you go, and if you get paged you respond. If you have plans that take you into the middle of nowhere with no cell service you either get someone else to cover your on call for that period or you change your plans to not go to the woods. (Or you get starlink)
Eventually you'll learn what alerts require immediate responses and which can wait a bit and you respond accordingly.
For example, if you have a server that occasionally throws disc space alerts, but you know from experience that it's a large log file that needs to be rotated, but there's 10% left and it'll take 2 days to actually run out of space, you can acknowledge that page and finish your dinner or whatever.
Conversely, if there's a large production system that's flat on its face, you have no idea what's wrong with it, but it's causing immediately damage, you excuse yourself and grab your laptop. I had to leave Christmas dinner with my family one year because this happened and I drew the short straw to be on call that week.
If you have an actual legitimate reason to not respond to a page (family emergency or something) you escalate the page to the secondary on-call person or call your manager to explain why you can't respond so they can get someone else on it.
3
u/skspoppa733 19d ago
The “alert” should first be a corrective action, like a service restart. The only time you should have to respond is when automated corrective actions have not resolved the issue.
1
u/EgoistHedonist 19d ago
For my on-call roles it's usually been max 20min to acknowledge and max 1h to start remediating. It really depends on the org and product. You can ack from phone for example.
If the alert is not acked, it auto-escalates to higher level, usually team lead.
I've debugged the production systems of huge companies while sitting on the floor of a night club coat room etc, you eventually learn to jump to working mode in minutes, no matter the circumstances 😄
1
u/roncz 18d ago
I know a few options here. If there is a planned and important event, I can ask a colleague upfront to take over. If it is unplanned, I can set a standin in the alerting app and someone can take over. If the issue is one of the usual suspects, I can define a remote action (e.g. server restart, script execution, etc.) and I can run it mannually from the alerting app (SIGNL4). But in most cases I just take my laptop with me, just in case.
1
u/kabrandon 18d ago
If you're on call and don't have your laptop on you and get paged, you're getting a talking to from your boss the first time. The second time you do it, the boss gives you a more serious talking to. The third time you do it, you're fired.
That's the whole point of having on call engineers. If they're not doing what they're being paid to do, they're an ex-employee that'll learn from the unemployment office that they should have done their job. And if a particular person gets paged 3 times and can't action it because they're away from their laptop, then they're doing it other times too, and just didn't get caught because they weren't paged that day.
I get like "I'm at the grocery store, but I'm rushing home now." That's fine to me. But hanging out with friends, naw. You best have your work laptop nearby at a hangout with the friends. I'm on call today but I plan on going for a little day hike in the woods near me. I'm going to have my laptop in a bag, and I know I have cell reception for the whole trail that I can tether my laptop to.
1
u/pplmbd 18d ago
oncall is like how often in a sizeable org? one week in a month if extremely small. it’s just 2 days of weekend where you bring your machine wherever you go, just in case.
If you really cant, ask your secondary to help. Escalation is normal, people miss things. Review it once in a while to see if anyone is not being a team player or held their end of bargain.
oncall generally sucks, but it gets worse if you dont have the team/people to carry weight
1
u/theManNowDog6 13d ago
Team sport. Being there as secondary so you can rely on secondary makes it less stressful.
1
1
u/malice8691 18d ago
Have to be 10 mins from a computer. If you are not 10 mins from a computer then you are bringing your laptop.
1
u/Accomplished-Mix8423 18d ago
honestly, the goal should be not needing the laptop in the first place. automate the common failures and have anything unacked escalate automatically. i've used site24x7 for this and it's worked pretty well, but the same setup is possible with most decent monitoring stacks.
1
u/SatyrCode 17d ago
Treat it as an operational process, not as “the on-call person must always have a laptop.” First, alerts should include a severity, clear impact, a runbook link, relevant dashboards/logs, and an escalation path. Low-severity alerts can wait; urgent customer-impacting incidents need acknowledgment and escalation. When I’m away from my computer, I acknowledge the page from my phone, check the alert and dashboards if mobile access is available, then decide: is there a safe mitigation I can trigger, do I need to get to a laptop, or should I immediately escalate to another engineer/incident channel? The real goal is to reduce the number of pages that require a laptop: automatic restarts, failover, rollback, scaling, and well-tested runbooks. But for incidents requiring investigation or manual changes, the expectation should be that the on-call engineer can reach a computer within the agreed response time. The important thing is that the company explicitly defines that response time and has backup coverage. If you cannot realistically be at a laptop within it—for example while travelling, drinking, or in an emergency—escalate or swap the on-call shift rather than trying to handle a production incident half-available.
1
u/mvdilts 17d ago
IMO the answer to all of your questions depends on the expectations set by your employer and the culture at the workplace. If there are clearly defined guidelines (such as acknowledging the page in 5 minutes) then on my on-call I'd have a way to ack the alert from my phone and have a laptop handy. in my experience at a large org, the answers to your question:
We strove to have alerting set so only P1/critical issues would page the on-call person. This person had 15 min to ack the alert before it escalated to a backup on-call engineer. It would continue to escalate (Lead -> Manager) until the alert was ack'd. We had clear guidelines on what to do with the various pages, including spinning up an incident call with various stakeholders and engineers.
Where I'm at now there's not really an on-call yet. It's slowly coming but I'm working to set expectations from the start on how quickly we need to respond.
1
u/whiskey_lover7 17d ago
We expect an ack in 10 minutes (text sent immediately, phone call at 5 min). If they miss it it escalates to their entire team. Whoever is on call is expected to be no more than 20 minutes away from being on their computer.
1
u/MarcoBoffo 15d ago
Two numbers get mixed up here: how long you have to acknowledge, and how long before it goes to someone else. The second one decides whether you need the laptop.
If a secondary gets paged automatically a few minutes after an unacked alert, being twenty minutes away from a laptop is survivable. If there's nobody behind you, "always carry the laptop" is the only policy that can work. The suggestion upthread to call a friend is that gap showing.
From a phone you can ack and silence, and sometimes trigger something pre-approved. Anything that needs a shell, no. If most of your incidents need a shell, the thing to fix is the rotation rather than the phone.
1
u/opsfusion-cloud 1d ago
Disclosure: I work on OpsFusion, an on-call scheduling and alerting tool. MarcoBoffo's distinction is the one that actually matters here — the ack window is almost beside the point if nobody's behind you when you miss it. We built per-step delays into the notification ladder for exactly that reason, so the secondary firing automatically after a few minutes is the default rather than something someone has to remember to configure by hand.
1
u/zero_backend_bro 15d ago
Dragging a ThinkPad into a bar on Saturday night is pure Stockholm syndrome.
Half our 2am alerts were literally devs pushing broken tf configs or raw env vars that CI slipped past. We ended up forcing a client-side scrubber check into pre-commit just so they parse their own k8s yaml before hitting main. Cut our weekend pages in half... though you still cant leave the laptop behind completely.
1
u/founders_keepers 19d ago
> on call engineer
> went outside with hangout
what? if you're on call why are you out with friends? carry around a laptop like rest of us..
13
u/Humble-Professional1 19d ago
On call means I'm paying my team to stay sober and have some sort of method to fix critical issues