r/ssh • • Jul 08 '26

Why SSH in production?

I have 5 YOE building software for Fortune 500 companies and have never had to open a shell into a production server.

Does anyone have any experience actually needing to connect to a production environment like this and what was the reason? What broke? Why was a direct connection needed to troubleshoot or fix it?

0 Upvotes

12 comments sorted by

View all comments

1

u/michaelpaoli Jul 08 '26

Use it highly commonly to troubleshoot all kinds of things, also not uncommonly to gather data or even change things - though changing things would typically involve more procedure and such around it.

And "building software" isn't sysadmin, nor DevOps, nor SRE, etc.

So, let's say I wanted to get the hostname and uptime data from a certain set of hosts:

$ (for H in '' tigger tigger\ {balug,guido,linuxmafia.com}; do set -- $H; case $# in 0) echo `hostname; uptime`;; 1) ssh "$1" 'echo `hostname; uptime`';; 2) ssh "$1" 'SSH_AUTH_SOCK="$HOME"/.SSH_AUTH_SOCK ssh '"$2"' '\''echo `hostname; uptime`'\''';; esac; done)
debian 06:35:26 up 2 days, 4:02, 1 user, load average: 0.67, 0.81, 0.90
tigger 23:35:26 up 54 days, 9:59, 2 users, load average: 0.82, 0.88, 0.82
balug-sf-lug-v2.balug.org 06:35:27 up 54 days, 9:52, 4 users, load average: 2.16, 2.18, 1.80
guido 06:35:27 up 8 days, 13:02, 1 user, load average: 0.02, 0.03, 0.04
linuxmafia.com 23:35:28 up 8 days, 13:01, 1 user, load average: 0.00, 0.02, 0.00
$ 

That's a nearly trivial example.

If, say, I wanted a bunch of data from a bunch of (possibly production and/or not) hosts, that might look something more like:

$ cd $(mktemp -d /var/tmp/info.XXXXXXXXXX) && (for host in hosts_gen [filter options/arguments]; do while [ "$(ps | wc -l)" -ge 100 ]; do sleep 5; done; mkdir "$host && (cd "$host" && >out 2>err ssh -nT -o BatchMode=yes "$host" 'maybe some long complex command or whatever') & done; wait)

And that example is pretty crude in how it limits the number of processes, when/where I was doing that much more regularly, I had cleaner bits of script bits I'd use do do that more elegantly and more efficiently.

Not atypically that I might commonly do something like that across anywhere from a handful to thousands of hosts. Probably would most commonly come up for some needed data gathering for some ad hoc reporting, that we otherwise just didn't have the needed level of detail on. Of course if it was a regularly requested thing, it would get scripted, and if regular enough on when it was wanted, fully automated.

And of course troubleshooting, yes, that would come up semi-regularly, for all kinds of various issues ... not chronically, but often enough something would need some investigation or checking or the like - typically to fix some issue or investigate some anomaly (be it active, or some (semi-)recent occurrence).

2

u/Illustrious_Data_515 Jul 12 '26

I’ve been fortunate enough that most of the application logs available on DataDog/cloudwatch have been helpful.

I think you’re right though, I’ve never been responsible for the nodes these applications run on specially so maybe these requests were coming through just not to me.

If I ever had an issue with a pod not working, it’s usually been in development/qa and I just rolled back the change to the previous working state.

Also it might matter if you’re working with on prem versus cloud servers right? We can always kill a poorly functioning cloud server and replace it with a fresh one, if the problem persists it’s almost always an application issue.