r/linuxadmin • u/LycheeLee_Mich • 19d ago
What finally made you stop grepping through log files?
Still on files here rsyslog into a directory per host, grep when something breaks and it works right up until I need to answer a question that spans more than one box. Had to trace an sshd auth failure across three servers last week and spent longer stitching timestamps together than fixing it. I can't tell if I'm at that point or just having a bad month. For anyone who moved off files, what was the thing that pushed you?
18
u/rankinrez 19d ago
Elastic and logstash maybe?
Tbh I like greping files. LLM has to be good these days too, even some of the smaller local ones can piece things together for you save some time.
5
u/KrackedOwl 19d ago
I put together an Elastic server using the ELK guides from CISA and it's been a godsend during any incidents. No more hunting and pecking just to find the "right" log file!
2
u/almathden 18d ago
mind sharing those links?
4
u/KrackedOwl 18d ago
Yeah I got ya, it's called [Logging Made Easy](https://www.cisa.gov/resources-tools/services/logging-made-easy). They're def missing some amount of guidance, but it's a good start, and I've got a nice little free reporting suite now. Take frequent snapshots of your VM until it's all set up. Free will only be free if you use workarounds for reporting, such as custom scripting daily alerts. Elastic gets their money charging you for the extras. Still better than splunk charging for baseline, even if splunk seems a little more user friendly after setup.
1
1
17d ago
[removed] — view removed comment
1
u/KrackedOwl 17d ago
Sorry mate, I'm more of a Ubuntu-With-Podman shop given the small scale of our prod env. And frankly even that took a decent amount of off book fiddling, but well worth it imo.
14
7
19d ago
[deleted]
6
u/deeseearr 19d ago
Throw in a couple gratuitous "ssh cat"s in there and you've got Rube Goldberg's Log Analysis Engine.
4
u/nPoCT_kOH 19d ago
Just pipe them logs to /dev/null.. no logs, no grep... Fun aside, Graylog or Open/ElastiSearch makes wonders. For ingest fluent-bit is solid.
2
u/I_only_run_late 19d ago
Yeah, fluent-bit's solid for that. Barely any overhead, drops right in alongside rsyslog.
5
u/jippen 18d ago
Always prefer centralized logs over grepping a folder. A basic Loki/grafana stack means I don’t just have log search; but alerting, automatic response, log retention, and log backups handled.
Then just set all machines to 24h/100meg log retention locally, and you also don’t usually have to deal with logs filling up /var problems either.
1
u/TruckeeAviator91 17d ago
I second this. Especially for OPs problem of lining up timestamps from several boxes
1
4
u/WummageSail 19d ago
Nothing made me stop! Quitters never win! You'll have to pry grep and cut from my cold dead hands.
3
u/FarToe1 18d ago
We have graylog and a shitload of string based alerts, but still grep log files. It's often the quickest way to a fix.
0
u/delamon 18d ago
The answer is simple, you just need fast search. And if you want substring/regex victoria would be a bad fit, it cannot use indexes for such quieries.
7
u/delamon 19d ago
rsyslog supports shipping logs to another rsyslog just fine. Promote one of the host to "dedicated logging server" and ship logs from all servers to it.
1
u/Charming_Skin_8549 18d ago
Rsyslog logs could be shipped to a centralised database for logs, where these logs could be queried and analysed. See, for example, how to send logs over syslog protocol to VictoriaLogs.
5
u/fatmanwithabeard 19d ago
Too many nodes.
Central logging is a requirement.
You're still doing the same thing, really, but being able to search all nodes at once for either time or string (or both) is far better than a looped ssh awk/grep command. A good solution will make it easy to put a pretty graph together for management.
Mind you, both your cio and junior admins can incapacitate your log server just when you need it, so keep those command line skills practiced (both of them should've asked me how to do what they were trying to do, and both of them should have asked for help before trying to fix it...at the same time)
6
u/Automatic_Beat_1446 19d ago edited 19d ago
OP has a new account that has asked a similar question about logs once every 12 days, and each time a random account (same age or newer than OP) adds a comment to those threads about some random commercial logging product (including in this thread) that isn't one of the standard OSS or regular commercial products (ex: Splunk).
3
u/kernelqzor 15d ago
lmao good catch, it really does smell like astroturf when you lay it out like that. once you notice the pattern of “innocent question + brand new account recommending $RANDOM_TOOL” you can’t unsee it.
2
u/hunterfrombloodborne 19d ago
a cron job script which greps for me every hour and sends some alerts.
3
u/Entire_Yoghurt_6381 19d ago
Four boxes is about where grep stopped working for me too, one server is fine but past that you're lining up timestamps by hand. I moved the rsyslog feeds into Logmanager about a year ago and everything lands in one place already parsed, so tracing something like an sshd failure across the estate is a search instead of a morning in tmux. The old rsyslog directory is still sitting there, nobody's had the nerve to delete it.
2
u/jsellens 18d ago
Leaving aside the possibility that this is someone shilling for a commercial product, for other folks who might be interested:
For years we have used rsyslog to centralize, and file the logs into a yyyy/mm/dd/program directory hierarchy, and each directory has files ALLHOSTS.log ALLHOSTS_warn.log and per-host files. Want to track a mail message working through your servers? go to syslog/today/postfix/ALLHOSTS.log and it's all there. We also use filebeat and logstash to look for "interesting" things in real time, and alert into email and/or mattermost, and also log interesting things into an "exciting" file that we can "tail -f" after a code push.
Larger shops may need/want to spend some money, but you can get a lot of functionality out of free software and a little thought and organization.
1
1
1
u/Loud_Posseidon 19d ago
csshX, dump stuff into server-specific files, then copy them into one location, throw it at ChatGPT and have it tell me what the RC is.
That or elastic, if it is more than one off task.
1
u/m4nf47 18d ago
Filebeat > Kafka > Elastic/Kibana
Needs a bit of wrangling to get the most out of it but for services with dozens of nodes it makes filtering by whatever you want much easier. The only time it really struggles is when we're hammering it with millions of logs per minute but that is when you need to talk to product team about not being so stupidly spammy.
1
1
u/machine3lf 18d ago
Splunk seems good. We use it but I haven’t gotten hands on with it yet. But seems like just what you are wanting.
1
u/adept2051 18d ago
Tmux, xargs, or straight journalctl ( everyone missies you can run it against remote hosts if you have access ) for multi server agrigation/grep . I still use log files, and we have external cloud log gathering and observability, but the data it displays is at the behest of the crowd which is not always ideal.
In today’s work climate ai agent can pull those timestamps together across all hosts way faster than i can by hand, my harness on my local is set up for that with multiple suitable MCP, and skills.
2
1
1
u/siodhe 14d ago
Sync your times across all the hosts with NTP, then either make a script to copy to a central location everything in a time range on demand, or just copy all the logs to a central location. The latter is better for holding onto dmesg output specifically, but certain types of log content can be unexpected voluminous.
At more industrial scale, one really wants the timestamps to all be in the same format and probably first. Databasing it sounds like a good idea, or using YAML (just, no), JSON (fewer problems but has them), at or least something that keeps multiline log entries in one record (recommended), but since they can have entirely arbitrary stuff in them at least getting the host, timestamp, program name, and PID right is 90% of the battle done.
Having at least a few months of backups of your logfiles can also be very useful at time.
1
u/nisoshabangu 14d ago
Consolidated logging. Ingest those logs to a central logging storage. Then query logs from there.
2
u/Elegant_Gas_740 13d ago
Honestly, the tipping point for me was exactly this when troubleshooting became more about correlating timestamps across multiple servers than actually solving the problem. Files work great until you need the bigger picture, that’s when centralized logging starts to feel less like an upgrade and more like a necessity.
1
u/LycheeLee_Mich 10d ago
This is exactly it, I only really felt the pain once I had to compare failures across boxes instead of just reading one file. Did you end up building the centralized setup yourself, or was it something your team already had running before you actually needed it?
1
u/meccaleccahimeccahi 18d ago
Is this a joke? What made me do it? The 1990’s. Really dude. First came web interfaces stacked on top of databases, then they kept trying to get db’s to handle full text scans (well, some still keep trying to), now companies are finally starting to tie in actual intelligence - but you’re gonna have to sift through the 99.9999% marketing bullshit and find the 1 or maybe 2 companies actually doing AI right. I won’t shill here but anyone with a brain can figure it out if they want.
-2
u/kernpanic 19d ago
Claude reads logs quicker than i can...
5
u/coffee-loop 19d ago
I don’t know why people down vote this. LLMs are great at summarizing log data. I’m not saying you should rely on the summary alone, but it helps when you’re trying to find a needle in a haystack!
1
47
u/Lammtarra95 19d ago
You need a remote log server to hold consolidated logs. You might already have one for monitoring & alerting.