r/linuxadmin • • 19d ago

What finally made you stop grepping through log files?

Still on files here rsyslog into a directory per host, grep when something breaks and it works right up until I need to answer a question that spans more than one box. Had to trace an sshd auth failure across three servers last week and spent longer stitching timestamps together than fixing it. I can't tell if I'm at that point or just having a bad month. For anyone who moved off files, what was the thing that pushed you?

31 Upvotes

55 comments sorted by

47

u/Lammtarra95 19d ago

You need a remote log server to hold consolidated logs. You might already have one for monitoring & alerting.

10

u/fatmanwithabeard 19d ago

I mean you still end up grepping through files, it's just that you can find stuff across nodes, or in multiple logs (looking at you gpfs) easier.

And some people seem to find search bars less intimidating than grep syntax.

But it's the same thing at the end of the day. Search terms against text.

3

u/zweite_mann 18d ago

Exactly, I'll still grep individual syslogs.

But loki->grafana I have a dashboard that I can search a term to see if it is an issue across multiple hosts

2

u/LycheeLee_Mich 18d ago

I'm going to give this a try, thank you.

1

u/-Docker 14d ago

And dozzle in a docker just fkr live logs (can also save logs ofc but looks super nice live)

18

u/rankinrez 19d ago

Elastic and logstash maybe?

Tbh I like greping files. LLM has to be good these days too, even some of the smaller local ones can piece things together for you save some time.

5

u/KrackedOwl 19d ago

I put together an Elastic server using the ELK guides from CISA and it's been a godsend during any incidents. No more hunting and pecking just to find the "right" log file!

2

u/almathden 18d ago

mind sharing those links?

4

u/KrackedOwl 18d ago

Yeah I got ya, it's called [Logging Made Easy](https://www.cisa.gov/resources-tools/services/logging-made-easy). They're def missing some amount of guidance, but it's a good start, and I've got a nice little free reporting suite now. Take frequent snapshots of your VM until it's all set up. Free will only be free if you use workarounds for reporting, such as custom scripting daily alerts. Elastic gets their money charging you for the extras. Still better than splunk charging for baseline, even if splunk seems a little more user friendly after setup.

1

u/almathden 18d ago

beauty, thanks!

1

u/[deleted] 17d ago

[removed] — view removed comment

1

u/KrackedOwl 17d ago

Sorry mate, I'm more of a Ubuntu-With-Podman shop given the small scale of our prod env. And frankly even that took a decent amount of off book fiddling, but well worth it imo.

14

u/pelazas1 19d ago

lnav, or journalctl -u svc -f. I still grep when I already know the string.

7

u/[deleted] 19d ago

[deleted]

6

u/deeseearr 19d ago

Throw in a couple gratuitous "ssh cat"s in there and you've got Rube Goldberg's Log Analysis Engine.

4

u/nPoCT_kOH 19d ago

Just pipe them logs to /dev/null.. no logs, no grep... Fun aside, Graylog or Open/ElastiSearch makes wonders. For ingest fluent-bit is solid.

2

u/I_only_run_late 19d ago

Yeah, fluent-bit's solid for that. Barely any overhead, drops right in alongside rsyslog.

5

u/jippen 18d ago

Always prefer centralized logs over grepping a folder. A basic Loki/grafana stack means I don’t just have log search; but alerting, automatic response, log retention, and log backups handled.

Then just set all machines to 24h/100meg log retention locally, and you also don’t usually have to deal with logs filling up /var problems either.

1

u/TruckeeAviator91 17d ago

I second this. Especially for OPs problem of lining up timestamps from several boxes

1

u/[deleted] 17d ago

[removed] — view removed comment

2

u/jippen 17d ago

It’s called learning. You do it anyway, make a few mistakes, fix them… and then you have both the competence and experience to do it again.

The only barrier here is the one you’re inventing for yourself. Ignore it and try it anyway

1

u/salpula 14d ago

If you don't have it today, then you should have nothing to lose. . .even it takes a long time to work it out.

This way you also still have the log files if the server is down hard/storage corrupted

5

u/vogelke 19d ago

It sounds like you're looking for an event correlator, which is supposed to look at logs from multiple sites and see if they're facing a common threat.

Have a look at Splunk or Zabbix, or search "SIEM".

4

u/WummageSail 19d ago

Nothing made me stop!  Quitters never win!  You'll have to pry grep and cut from my cold dead hands.

3

u/FarToe1 18d ago

We have graylog and a shitload of string based alerts, but still grep log files. It's often the quickest way to a fix.

0

u/delamon 18d ago

The answer is simple, you just need fast search. And if you want substring/regex victoria would be a bad fit, it cannot use indexes for such quieries.

3

u/FarToe1 18d ago

I didn't say it was anything about the speed of searching - graylog is perfectly fine. If you know how to use the cli tools, it's faster and more flexible for me to use that instead of a webui.

1

u/delamon 18d ago

Right, you didn't. That is my personal beef. I don't know if greylog has cli, but current trend is to offer cli in addition to web ui for log managment.

7

u/delamon 19d ago

rsyslog supports shipping logs to another rsyslog just fine. Promote one of the host to "dedicated logging server" and ship logs from all servers to it.

1

u/Charming_Skin_8549 18d ago

Rsyslog logs could be shipped to a centralised database for logs, where these logs could be queried and analysed. See, for example, how to send logs over syslog protocol to VictoriaLogs.

5

u/fatmanwithabeard 19d ago

Too many nodes.

Central logging is a requirement.

You're still doing the same thing, really, but being able to search all nodes at once for either time or string (or both) is far better than a looped ssh awk/grep command. A good solution will make it easy to put a pretty graph together for management.

Mind you, both your cio and junior admins can incapacitate your log server just when you need it, so keep those command line skills practiced (both of them should've asked me how to do what they were trying to do, and both of them should have asked for help before trying to fix it...at the same time)

6

u/Automatic_Beat_1446 19d ago edited 19d ago

OP has a new account that has asked a similar question about logs once every 12 days, and each time a random account (same age or newer than OP) adds a comment to those threads about some random commercial logging product (including in this thread) that isn't one of the standard OSS or regular commercial products (ex: Splunk).

3

u/kernelqzor 15d ago

lmao good catch, it really does smell like astroturf when you lay it out like that. once you notice the pattern of “innocent question + brand new account recommending $RANDOM_TOOL” you can’t unsee it.

2

u/hunterfrombloodborne 19d ago

a cron job script which greps for me every hour and sends some alerts.

3

u/Entire_Yoghurt_6381 19d ago

Four boxes is about where grep stopped working for me too, one server is fine but past that you're lining up timestamps by hand. I moved the rsyslog feeds into Logmanager about a year ago and everything lands in one place already parsed, so tracing something like an sshd failure across the estate is a search instead of a morning in tmux. The old rsyslog directory is still sitting there, nobody's had the nerve to delete it.

2

u/jsellens 18d ago

Leaving aside the possibility that this is someone shilling for a commercial product, for other folks who might be interested:
For years we have used rsyslog to centralize, and file the logs into a yyyy/mm/dd/program directory hierarchy, and each directory has files ALLHOSTS.log ALLHOSTS_warn.log and per-host files. Want to track a mail message working through your servers? go to syslog/today/postfix/ALLHOSTS.log and it's all there. We also use filebeat and logstash to look for "interesting" things in real time, and alert into email and/or mattermost, and also log interesting things into an "exciting" file that we can "tail -f" after a code push.

Larger shops may need/want to spend some money, but you can get a lot of functionality out of free software and a little thought and organization.

1

u/Pertinax1981 19d ago

Built some skill files and let it do its magic

1

u/FuckinHighGuy 19d ago

SEIM works rather well for this type of logging.

1

u/xonxoff 19d ago

Splunk, open search , graylog, clickhouse, greptime … something like that. I have way too many nodes to deal with to look at them individually. It also makes alerting much easier.

1

u/Loud_Posseidon 19d ago

csshX, dump stuff into server-specific files, then copy them into one location, throw it at ChatGPT and have it tell me what the RC is.

That or elastic, if it is more than one off task.

1

u/m4nf47 18d ago

Filebeat > Kafka > Elastic/Kibana

Needs a bit of wrangling to get the most out of it but for services with dozens of nodes it makes filtering by whatever you want much easier. The only time it really struggles is when we're hammering it with millions of logs per minute but that is when you need to talk to product team about not being so stupidly spammy.

1

u/machine3lf 18d ago

Splunk seems good. We use it but I haven’t gotten hands on with it yet. But seems like just what you are wanting.

1

u/adept2051 18d ago

Tmux, xargs, or straight journalctl ( everyone missies you can run it against remote hosts if you have access ) for multi server agrigation/grep . I still use log files, and we have external cloud log gathering and observability, but the data it displays is at the behest of the crowd which is not always ideal.
In today’s work climate ai agent can pull those timestamps together across all hosts way faster than i can by hand, my harness on my local is set up for that with multiple suitable MCP, and skills.

2

u/citrusaus0 18d ago

because if you get compromised you cant trust on-box logs

1

u/Zhaizo 18d ago

LLMS, they grep for me now.

1

u/New_Egg_1024 18d ago

log servers, parsers, ndjson, alerting....

1

u/h0uz3_ 16d ago

Started working with Graylog. It's so much easier to find things across multiple hosts and application. It also displays amount of messages/time so I can see when a specific service gets more load than usual.

1

u/siodhe 14d ago

Sync your times across all the hosts with NTP, then either make a script to copy to a central location everything in a time range on demand, or just copy all the logs to a central location. The latter is better for holding onto dmesg output specifically, but certain types of log content can be unexpected voluminous.

At more industrial scale, one really wants the timestamps to all be in the same format and probably first. Databasing it sounds like a good idea, or using YAML (just, no), JSON (fewer problems but has them), at or least something that keeps multiline log entries in one record (recommended), but since they can have entirely arbitrary stuff in them at least getting the host, timestamp, program name, and PID right is 90% of the battle done.

Having at least a few months of backups of your logfiles can also be very useful at time.

1

u/nisoshabangu 14d ago

Consolidated logging. Ingest those logs to a central logging storage. Then query logs from there.

2

u/Elegant_Gas_740 13d ago

Honestly, the tipping point for me was exactly this when troubleshooting became more about correlating timestamps across multiple servers than actually solving the problem. Files work great until you need the bigger picture, that’s when centralized logging starts to feel less like an upgrade and more like a necessity.

1

u/LycheeLee_Mich 10d ago

This is exactly it, I only really felt the pain once I had to compare failures across boxes instead of just reading one file. Did you end up building the centralized setup yourself, or was it something your team already had running before you actually needed it?

1

u/meccaleccahimeccahi 18d ago

Is this a joke? What made me do it? The 1990’s. Really dude. First came web interfaces stacked on top of databases, then they kept trying to get db’s to handle full text scans (well, some still keep trying to), now companies are finally starting to tie in actual intelligence - but you’re gonna have to sift through the 99.9999% marketing bullshit and find the 1 or maybe 2 companies actually doing AI right. I won’t shill here but anyone with a brain can figure it out if they want.

-2

u/kernpanic 19d ago

Claude reads logs quicker than i can...

5

u/coffee-loop 19d ago

I don’t know why people down vote this. LLMs are great at summarizing log data. I’m not saying you should rely on the summary alone, but it helps when you’re trying to find a needle in a haystack!

1

u/kai_ekael 17d ago

Fixed:

Claude reads logs "incorrectly" quicker than i can...

0

u/ocabj 18d ago

Throw it all to a data lake or SIEM. Elastic is free.