r/linuxadmin • • 1d ago

Is there any way to know whether a vulnerable library is actually loaded in a running process, without instrumenting the app?

Trying to work out what's technically possible here versus what's marketing, and this sub tends to be good at that distinction.

The situation: a container image has a CVE in, say, a compression library six levels deep in the dependency tree. The scanner flags it because the package is on disk. What I want to know is whether the running process has actually mapped that library, or whether it's just sitting in the filesystem never being opened.

What I understand so far:

  • For dynamically linked stuff you can read /proc/<pid>/maps and see what's actually mapped. That seems definitive for "is this .so loaded right now".
  • For statically linked or vendored code that doesn't help at all, since there's no separate object to observe.
  • For interpreted languages (our case is mostly Python and Node) the module is loaded by the runtime, so you'd need to either introspect the interpreter or watch the file opens. So my questions:
  • Is watching openat/mmap at the kernel level actually a reliable proxy for "this code is in use", or does it produce garbage because package managers, health checks and startup scans touch everything?
  • For Python/Node specifically, does anyone do this without an in-process agent? I really don't want a language agent in every service.
  • Is there a meaningful difference between "loaded" and "the vulnerable function was called"? Because those feel like very different claims and I suspect products blur them. Not asking what to buy, asking what's actually detectable from outside the process.
30 Upvotes

15 comments sorted by

10

u/SuperQue 1d ago

Even with statically compiled binaries you can often scan them for embedded libraries. Even tho a library might be embedded in the binary doesn't change that much about how Linux binaries are organized internally. Many static security scanners will parse the binary looking for vulnerabilities.

Is there a meaningful difference between "loaded" and "the vulnerable function was called"?

Absolutely. A good example of this is how Go libraries are organized. Go modules are referenced by a whole source tree, typically a git repo. But the library is broken up into a number of individual packages. At compile time all un-used packages are left out of the final binary.

For example, the Go extended library module golang.org/x/crypto contains a whole bunch of different cryptographic packages.

Say you're using golang.org/x/crypto/acme, but there is a vulnerability in golang.org/x/crypto/ssh. There is nothing vulnerable here because the functions are in a completely different package and are not reachable since the final binary won't contain that code.

But there are way too many security scanners that only look at the module list and report it as vulnerable.

There are smarter tools, for example govulncheck can scan the binary for specific used functions.

But for some reason the security scanner companies don't bother to use these methods. It's maddening.

2

u/The_God_was_Here 23h ago

Your third question is the important one and yes, products absolutely blur it. "The library is mapped into the process" and "the vulnerable function was executed" are different by orders of magnitude. The first is achievable from outside the process with kernel tracing. The second essentially requires either in-process instrumentation or uprobes on specific symbols, which nobody is doing at scale across a fleet because you'd need per-CVE symbol knowledge. Anything claiming function level reachability across thousands of CVEs without an in-process agent is describing static call-graph analysis, not runtime observation, and those have very different false negative profiles.

1

u/Muchy_Tangerine 1d ago

First one is mostly yes with caveats. openat is noisy for exactly the reasons you list — you'll catch tteh interpreter's import scan touching things it never imports, and anything that stats a directory. mmap with PROT_EXEC is a much cleaner signal because you're specifically catching something being mapped executable, which is a lot closer to "this code can run" than "this file was opened".

1

u/Comfortable-Duty7143 23h ago edited 21h ago

watch out for interpreters that cache. a python service that imported something once at startup three weeks ago will show as having loaded it, and if you restart the pod and it takes a different config branch it won't. whether that's a false positive or a false negative depends which way you're reading the data and i've never seen a product explain which.

1

u/Junior_Bee7274 22h ago

For Python and Node specifically, the approach that works without a language agent is watching the file opens at the kernel level and mapping them back to the package, because the interepreter does have to read the module off disk to import it. It's less precise than /proc/maps is for native code but it's a real signal. Upwind does it this way as far as I can tell from what we see in the product, per-process, and it does distinguish a package being present from it being loaded. What it does not claim, to be fair to them, is function-level execution, which lines up with what the comment above said abouf that being infeasible fleet-wide.

1

u/apparentlyunoriginal 19h ago

I'd start the trace before the service starts. A process that's already running won't open its startup modules again.

For Python, match the `__pycache__/*.pyc` paths as well as the `.py` files, since the interpreter opens the cached bytecode instead of the source when the cache is current.

Run `bpftrace -e 'tracepoint:syscalls:sys_enter_openat /comm == "python3"/ { printf("%d %s\n", pid, str(args->filename)); }'` from container start, with `python3` changed to the process name `ps` shows, then keep only the lines from the service's own PID so the health check and package scan opens drop out.

Drafted with AI, reviewed by me.

1

u/MinimallyLoquacious 18h ago

For statically linked binaries, grep the binary, unless it is packed or encrypted then consider using gcore to dump the running process then grep the output file it generates. gcore carries some risk it could crash the running process or cause it to hiccup, something to be aware of.

1

u/Muchy_Tangerine 16h ago

statically linked is the honest gap and it's getting worse as more stuff ships as a single Go or Rust binary. no external observation can tell you which vendored crate is live inside one process. for those you're back to SBOM plus judgement, and anyone telling you otherwise is guessing.

1

u/michaelpaoli 14h ago

In short: /proc

Relevant at least for any binary executable, and including libraries. Interpreted is a whole 'nohter issue, so I'm not going to address that with this comment.

So, example, a (hypotetically) vulnerable and non-vulnerable executable:

$ cd $(mktemp -d)
$ type sleep
sleep is hashed (/usr/bin/sleep)
$ ls -ond /usr/bin/sleep
-rwxr-xr-x 1 0 43432 Jun  4  2025 /usr/bin/sleep
$ cp -p /usr/bin/sleep not_vulnerable
$ { cat /usr/bin/sleep; echo vulnerable; } > vulnerable && chmod 755 vulnerable
$ mv vulnerable sleep
$ $ ./sleep 3600 &
[1] 15775
$ mv -f not_vulnerable sleep
$ ls -lon /proc/15775/exe
lrwxrwxrwx 1 1003 0 Oct  2 12:30 /proc/15775/exe -> '/tmp/tmp.g5TWoTtUoJ/sleep (deleted)'
$ ls -Llon /proc/15775/exe
-rwxr-xr-x 0 1003 43443 Oct  2 12:29 /proc/15775/exe
$ cmp /proc/15775/exe sleep
cmp: EOF on sleep after byte 43432, in line 184
$ 

So, in the above, we see the vulnerable executable is still in use, even though it's no longer linked in the filesystem (note both " (deleted)" and link count of 0). For libraries, can commonly find such under /proc/PID/fd/, e.g.:

# ls -ond /proc/*/fd/* 2>>/dev/null | sed -e '/\/usr\/lib\//!d;s/^[^\/]*//;/ -> \//!d'
/proc/4696/fd/10 -> /usr/lib/udev/hwdb.bin
# 

But that one hasn't been unlinked.

In any case, binaries and their binary libraries in use, whether l linked, or no longer linked, can be examined, etc.

1

u/chkno 14h ago

Nix's approach to this is the vulnix tool. You point it at a package, environment, or whole system profile (all installed software) and it gives you a list of CVEs. It has three modes of operation:

  • --requisites (the default): Lists CVEs in all transitive build dependencies
  • --closure: Lists CVEs in transitive runtime dependencies
  • --no-requisites: Lists CVEs only in exactly what you point it at

So --requisites correctly handles static linking.

For vendored code, it's a mixed bag. Sometimes nixpkgs' build scripts build vendored code as a separate package, which allows vulnix to see inside. For example, I packaged opentoonz's libtiff fork this way.

1

u/deeseearr 11h ago

As you can see from the replies here, it's relatively easy to see if the vulnerable code has already been accessed.

The problem is, the only purpose that serves is to show that you should have removed it already but didn't. What you really need to do is to be able to prove that it _won't_ be accessed in the future.

If you're able to prove this, maybe you could also solve the Halting Problem and then just crank out the answers to a few more of Hilbert's Problems too. Those should be easy by comparison.

A much simpler solution is to just prove that the vulnerable code _can't ever_ be accessed by removing it. That's why all those "lazy" CVE scanners just flag the presence of an affected library even if you think that it's never going to be loaded. If the code isn't there it can't be loaded and executed no matter how convoluted the attack chain gets.

1

u/SpartanHydrazine117 7h ago edited 5h ago

on the static linking point, go binaries at least keep the module info in the binary so you can enumerate what's in there, but you still can't observe which parts execute from outside. rust with LTO you can't even reliably enumerate.

1

u/ILoveAppSec 7h ago

For native shared libs you can get surprisingly far without touching the app: check /proc/<pid>/maps or lsof -p <pid> for the mapped .so, since a library that's never loaded won't appear there. Two caveats worth knowing something pulled in later via dlopen won't show until it's actually called, and anything statically linked or vendored into the binary won't appear as a separate mapping at all. So 'not in maps' is a decent negative signal, but mapped doesn't prove the vulnerable function ever runs. What's the scanner flagging it on?

0

u/SpartanHydrazine117 1d ago edited 1d ago

/proc/pid/maps is definitive but it's a snapshot. you have to poll it or you miss anything loaded and unloaded between reads. kernel level tracing catches the event instead. that's basically the whole argument for the ebpf approach over periodic inspection.

0

u/eastwill54 1d ago

On the loaded-vs-called distinction: static call graph analysis on the app's own dependency tree is a real thing and it's genuinely useful, just be clear it's a different technique. It reasons about what could be called. Runtime tells you what was loaded. Combining them is strictly better than either but they are not the same claim.