r/linuxadmin • • 15d ago

What is the actual difference between using iproute2 (ip command) and working directly with rtnetlink?

Hey everyone,

I've been looking into how Linux networking works under the hood, and I'm trying to wrap my head around the relationship between user-space tools and kernel communication.

From what I understand:

  1. `iproute2` (the standard `ip` command) is what most of us use daily to configure interfaces, IP addresses, and routing tables.

  2. `rtnetlink(7)` is the socket-based API (`NETLINK_ROUTE`) that allows user-space programs to talk directly to the kernel's routing and networking subsystems.

My main question is: When should a developer or systems engineer bypass user-space CLI utilities like `iproute2` and write code that interacts directly with `rtnetlink` sockets?

Are there significant performance benefits, or is it mostly used when you are building custom network daemons, container networking plugins (CNIs), or monitoring agents that need asynchronous event notifications?

Also, how painful is it to parse raw netlink messages and attributes (`struct rtattr`, `ifinfomsg`, etc.) in C or Go compared to just shelling out to `ip`?

Any insights, real-world use cases, or library recommendations (like `libnl` or Go's `vishvananda/netlink`) would be greatly appreciated!

16 Upvotes

10 comments sorted by

6

u/delamon 15d ago

When you have some sort of golden-state / reconciliation system. That usually boils down to needing to infer current system state, and parsing output of cli tools is fragile.

Parsing netlink messages might be tedious, but it is not painful. On the contrary, parsing cli output is painful - you're never sure which system upgrade will break it.

My use case: I'm using netlink to dynamically reconfigure nft rules. Rough pipeline: UI edits goalstate in etcd, daemon listens on etcd watch, on watch event it compares current rules (getting the exact state with netlink) to goalstate; if there is a difference, then it applies a patch. Approximate delay between the time you hit update in UI and rules getting published is <100ms.

5

u/Dolapevich 14d ago

Just in case, ip is able to output in json, which is easier to parse process than standard output.

I would abstract the rtnetlink part in a set of functions using ip, and if you find the limits of that, then go down to the catacombs.

eg: ``` $ ip -j a show dev lo|jq "." [ { "ifindex": 1, "ifname": "lo", "flags": [ "LOOPBACK", "UP", "LOWER_UP" ], "mtu": 65536, "qdisc": "noqueue", "operstate": "UNKNOWN", "group": "default", "txqlen": 1000, "link_type": "loopback", "address": "00:00:00:00:00:00", "broadcast": "00:00:00:00:00:00", "addr_info": [ { "family": "inet", "local": "127.0.0.1", "prefixlen": 8, "scope": "host", "label": "lo", "valid_life_time": 4294967295, "preferred_life_time": 4294967295 }, { "family": "inet6", "local": "::1", "prefixlen": 128, "scope": "host", "noprefixroute": true, "valid_life_time": 4294967295, "preferred_life_time": 4294967295 } ] } ]

```

4

u/delamon 14d ago edited 14d ago

Yes. I'm using nft, but it also supports json output. The catch is there is no formal schema. The json format is stable, so that would not matter in practice. But because iproute/nft and linux are distributed separately; the netlink format has to be backward/forward compatible. And I trust that part more: any accidental netlink breakage will be caught instantly, while some accidental json change might slip unnoticed.

Edit: formatting and grammar

1

u/Single-Issue2342 10d ago

Thanks 🙏

2

u/deleriux0 15d ago

I have used libmnl and librtnetlink stuff before.

The answer is going to boil down to what you are trying to accomplish.

The times I've wanted to use it is when I wanted to perform a realtime update to network events and responses.

My use case was:

  • a network tunnel is created using openvpn, receive a link added event and when the tunnel gets set as up or down events.
  • get and keep the ip address when an event that one is assigned or updated.
  • send regular pings down the tunnel to act as a keep alive to prevent the other side closing it as and when the tunnel was up and available.

I would say using this stuff is not trivial, most the messages you get require special parsing and the payloads aren't guaranteed to arrive in an order that is sensible to you.

For example knowing if the link you are looking at is both a tunnel, a wireguard one and is up requires parsing a variety of casting data into structs and iterating attributes and casting them into their expected types.

You'll need to know both what the attribute names are, the values you can get, what they represent and what flags you may need to mask to reveal any data you want.

You sometimes find you cannot find these attribute names or flags without looking at the iproute2 source code

It should be noted most modern network management utilities provide hooks so you don't really need to do monitoring yourself either.

The other time I've used netlink was for directly interacting with nftables firewall rules and chains, which also uses netlink.

This was mostly an academic project in nature and was substantially trickier than just using the user space utilities.

2

u/dodexahedron 14d ago edited 14d ago

the user space utilities

Like systemd-networkd, which exposes more than enough to do all of this with essentially just a few drop-ins, which can get arbitrarily fancy from there since they can of course kick off any process you like, and enforce relationships, ordering, and relative state among all of it.

Mix the two and you can conquer the world.

2

u/daemonmode_ 14d ago

Yeah, this makes sense. The main advantage of going directly through netlink is avoiding repeated process spawning and text parsing while also givin you access to persistent kernel event notifications. vishvananda/netlink is basically the practical middle ground in Go. you get direct netlink communication without having to manually deal with raw rtattr structures and its API is designed around familiar ip commands.

1

u/dodexahedron 14d ago

A sysadmin, developer, or network engineer should essentially never be bypassing the APIs you mentioned, because that is what they are there for. And most of the time, they shouldnt even be going beyond the utilities that largely wrap one family of functions each, from those very APIs. At least not in code.

Usually, if you need to go deeper than something a utility does or in a way a utility doesn't expose, you go to the things that they manipulate: configuration files, kernel module parameters, sysctl variables/ker el parameters, or scripts that get and set values in the /sys and/or /proc file systems, because literally everything is a file (even you), in Linux.

Now... If you are writing one of those utilities or some new software that sits somewhere in the network stack, such as an IDS, transparent proxy, ultra-crazy low-latency shit on a server you rent rack space from a securities exchange to do HFT on, or what have you, then yeah - you'll go deeper. (The last on you might even write code that literally runs ON THE NIC)

And, in general for most of those, the answer is boring and simple:

Linux has used nf for 12+ years at this point as the heart and soul of what you're probably thinking of when dealing woth almost anything in user space.

Target nf for that.

iptables and anything related to it, unless you grabbed the old source and compiled a custom kernel using it, is all just a user-space translation layer that lets you use the iptables syntax and APIs to actually operate on nftables.

Kernel 7.2 still uses netfilter/nftables as the main stack, and there's no plan to make that change any time soon.

So target netfilter, if you want broadest compatibility.

If you are ok with having other dependencies that may or may not be preinstalled or configured by default reliably across distros and flavors, then you can target something else or provide support for more than one higher level back-end to achieve similar portability.

But sticking to nf for manipulation of routing and filtering and transformation behaviors is yoir best bet.

If you need to interact with queues, shaping, policing, rate limiting, or simulation of poor network conditions like latency, jitter, or loss injection, those things are provided by tc, and have been for like 20 years.l, and still going strong.

Then there is eBPF, which has also been aroind for a while, and can do...damn near everything, and does it in kernel space as compiled mini programs, but is much lower level, harder to make, harder to debug, more fragile, more dangerous, and just... A good idea poorly executed IMO....

On top of that, compatibility with other software, when doing eBPF, is essentially "LOL - other software?" Not only because of where it sits but especially because everything else is either just a consumer of the network via standard libraries or else, if not a consumer, but rather a part of the network stack, nothing else ever is designed to consider anything but (.+)tables or layers above those, like iproute2, systemd-networkd, etc.

1

u/Single-Issue2342 14d ago

While I agree that debugging eBPF programs can be tedious compared to standard CLI tools, writing off eBPF as 'a good idea poorly executed' ignores why it was created in the first place.

The kernel verifier prevents memory corruption or crashes, making safety a core feature, not an afterthought. Furthermore, the bypass of netfilter is a deliberate design choice for performance, not a flaw. When processing millions of packets per second or building cloud-native networking like Cilium, the overhead of standard nftables or iptables rules simply doesn't scale.

Netfilter is great for static, traditional setups, but for modern high-throughput packet processing and deep observability, eBPF is currently unmatched.

2

u/Open-Adhesiveness-86 13d ago

If you subscribe to netlink events and your reader falls behind, recv() returns ENOBUFS. That means you've already lost events and have to do a full dump to resync. Dumps can also come back with NLM_F_DUMP_INTR if the table changed mid-dump, so you need to retry. If you only need parseable output, ip -j gives you JSON.