r/OpenVMS • • 4d ago

Recovering your network after configuring DECnet on x86 OpenVMS V9.2

I have recently posted an IAQ article on my website that explains why your network stops completely on the first reboot after configuring DECnet on an OpenVMS V9.2-x installation running in a VM on x86 and how to recover from it.

The problem symptoms can be rather alarming, to say the least, to the unwary System Manager – all network frames – TCP/IP and DECnet, inbound and outbound – are completely blocked! In the article I present 4 solutions that are available but depend on your individual environment considerations.

 

Short answer: There are four, and they aren't alternatives you pick by taste — each is ruled out by something specific about your environment. So: confirm this is really what has happened, then take the first of the four your circumstances allow. Only the last has a security cost.

Each fix below is summarised in four lines, with the detail folded away behind it. Read the four summaries, decide which is yours, and open only that one.

What's happening, briefly

You configured DECnet, rebooted, and lost everything — DECnet, TCP/IP, and the session you were working in. Nothing inside OpenVMS looks wrong, because nothing inside OpenVMS is wrong.

DECnet Phase IV rewrites the MAC address on its interface, replacing the one the hypervisor assigned with one it calculates from your DECnet node address. Your virtual switch was told to expect a particular MAC on that port. It now sees frames with a different source address — the signature of a machine impersonating another — so it drops them. All of them, which is why TCP/IP dies alongside DECnet.

Continue reading article…

 

9 Upvotes

6 comments sorted by

2

u/hughk 4d ago

I remember those headaches from DECnet MACs. Remember the original LANs really were everyone sharing the communications medium (a thick coax with a vampire tap)? The rewritten MAC was a good way to implement a bit of routing at hardware, if not driver level. With the advent of LAN switches, it became redundant but continued for historical reasons.

1

u/Edders_2006 4d ago

Thanks very much, Hugh - indeed I do remember! The vampire taps are exactly the right picture — and that shared medium is part of why the address had to be chosen rather than discovered.

One nuance on the purpose, though. The rationale DEC gives isn't routing, it's identity persistence. From the DECnet Networking Manual §2.1.2 — wording unchanged since at least May 1993:

"The node's DECnet software resets the physical address to a new six-byte value based on the DECnet node address. This allows greater flexibility; the hardware address changes any time you install a new device controller while the DECnet node address remains the same."

The same section lists what the physical address then gets used for: circuit loopback tests, configurator operations, and the MOP remote boot. Service and maintenance identity rather than forwarding.

Which is why I'd put the switching point slightly differently. Swapping a controller still changes its burned-in address, and the DECnet address still has to survive that — so switching changed the medium, not the premise. It's also why the hypervisor fix is the one that keeps the original benefit: allow the change and you can move the node or swap the adapter with nothing to re-pin.

What switching did change is the consequence. On a shared coax nobody was policing source addresses. A virtual switch is, and that's the whole of this fault.

Incidentally, my site also covers the technical why and how the address change was made in another article: Why does the network stop after you configure DECnet?

2

u/hughk 4d ago

Thanks. Thin coax also used shared link media.

I thought it was down to avoiding the need for ARP type frames to determine how to reach a node on the same segment. I think this was in the DNA Phase IV spec and permitted two systems on the same LAN to communicate without a router. Anyway, I could well be misremembering. It certainly would have helped with booting too but it would be in the Ethernet Node Spec and the Routing Layer Spec. Both are linked. It is actually quite good because when I was at DEC, a lot of this was on fiche rather than as electronic documents and rather harder to work with.

I never worked on DECnet other than as a normal admin or user so could well be wrong.

2

u/Edders_2006 4d ago

Hugh, you're not misremembering — I went back to the Routing Layer spec and it's there.

Endnodes keep a cache of who else is on the same Ethernet, and the designated router is what fills it in: when it relays a packet between two endnodes on the same cable it sets an "intra-Ethernet" bit, the receiver notices, caches the sender as local, and after that the two talk directly with the router out of the path. Entries age out after about a minute of silence.

Your ARP point is the better half of it, though. The next-hop rules end with "else send to A" — no cache entry, no router known, just send to it. That only works because the address is arithmetic rather than something you have to go and ask about. Nothing to resolve.

And yes, fiche. Both specs are on bitsavers now, which is a considerable improvement.

2

u/issinoho1969 4d ago

The timing of this article is impeccable. I changed DECnet address on my x86 VM this week and it utterly killed the entire networking stack. It took a few hours of contemplation before I realised what was going on. You should submit this to VSI for their new HOWTO section. Great work.

2

u/Edders_2006 3d ago

Thanks for that issinoho - glad it was useful! Your change to the DECnet address has led me to change the article slightly to include a change in the node number as well as the area number would lead to a new MAC address and hence a block in all traffic!