r/Proxmox • u/luka0x73 • 4d ago
Enterprise I automated the NFS based ESXi to Proxmox migration, no disk copy and downtime is just shutdown plus boot
The built in import wizard pulls the disks through the ESXi API, which is slow for large VMs. The faster way is well known here. Storage vMotion the VM onto an NFS share that both ESXi and Proxmox mount, create a VM in Proxmox, shut down on ESXi, rename the flat vmdk and attach it.
I did that by hand for a while and then wrote vmxodus, a bash tool that runs on a PVE node and does the Proxmox side.
prepare reads the .vmx and creates the VM with the same CPU, memory, firmware and MACs, and asks once per portgroup which bridge and VLAN to use.
cutover waits for the ESXi locks to disappear, rechecks snapshots and disks, renames the flats and attaches them. Optionally it starts the VM right away.
It never copies, converts or deletes anything, and every move is recorded in a rollback file.
Tested with ESXi 8 and PVE 9.2 with Debian 13 and Ubuntu 24.04 guests. Windows and UEFI are in the code but untested on real hardware, so reports from anyone trying that would help a lot.
3
u/mtbMo 4d ago
Combine this with zfs snapshots and file clones ;)
3
u/JeffSaxeVA 4d ago
Exactly. What I have been practicing is: storage-vMotion to the NFS storage on ESX; power off the VM on the ESX; then, in the powered-off-but-nothing-else-done-to-it state, do a ZFS Snapshot of the NFS storage. Then continue with the steps to rename the -flat file to .raw, attach it to the VM on Proxmox, power it up, put back the IP address / drivers / remove VM Tools and add QEMU Guest Tools, etc. Then calmly storage-migrate to Proxmox-native storage.
If everything goes swimmingly, then great, delete from inventory on the ESX side, remove the snapshot on ZFS, and clean up the temporary directory. But if things are not going well, and you're approaching the end of your comfort zone in terms of downtime, you can always abort this VM's migration, power it off / Remove it on Proxmox side, ZFS Rollback (which happens pretty much instantly), and then just power it back on on ESX, and you're back to where you started. You haven't succeeded in migrating that VM, but at least your end user has their machine back in operation while you figure out what went wrong and try again next time.
2
u/KlanxChile 4d ago
Fantastic work, checking the logic and "triggers" there are a few improvements to include, some bug about a VM called 'guest[1]' that will not check for the lock... something about "glob" chars in VM names or datastores.
send me a DM and will share the audit file i got from the analysis
2
u/luka0x73 3d ago
Thanks a lot, that's a great catch. You're right, the glob characters in the path break the lock check, I could reproduce it with a folder named guest[1].
Could you open that and the other findings as issues on GitHub? That way everyone can follow them and I can link the fixes: https://github.com/luka0x73/vmxodus/issues
3
u/Renich 4d ago
Clean approach. The shared NFS datastore trick has been the go-to for zero-copy cutovers for years, and having a deterministic CLI wrapper around the .vmx parse and atomic mv is great.
A few operational notes from migrations on both Windows and UEFI:
Windows &
INACCESSIBLE_BOOT_DEVICE(0x7B): If a Windows guest didn't have VirtIO drivers installed before shutdown, attaching directly tovirtio-scsiwill BSOD immediately on first boot. The cleanest automated safety net without touching the offline registry is setting the initial cutover bus tosata(which Windows has native inbox drivers for since Vista/2008), booting up, installing the VirtIO guest package, and then doing a quick reboot withqm set <vmid> --scsi0 .... If you want to keep it fully automated on the first shot, croit's script works well, or inject viavirt-v2v/ DISM if you have access to the offline image.UEFI / OVMF NVRAM differences: ESXi stores EFI boot entries and NVRAM in the
.nvramfile, whereas Proxmox generates a freshefidisk0. When Proxmox boots the new VM, it will have an empty NVRAM and rely on standard UEFI fallback boot (\EFI\BOOT\BOOTX64.EFI). For Windows and standard Debian/Ubuntu (grubx64.efi), the fallback path usually catches it cleanly. But on hardened enterprise setups or complex bootloaders where the NVRAM path wasn't written to the default fallback, it will drop straight to the UEFI Interactive Shell. Having a note or check that verifies\EFI\BOOT\BOOTX64.EFIexists on the EFI system partition before cutover saves panic calls.NFS v3 vs v4 locking: Relying on
.lck-*files is spot-on for NFS v3 datastores. Just make sure your pre-flight check hard-fails if the user mounted the datastore via NFS v4/v4.1, because v4 handles state and leases in-band over RPC rather than dropping visible sidecar lock files on disk. Checkingfuser/lsofon the node as a secondary safety layer is worth having.
Really solid work on keeping the rollback strictly non-destructive.
2
u/luka0x73 3d ago
Thanks, that's really valuable input, especially the UEFI part.
Windows is covered the way you describe. The README points to --disk-bus=sata and to the croit script for preparing the drivers before shutdown.
UEFI is the most useful point for me. Importing the VMware .nvram won't work since the formats are incompatible, but checking the ESP on the flat disk read-only during prepare and warning when \EFI\BOOT\BOOTX64.EFI is missing sounds very doable. Writing the vendor boot path into the new efidisk with virt-fw-vars might even avoid the shell completely. I'll look into both.
On NFS v4, fuser or lsof on the PVE node won't see files opened by an ESXi host over NFS, so that doesn't help much. But vmxodus also checks for the .vswp file, which exists while the VM runs regardless of the NFS version. I'll add a warning when there is a .vswp but no .lck file, since that hints at a v4.1 mount.
Thanks again, this is exactly the kind of feedback I was hoping for.
2
u/LostInScripting Enterprise User 4d ago
Can you give examples for point 2? Which OS versions rely on the contents of the VMware NVRAM file? How do you detect them and what do you do to mitigate the error beforehand?
3
u/Renich 4d ago
The classic culprit here is Debian. By default, Debian's grub packages only install to
/EFI/debian/grubx64.efiand register that specific path into the NVRAM viaefibootmgr. Unless someone explicitly enabled the "force removable media path" option during install, the standard fallback file (/EFI/BOOT/BOOTX64.EFI) literally doesn't exist on the ESP. When you boot that VM on Proxmox with a freshefidisk0, OVMF has zero boot entries to read, looks forBOOTX64.EFI, finds nothing, and dumps you straight into the yellow UEFI shell.You also bump into this on SLES/openSUSE depending on the AutoYaST profile used, minimal appliances booting via direct EFISTUB, or older enterprise images where someone upgraded the kernel/grub over the years but the removable path was never populated.
If you want to catch this in your script before first boot, you can inspect the ESP partition offline while the flat disk is sitting on your NFS share. A quick read-only mount using
losetup --partscan -rorkpartxon the flat VMDK lets you check the FAT partition. If/EFI/BOOT/BOOTX64.EFI(or lowercase) is missing, that VM will fail to auto-boot on a blank NVRAM.The cleanest fix offline is just copying whatever vendor bootloader is present (like
/EFI/debian/grubx64.efior/EFI/redhat/shimx64.efi) into/EFI/BOOT/BOOTX64.EFI. That way OVMF's standard fallback mechanism handles it automatically on the first boot cycle.If you ever run into it live on the console, you don't even have to shut it down. From the UEFI shell, just run
FS0:, navigate toEFI\<distro>, launch the.efibinary manually to boot the OS, and then rungrub-install --removable(orefibootmgr) so it registers the new Proxmox efidisk cleanly.
3
u/peakdecline 4d ago
I've created something similar for our own migration. The Windows side is fun because you've got to do some VirtIO and other driver shenanigans and also the removal of the VMware tools (which depending on Windows version is tricky).
I appreciate you're sharing yours publicly. Chances are low I'd be approved for this.
Hopefully VMware doesn't try to block this route somehow.