Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Proxmox
    VMware
    Migration
    ZFS
    Ceph
    Storage

    Why Your Proxmox Migration Failed (Hint: It Wasn't Proxmox)

    January 23, 2026
    9 min read

    There's a familiar moment happening in a lot of IT meetings right now.

    Someone pulls up a spreadsheet with the licensing costs circled in red. Broadcom gets mentioned, usually with a sigh. And then someone asks the question that sounds obvious, almost inevitable:

    "Why don't we just move everything to Proxmox? It's basically free VMware."

    On the surface, it makes sense. Proxmox looks ready, the feature list checks out, and plenty of people are running it in production without issue. Meanwhile VMware vSphere has gone from default choice to uncomfortable budget conversation.

    So the decision gets made. VMs are converted, hosts are built, and the cutover happens. And then the weird stuff starts.

    Databases fall over under load. Latency jumps for no obvious reason. Entire nodes feel sluggish even though monitoring says everything is "fine." Before long, the conclusion lands:

    "Proxmox is unstable."

    "Open source isn't enterprise-ready."

    "We should've stayed on VMware."

    That isn't what actually happened, though. The thing that failed was an assumption VMware spent 15 years teaching us to make.

    VMware made storage invisible more than easy

    VMware's biggest accomplishment was insulation, even more than virtualization itself.

    VMFS, SAN integrations, vSAN, and years of engineering work smoothed over the sharp edges of storage. Locking behavior, alignment issues, and cache pressure all lived behind a clean abstraction layer. You didn't need to think about it most days, and that was the point.

    That didn't make admins lazy. It made the platform successful. But it also meant a lot of teams learned which buttons to click without learning why things worked. You sized hosts, followed best practices, and trusted the system to absorb the complexity.

    Proxmox doesn't play that role. Instead, it hands you the raw materials: ZFS, Ceph, LVM, XFS. These are powerful tools with fewer guardrails and much more honesty. If something goes wrong, you're going to see it and feel it.

    That difference alone is enough to derail migrations that look perfectly fine on paper.

    ZFS and the memory you thought you had

    One of the most common post-migration stories sounds almost boring at first. A VM crashes. There's plenty of RAM installed, the host isn't overloaded, and nothing obvious is wrong, so Proxmox gets the blame.

    Dig a little deeper and ZFS is usually sitting right there, doing exactly what it was designed to do.

    ZFS loves memory. It uses ARC to turn spare RAM into faster reads and better performance, and if you don't tell it otherwise, it will take what it can get, which can be a lot.

    If you migrated from ESXi and sized your hosts the same way, that's where things start to unravel. On VMware, that memory belonged almost entirely to guests. On Proxmox with ZFS, the filesystem is now competing for it too.

    When pressure hits, the Linux OOM killer doesn't care that your SQL VM is "important." It just sees a system under stress and makes a decision, usually the wrong one from your point of view.

    Yes, newer ZFS versions are better about releasing memory, and ARC is now more reclaimable than it used to be. The core lesson still holds: if you don't cap it intentionally, you're trusting adaptive behavior during the exact moments when predictability matters most.

    That's storage doing storage things, and calling it a Proxmox bug misses the point.

    Ceph on 1GbE: technically possible, practically a disaster

    Then there's Ceph, or rather, how people try to run it.

    Ceph is distributed storage, which means your disks are constantly talking to each other over the network. When everything is healthy, it's manageable. When a disk fails, that chatter turns into a firehose, and on a 1GbE network, that firehose becomes a brick wall.

    The deceptive part is that the cluster doesn't necessarily go down. It stays "up," and health checks might even look okay. But latency skyrockets, rebuilds crawl, and every VM feels like it's wading through mud. From the outside, it looks like Proxmox fell apart, when in reality the network did.

    Ceph documentation has been blunt about this for years: 10GbE is the minimum, the floor rather than a suggestion. In real production environments, 25GbE is what actually gives you breathing room when things break, because things always break.

    Running Ceph on 1GbE because "it works" is like discovering your car can technically drive with the parking brake on. You'll get moving. You just won't like what happens next.

    The migration mistake that goes unnoticed until it's too late

    Most migrations follow the same script: convert the disk, import the VM, boot it up, and if it starts, call it good. That's how you end up with systems that feel inexplicably slow.

    If you move virtual disks without paying attention to sector alignment, especially when crossing from legacy 512-byte layouts into 4K-backed storage, every write can turn into extra work. One logical write becomes multiple physical operations, IOPS drop, and latency climbs.

    Nothing crashes and nothing throws an error. Performance just degrades in the background. And because the hypervisor changed at the same time, the blame lands there instead of on the invisible storage mismatch dragging everything down.

    VMware absorbed a lot of that pain for you. Proxmox doesn't. It assumes you know what you're doing, and if you don't, it won't stop you; it'll just let you learn the hard way.

    "Free" was never the real cost

    The part people don't like to say out loud is that removing VMware licensing relocates complexity instead of removing it.

    With VMware, you paid money to avoid thinking about certain problems. With Proxmox, you save money and take those problems back in-house: memory behavior, network saturation, disk geometry, failure domains, and rebuild dynamics.

    That trade can be fantastic, and plenty of teams are making it work. They're not doing it by pretending Proxmox is a drop-in replacement for vSphere, though.

    If you want set-it-and-forget-it storage, buy a SAN. If you want performance, buy RAM and tune ZFS. If you want hyperconverged infrastructure, build a real network and accept what that implies. What you can't do is assume nothing else needs to change.

    A reality check more than a Proxmox problem

    VMware didn't make engineers worse. It made certain mistakes harder to see, and Proxmox removes that cushion.

    That's uncomfortable, especially when migrations are rushed and expectations are unrealistic. But the failures people are seeing don't prove that Proxmox isn't ready. They show that abstraction hid more than we realized.

    The teams that slow down, relearn the fundamentals, and design intentionally will be fine. The ones that rush, convert disks, and hope muscle memory carries them through will keep blaming the wrong thing.

    Your migration failed because Proxmox stopped lying about how your infrastructure actually works, and that can be a rude awakening.