Mr.PlanB Logo

    Newsletter

    Subscribe our newsletter

    Get new infrastructure guides, comparison reports, and migration notes in your inbox.

    Infrastructure notes, guides, and new tools. Unsubscribe anytime.

    Back to Blog
    Proxmox
    Cluster
    High Availability

    Proxmox Quorum Failed With 2 of 3 Nodes: Why

    October 2, 2026
    8 min read

    A three-node Proxmox cluster with one vote per node should normally survive losing one node, since two votes out of three is a majority. If the remaining two nodes lose quorum anyway, stop counting physical servers and check how many voters Corosync thinks it has.

    That's what cleared up a confusing incident in August 2026. The cluster had three physical nodes, yet taking one offline made quorum disappear. The cause turned out to be a QDevice nobody expected, still sitting in the voting configuration. Corosync expected four votes, so two surviving nodes weren't enough. The number of servers in a cluster and the number of voters can differ.

    Should two nodes keep quorum in a three-node Proxmox cluster?

    Yes. In a normal three-node Proxmox VE cluster where each node has one vote, two surviving nodes keep quorum because two is more than half of three.

    Most administrators start from that mental model, and here the administrator's expectation was right. Three physical nodes were visible, each apparently with one vote, so one node should have been able to go away without taking quorum with it. The visible topology just didn't show everything.

    Quorum counts votes, not server chassis on a shelf. Corosync tracks the expected number of votes and works out the quorum threshold from that configuration. If there's another voting participant, even one the administrator has forgotten about, the arithmetic changes. That's how a cluster that looks like three nodes ends up behaving as if it has four voters.

    If you're still getting comfortable with the basic architecture, the Mr.PlanB Proxmox resources explain how clustering, storage, and host management fit together, which helps before you start troubleshooting subtler HA behavior.

    What did pvecm status reveal?

    The graphical cluster view didn't show the problem, but pvecm status did. With all three physical nodes online, the output reported:

    • Nodes: 3
    • Expected votes: 4
    • Highest expected: 4
    • Total votes: 3
    • Quorum: 3

    The membership section listed the three normal nodes plus a QDevice entry. That showed the contradiction right away: three servers, with the quorum system still counting against four expected votes.

    The QDevice showed as not registered, but it was still part of the expected voting configuration, and that's enough for a nasty failure mode. With all three servers up, the cluster had three votes and met the quorum requirement of three, so everything looked healthy. Take one physical node offline and only two server votes remain against a threshold that's still three. The cluster has no quorum.

    So when cluster behavior doesn't match the topology you think you built, pvecm status should be one of the first commands you run. The GUI tells you which nodes are present, and the vote output tells you what the consensus system believes.

    How can an old QDevice break a three-node cluster?

    A QDevice adds an external vote for quorum decisions. It's most often used when an even number of cluster nodes could otherwise end in a tie. Set up on purpose, it's useful. It gets confusing when the cluster topology changes and the voting configuration doesn't.

    Proxmox documents QDevice support as part of Corosync quorum handling. Its example shows a QDevice listed in pvecm status next to the cluster nodes, and the documentation explains that the QDevice can cast a vote during partition scenarios. The same documentation warns about topology changes: if you add or remove nodes in a cluster that uses a QDevice, remove the QDevice first, then set it up again afterwards if the new topology still needs one. This incident is what happens when that step gets skipped.

    The administrator said the environment had been a two-node cluster for several years before it became a three-node setup. They didn't remember configuring the mystery QDevice, so the thread never worked out when or why it was added. What mattered was that the voter registration still included it.

    Old cluster state can outlast anyone's memory of it. That's especially common in homelabs, where hardware gets swapped, nodes get renamed, experiments turn into permanent infrastructure, and a configuration from years ago can stay active long after the reason for it is gone.

    Why did the cluster work until a node went offline?

    The stale configuration was wrong, but it didn't break anything straight away. With three servers up, the system reported three total votes and needed three for quorum. That met the threshold exactly, so the cluster stayed quorate. Nothing in day-to-day operation made the administrator notice that the expected vote count was one higher than the number of physical nodes. It only showed up during failure testing.

    High availability setups often go wrong this way. A configuration can look healthy for as long as every component is up, and the mistake only shows once something fails and the system actually has to use the redundancy you thought you had. Here the three-node design was meant to tolerate losing one node, and the stale fourth expected vote silently took that tolerance away. On paper the cluster still looked redundant, but the vote math said otherwise.

    That's why planned failure testing is worth doing. Shut down one node while everything else is healthy and watch what happens. Check that the remaining nodes stay quorate, HA actions behave as expected, and storage and networking stay healthy. Until you've tested a redundancy design, part of it is still an assumption.

    How was the stale QDevice removed?

    The suggested cleanup command was pvecm qdevice remove, and it didn't finish cleanly. It reported that corosync-qdevice.service was not loaded and returned an error while trying to stop the service on one node. Even so, the administrator saw Expected votes drop from 4 to 3, which confirmed the fix had taken.

    With the cluster expecting three votes, the quorum threshold became two. Another participant pointed out that with expected votes at three, two surviving nodes should be enough. The administrator then confirmed the problem was solved and updated the original post to name the unknown QDevice as the cause. Proxmox's current administration guide documents the same removal command for a QDevice set up through pvecm.

    The failed service stop is a good reminder not to judge a maintenance command only by its exit status. After any change to cluster membership or quorum, run pvecm status again and check the state that matters: whether Expected votes changed, whether the unwanted QDevice disappeared, and whether the cluster is quorate. If the answers are right, the change worked.

    Why does Proxmox quorum care so much about majorities?

    Quorum stops disconnected parts of a cluster from each deciding they're the authoritative side at the same time. A dead server is the obvious failure. A network partition is the harder one. Two machines can both be alive and still lose contact with each other, and from either node's point of view there may be no way to tell whether the other host died or only the network path between them did. If both sides keep running the same resources independently, you can get competing instances and diverging state, which is far worse than a temporary outage. So cluster consensus is deliberately conservative. A partition has to hold enough votes to show it's the authoritative side before it carries on with normal clustered operation.

    The community thread about this incident got tangled because several people focused on the general rule that even-sized clusters can have tie problems. That rule is real, but it wasn't the cause here. Two nodes are a majority of a correctly configured three-vote cluster, and the unexpected fourth expected vote was what changed the answer. When you're debugging quorum, do the exact arithmetic instead of leaning on rules of thumb.

    What should you check before trusting a Proxmox cluster?

    Seeing every node name in the cluster interface isn't enough. Check the voting state while everything is healthy.

    Run pvecm status and look closely at four things:

    • Expected votes
    • Total votes
    • Quorum
    • Membership information

    For a basic three-node cluster with one vote per node and no QDevice, those numbers should add up in an obvious way. If you have three nodes and four expected votes, find the fourth voter before you count on surviving a one-node failure.

    Run the check again after any topology change. Adding or removing a node, retiring an old witness, rebuilding a host, or turning a two-node cluster into a three-node one can all undo the assumptions behind your original quorum design.

    If high availability was worth building, it's worth testing. Mr.PlanB's Proxmox Health Check can be part of a regular habit of checking cluster state before an outage makes you do it.

    The number to trust is Expected votes

    What surprised me about this incident is that the administrator's understanding of a three-node cluster was basically correct. One node should have been able to fail. The mistake was assuming three nodes meant three voters, when Corosync was expecting four.

    That one number explains the whole failure. With four expected votes, quorum was three, and with one server offline only two votes remained. Removing the stale QDevice brought expected votes back to three and restored the one-node failure tolerance.

    When a Proxmox cluster won't behave the way the diagram in your head says it should, check the vote table before you redesign anything. The problem may be sitting in the configuration while the hardware is fine.

    Frequently Asked Questions

    Should a 3-node Proxmox cluster keep quorum if one node fails?

    Yes, if each node has one vote and there are no extra voters. Two surviving nodes are a majority of three and should keep quorum. If `pvecm status` shows more than three expected votes, check the voter list for a QDevice or stale cluster configuration.

    Why did this 3-node Proxmox cluster require three votes?

    The cluster still had a QDevice registered, which raised expected votes to four. With four expected votes the reported quorum threshold was three, so after one physical node went down only two votes were left and the cluster lost quorum.

    How do I check quorum votes in Proxmox?

    Run `pvecm status` while the cluster is healthy and look at Expected votes, Total votes, Quorum, and Membership information. Those fields show whether Corosync sees only the nodes you expect or an extra QDevice as well.