What Is RabbitMQ High Availability (HA) and How Is It Configured?
RabbitMQ high availability has changed substantially. For most of the broker’s history, HA meant classic mirrored queues. As of RabbitMQ 4.0, classic mirroring is gone. Quorum queues are now the only built-in option for replicated, highly available queues, with streams as a separate replicated data structure for append-only patterns. Teams upgrading from 3.x without migrating their mirroring policies will discover their queues are no longer replicated at all.
This guide explains what high availability means in current RabbitMQ, how to configure it correctly, and what the migration from mirrored queues looks like.
Quick Answer
In RabbitMQ 4.x, high availability for queues is provided by quorum queues, which use Raft consensus to replicate messages across multiple cluster nodes. Streams are a separate replicated data structure for log-style workloads. Classic mirrored queues were deprecated in 3.9 and removed in 4.0. HA requires a cluster of at least three nodes, queues declared with x-queue-type=quorum, durable exchanges, persistent messages, and clients configured for automatic failover. Capacity planning must account for the disk and memory cost of replication.
What HA Actually Means
In a RabbitMQ context, high availability covers three distinct concerns:
- Queue availability: the queue continues to accept and deliver messages when a node fails.
- Message durability: messages survive node failure without loss.
- Cluster availability: the broker as a whole continues to function with majority of nodes alive.
Classic non-replicated queues fail on the first concern: when the node hosting the queue is down, the queue is unavailable. They can satisfy the second concern (persistence on disk) but only if the node comes back. Quorum queues address all three.
How Quorum Queues Provide HA
Quorum queues store their state as a Raft log replicated across queue members (typically three or five nodes). Writes are confirmed only after a majority of members have persisted them. Reads come from the leader. If the leader fails, the remaining members elect a new leader from the surviving replicas. The queue continues to operate as long as a majority of members are reachable.
Key properties:
- All data is persisted to disk before being confirmed to producers.
- Replication is built into the queue type, not bolted on via policy.
- Failover is fast and predictable; clients see at most a brief unavailability during leader election.
- Throughput in steady state is higher than classic mirrored queues were.
Setting Up an HA Cluster
Step 1: Run at least three nodes
Quorum needs an odd majority. Three nodes tolerates one node failure; five tolerates two. Two nodes is worse than one node for HA, because a network partition leaves both sides unable to form a majority.
Step 2: Cluster the nodes
# On node 2 and 3
rabbitmqctl stop_app
rabbitmqctl reset
rabbitmqctl join_cluster rabbit@node1
rabbitmqctl start_app
Verify with rabbitmqctl cluster_status.
Step 3: Declare quorum queues
Producers and consumers should declare queues with x-queue-type=quorum:
channel.queue_declare(
queue='orders.new',
durable=True,
arguments={'x-queue-type': 'quorum'},
)
Alternatively, set the default queue type per virtual host so consumers do not need to know:
# New vhost with quorum as the default queue type
rabbitmqctl add_vhost orders --default-queue-type quorum
# Or update an existing vhost in place
rabbitmqctl update_vhost orders --default-queue-type quorum
Step 4: Configure dead-lettering
Quorum queues in 4.x default to x-delivery-limit=20. Without a dead-letter exchange, messages exceeding that limit are silently dropped. Always configure a DLX for production quorum queues.
Step 5: Configure clients for failover
The broker handles failover internally. Clients also need to know how to connect to whichever node is reachable. Use the modern client library’s connection list feature, providing all node addresses, and enable automatic recovery.
Step 6: Plan disk and memory capacity
Quorum queues use more disk than classic queues because all messages are persisted and the Raft log is maintained per queue. The RabbitMQ production checklist recommends overprovisioning disk and verifying that the disk free limit is set to a substantial value (e.g., comparable to the memory watermark), not the default 50 MB.
Streams as an Alternative for Some Workloads
Streams are a separate replicated data structure introduced in 3.9. They are append-only, support replay, and scale to higher throughput than quorum queues for log-style workloads. Streams are not a drop-in replacement for queues; they have different semantics (offset-based consumption, no per-message ack and redelivery model).
Use quorum queues for work-queue patterns where messages should be consumed once and acked. Use streams for event-log patterns where multiple consumers replay from offsets and retention is time-based.
Migrating From Classic Mirrored Queues
If you are running RabbitMQ 3.13 with mirrored queues and planning to upgrade to 4.x, migration is mandatory before the upgrade. After 4.0, mirroring-related policy keys (ha-mode, ha-params, ha-sync-mode, etc.) have no effect.
Blue-green migration
The recommended path. Stand up a new 4.x cluster (green), use rabbitmqadmin definitions export to extract definitions from the 3.x cluster (blue), edit the exported JSON to convert mirrored classic queues to quorum queues, import into green with rabbitmqadmin definitions import, then switch producers and consumers over. The transformation step is performed outside rabbitmqadmin (in a script or by hand); the tool itself handles only export and import.
The RabbitMQ team has documented this migration path explicitly for 3.13 to 4.x.
In-place migration
Possible for workloads that can tolerate a queue swap. Create a new quorum queue alongside the existing mirrored queue, switch consumers to the new queue, drain the old queue, switch producers, delete the old queue. More disruptive than blue-green and harder to roll back.
What does not work
You cannot simply upgrade 3.x to 4.x and expect mirrored queues to “become” quorum queues. The upgrade leaves mirrored queues as plain classic queues without replication. Plan the migration before the upgrade.
Configuration Mistakes That Defeat HA
Two-node clusters
A two-node cluster has no quorum majority on partition. One node down means the surviving node still cannot form a majority. Always run three or more nodes for HA workloads.
Quorum queues with replicas on only one node
Default behaviour places replicas across cluster nodes, but explicit policies or misconfiguration can constrain a quorum queue to fewer replicas. Verify with rabbitmq-queues quorum_status <queue>.
Clients pointing at a single node
A client that knows only one node’s address will fail when that node fails, even if the cluster as a whole is healthy. Configure clients with all node addresses or a load balancer that routes to surviving nodes.
Persistent messages on non-durable queues
Persistence is a per-message property. The message must be persistent AND the queue must be durable AND the queue must be replicated. All three. Missing any one defeats the durability guarantee.
Disk too small for quorum queues
Classic mirrored queues stored less on disk than quorum queues do. Migrating without re-sizing disk is a recipe for unexpected disk alarms.
Common Mistakes
- Treating clustering as HA. Clustering provides metadata replication and a unified namespace. By itself it does not replicate queue contents. Quorum queues do.
- Using mirroring policies in 4.x. They have no effect; the queues are not replicated.
- Skipping dead-letter routing on quorum queues. The default delivery limit of 20 means messages are silently dropped without a DLX.
- Running two-node clusters for HA. No majority means no quorum.
- Forgetting client-side failover. A perfectly available cluster does not help if clients are pinned to a dead node.
Version Notes
- 3.x: classic mirrored queues exist and work. Quorum queues are available from 3.8 onward. Begin migration planning.
- 4.0: classic mirroring removed. Quorum queues default
x-delivery-limit=20. vm_memory_high_watermark.relativedefault changed from 0.4 to 0.6.queue_master_locatordeprecated in favour ofqueue_leader_locator. - 4.1+: continued improvements to quorum queue performance and feature parity. Streams continue to evolve.
- 4.2+: Khepri replaces Mnesia as the default metadata store. Migrations from earlier versions should test against the metadata store path explicitly.
- 4.3+:
queue_master_locatoris no longer merely deprecated. It is denied by default, and a policy carrying it is rejected outright with “use of deprecated queue-master-locator argument is not permitted.” Any policy still using it must be rewritten to queue-leader-locator before upgrading. Check withrabbitmqctl list_deprecated_features.
Summary Table
| HA concern | Solution | Configuration |
|---|---|---|
| Queue availability after node failure | Quorum queues | x-queue-type=quorum |
| Message durability | Persistent messages on durable, replicated queue | delivery_mode=2, durable=true, quorum |
| Cluster availability | Three or more nodes | join_cluster, verify with cluster_status |
| Client failover | All node addresses, recovery enabled | Client library configuration |
| Poison message handling | DLX | Policy with dead-letter-exchange |
| Log-style workload HA | Streams | Stream declaration, not quorum queue |
Metrics to Monitor
| Metric | What it tells you | Action |
|---|---|---|
| Cluster partition state | Split-brain | Page immediately |
| Quorum queue leader distribution | Cluster balance | Rebalance if heavily skewed |
| Replica state per quorum queue | Replication health | Investigate if replicas behind |
| Per-node connection count | Failover behaviour of clients | Verify clients distributed |
| Disk usage on quorum subdirectory | Raft log footprint | Plan capacity |
FAQs
What replaced classic mirrored queues in RabbitMQ 4.0?
Quorum queues. They use Raft consensus to replicate messages across cluster nodes, with stronger guarantees and higher throughput than mirrored queues. Streams are a separate replicated data structure for append-only workloads. Classic queues without mirroring still exist and are supported, but they are not replicated.
How many nodes do I need for high availability?
At least three. Quorum needs a majority. Two nodes cannot form a majority on partition (each side is one of two, not a majority). Three nodes tolerate one node failure; five tolerate two.
Can I just upgrade from 3.x to 4.x and expect my mirrored queues to work?
No. Classic mirroring was removed in 4.0. After upgrade, queues that had mirroring policies become plain classic queues without replication. Plan the migration to quorum queues before the upgrade, using the blue-green migration path the RabbitMQ team has documented.
Do quorum queues replace streams?
No. They solve different problems. Quorum queues are for work-queue patterns where messages are consumed once. Streams are for log patterns where multiple consumers replay from offsets. Use the right one for the workload.
What is the difference between clustering and HA?
Clustering creates a unified RabbitMQ namespace across multiple nodes, with metadata replicated. By itself, clustering does not replicate queue contents; only metadata. For queue contents to be replicated, you need a replicated queue type: quorum queues or streams.
What happens if I lose a quorum queue replica?
If a minority of replicas fails, the queue continues operating normally. The lost replica catches up when it returns. If a majority fails, the queue becomes unavailable until enough replicas are restored to form a majority.
When to Get Expert Help
If you are running RabbitMQ in production and have not yet migrated from classic mirrored queues, or you are planning a 3.x to 4.x upgrade, an HA review can identify whether your current topology, queue declarations, client configuration, and capacity will deliver the availability you expect. Seventh State reviews cluster topology, queue types, client failover, and the migration plan, then provides a step-by-step path with rollback options and capacity guidance for the new disk footprint.
| Seventh State Team




