Pacemaker, Corosync and GFS2 shared storage on RHEL 8: build guide

Build a multi-node Pacemaker/Corosync cluster on RHEL 8 with DLM, lvmlockd and GFS2 shared filesystems: fencing, journals, resources, constraints and failover tests.

By Rami Chiha11 min read
Quick engineering answer

What this guide covers

What this solves
Build a multi-node Pacemaker/Corosync cluster on RHEL 8 with DLM, lvmlockd and GFS2 shared filesystems: fencing, journals, resources, constraints and failover tests.
Applies to
Linux / RHEL · RHEL 8 · Pacemaker · Corosync · GFS2
Prerequisites
Review the guide's prerequisites, architecture, and decision sections before execution.
Risk
Follow every warning and validate environment-specific commands before a change window.
Expected result
GFS2 lets several RHEL nodes mount the same block device at the same time, for example for shared application homes, shared configuration or shared binaries across an application tier. It depends on a working cluster: Pacemaker and Corosync for membership, DLM for distributed locking, `lvmlockd` for shared LVM, and **fencing** so that a misbehaving node can be cut off.
How to verify
Use the guide's validation, test, acceptance, or handover steps.
Last verified
2026-10-04

GFS2 lets several RHEL nodes mount the same block device at the same time, for example for shared application homes, shared configuration or shared binaries across an application tier. It depends on a working cluster: Pacemaker and Corosync for membership, DLM for distributed locking, lvmlockd for shared LVM, and fencing so that a misbehaving node can be cut off.

This guide shows the build pattern for a four-node cluster on RHEL 8 (virtual machines on a hypervisor, with shared LUNs presented to every node). Names and addresses are placeholders.

Architecture#

  • 4 cluster nodes: app01 to app04, two NICs each: management and a private cluster network.
  • Shared LUNs presented to all nodes, one per shared filesystem.
  • One cluster per application tier, so a problem in one tier cannot freeze the others.
  • Shared filesystems mounted by Pacemaker, not by /etc/fstab.

0. Before you start#

  • Time sync on all nodes (chronyc sources).
  • Forward and reverse DNS, and a consistent /etc/hosts for cluster names.
  • The same shared devices are visible on every node. Prefer stable names (/dev/disk/by-id/... or multipath aliases) over /dev/sdX, which can differ per node.
  • A fencing topology decided in advance. It must fence every node through a path that does not depend on the failed node. GFS2 and DLM without working fencing are unsafe.
  • Supported package repositories enabled for the installed RHEL 8 release family. Commands and resource defaults can differ between RHEL 8 minor releases; use the matching Red Hat documentation and installed pcs --help, not syntax copied from another family.

1. Install the packages#

bash
sudo dnf install -y pcs pacemaker fence-agents-all lvm2-lockd gfs2-utils dlm
sudo passwd hacluster                      # same password on every node
sudo systemctl enable --now pcsd.service

Open the firewall for high availability on each node:

bash
sudo firewall-cmd --permanent --add-service=high-availability
sudo firewall-cmd --reload

2. Create the cluster#

bash
sudo pcs host auth app01 app02 app03 app04 -u hacluster

sudo pcs cluster setup app_clust --start \
  app01 addr=10.60.1.11 addr=10.60.2.11 \
  app02 addr=10.60.1.12 addr=10.60.2.12 \
  app03 addr=10.60.1.13 addr=10.60.2.13 \
  app04 addr=10.60.1.14 addr=10.60.2.14

sudo pcs cluster enable --all
sudo pcs cluster start --all

Using two addr= values per node gives Corosync two links, so one network problem does not split the cluster. Put each on a separate physical or virtual network where you can.

3. Configure fencing first#

Fencing is what makes shared storage safe. Choose Red Hat-supported agents and topology for the actual platform. Power fencing and storage/SCSI fencing have different failure assumptions; SCSI fencing alone is not a generic substitute for power fencing. Verify agent support, credentials, mappings, timeouts, redundant fence paths and any watchdog requirements against the RHEL 8 minor-release documentation.

After creating the STONITH devices with credentials held outside shell history, verify that every cluster node maps to the intended physical or virtual machine. Then test each node in a maintenance window, one at a time:

bash
sudo pcs stonith status
sudo pcs property set stonith-enabled=true
: "${TEST_NODE:?set one verified cluster node name}"
[[ "$TEST_NODE" =~ ^[A-Za-z0-9][A-Za-z0-9.-]*$ ]] || exit 2
sudo pcs stonith fence "$TEST_NODE"

Do not continue until fencing is enabled, pcs stonith status covers every node, each fence test powers off or isolates only its intended target, and the node rejoins cleanly. Disabling STONITH to make resources start is not an acceptable workaround.

4. Quorum behaviour#

bash
sudo pcs property set no-quorum-policy=freeze

With freeze, resources on a partition that loses quorum stay where they are but are not touched, which is what GFS2 requires. Monitor quorum; it is the first thing to check when the cluster is "stuck".

5. DLM and lvmlockd resources#

bash
sudo pcs resource create dlm --group locking ocf:pacemaker:controld \
  op monitor interval=30s on-fail=fence
sudo pcs resource clone locking interleave=true
sudo pcs resource create lvmlockd --group locking ocf:heartbeat:lvmlockd \
  op monitor interval=30s on-fail=fence
sudo pcs status --full

Confirm dlm and lvmlockd run on every node before moving on. Also confirm use_lvmlockd = 1 is set in /etc/lvm/lvm.conf on all nodes, as required by lvmlockd, and follow Red Hat's procedure for your minor release regarding the LVM devices file (the other nodes may need the shared devices added).

6. Shared volume groups and logical volumes#

Run vgcreate and lvcreate once, on one node. Before any destructive command, record the array LUN ID and expected WWID in the approved change, then prove that every node resolves that WWID to the intended multipath map. Do not proceed with a raw /dev/sdX path, inconsistent WWIDs, a degraded map, existing signatures, mounts, holders, active LVs, or any unknown consumer.

The following is a preflight pattern, not a device-discovery shortcut. Substitute the approved WWID and alias; quote variables, require absolute mapper paths, and abort on ambiguity. Run the read-only checks on every node and peer-review the collected output.

bash
: "${EXPECTED_WWID:?set the approved array WWID}"
: "${MAP:?set the approved /dev/mapper path}"
[[ "$EXPECTED_WWID" =~ ^[0-9A-Fa-f]{16,128}$ ]] || exit 2
case "$MAP" in /dev/mapper/*) ;; *) exit 2 ;; esac
[[ "$MAP" =~ ^/dev/mapper/[A-Za-z0-9+_.-]+$ ]] || exit 2
test -b "$MAP"
MAP_NAME=${MAP##*/}
test "$(multipathd show maps raw format '%n %w' | awk -v n="$MAP_NAME" -v w="$EXPECTED_WWID" '$1==n && $2==w{c++} END{print c+0}')" -eq 1
multipath -ll "$MAP_NAME"
test -z "$(lsblk -nr -o MOUNTPOINTS "$MAP" | sed '/^$/d')"
test -z "$(lsblk -nr -o TYPE "$MAP" | awk '$1=="lvm"{print}')"
test -z "$(pvs --noheadings -o pv_name 2>/dev/null | awk -v d="$MAP" '$1==d{print}')"
test -z "$(wipefs -n --output TYPE --noheadings "$MAP" | sed '/^$/d')"
lsblk -e7 -o NAME,PATH,SIZE,TYPE,FSTYPE,MOUNTPOINTS,WWN "$MAP"

Also inspect lsblk -o NAME,PKNAME,HCTL,TYPE,FSTYPE,MOUNTPOINTS, dmsetup ls --tree, fuser -vm "$MAP", lsof "$MAP", array host mappings and multipath path health. Treat inconclusive output as a failed gate. wipefs -n is read-only; do not use wipefs -a to make a gate pass. Stop cluster resources or applications that could claim the device before proceeding.

Back up configuration before changing storage:

bash
BACKUP_DIR='/root/cluster-change-backup'
install -d -o root -g root -m 0700 "$BACKUP_DIR"
sudo pcs config backup "$BACKUP_DIR/pre-storage-change.tar.bz2"
sudo vgcfgbackup -f "$BACKUP_DIR/vg-%s.conf"

Verify the backup files are non-empty and move a protected copy off-node. The exact pcs config backup archive suffix and restore command are release-dependent; confirm them with installed help.

bash
: "${APPS_MAP:?set the approved applications multipath map}"
: "${CONF_MAP:?set the approved configuration multipath map}"
[[ "$APPS_MAP" =~ ^/dev/mapper/[A-Za-z0-9+_.-]+$ ]] || exit 2
[[ "$CONF_MAP" =~ ^/dev/mapper/[A-Za-z0-9+_.-]+$ ]] || exit 2
test -b "$APPS_MAP" && test -b "$CONF_MAP" && test "$APPS_MAP" != "$CONF_MAP"
# Proceed only after the full gate above passed for both maps on every node.
sudo vgcreate --shared vg_apps "$APPS_MAP"
sudo vgcreate --shared vg_conf "$CONF_MAP"

sudo lvcreate --activate sy -l+100%FREE -n lv_apps vg_apps
sudo lvcreate --activate sy -l+100%FREE -n lv_conf vg_conf

7. Create the GFS2 filesystems#

The cluster name in -t must match the cluster, and the journal count (`-j`) must be at least the number of nodes that will mount the filesystem. A four-node cluster needs -j4; a five-node cluster needs -j5.

bash
sudo mkfs.gfs2 -j4 -p lock_dlm -t app_clust:gfs2-apps /dev/vg_apps/lv_apps
sudo mkfs.gfs2 -j4 -p lock_dlm -t app_clust:gfs2-conf /dev/vg_conf/lv_conf

Formatting erases the volume. Immediately before each mkfs.gfs2, repeat the signature, mount, holder and in-use checks against the LV and confirm its VG/PV lineage points to the approved WWID. Adding a node later means following the release-matched Red Hat gfs2_jadd procedure while the filesystem is mounted; add capacity before allowing the new node to mount it.

8. Pacemaker resources and constraints#

Create /apps and /conf as empty mount-point directories on every node before defining the Filesystem resources. Confirm they are not listed in /etc/fstab and are not already mounted. Pacemaker must be their only mount owner.

bash
# Run on every node before creating the resources:
sudo install -d -o root -g root -m 0755 /apps /conf

# Activate shared logical volumes on every node
sudo pcs resource create lv_apps --group vg_apps ocf:heartbeat:LVM-activate \
  lvname=lv_apps vgname=vg_apps activation_mode=shared vg_access_mode=lvmlockd
sudo pcs resource create lv_conf --group vg_conf ocf:heartbeat:LVM-activate \
  lvname=lv_conf vgname=vg_conf activation_mode=shared vg_access_mode=lvmlockd
sudo pcs resource clone vg_apps interleave=true
sudo pcs resource clone vg_conf interleave=true

# Mount the filesystems
sudo pcs resource create fs_apps --group vg_apps ocf:heartbeat:Filesystem \
  device=/dev/vg_apps/lv_apps directory=/apps fstype=gfs2 options=noatime \
  op monitor interval=10s on-fail=fence
sudo pcs resource create fs_conf --group vg_conf ocf:heartbeat:Filesystem \
  device=/dev/vg_conf/lv_conf directory=/conf fstype=gfs2 options=noatime \
  op monitor interval=10s on-fail=fence

# Locking must start before storage, on the same nodes
sudo pcs constraint order start locking-clone then vg_apps-clone
sudo pcs constraint order start locking-clone then vg_conf-clone
sudo pcs constraint colocation add vg_apps-clone with locking-clone
sudo pcs constraint colocation add vg_conf-clone with locking-clone

pcs syntax and generated clone names vary across RHEL 8 minor releases. Review the resulting CIB with pcs config and confirm the mandatory order and colocation constraints before enabling application access; do not infer safety from a resource merely showing Started.

9. Validate#

AreaTest
Clusterpcs status --full shows all nodes online, all resources started
Lockingdlm and lvmlockd running on every node
Storagevgs, lvs show shared VGs; df -h shows GFS2 mounted everywhere
Shared accessCreate a file on one node and read it on the others
FailoverReboot one node; confirm the remaining mounts stay healthy and the node rejoins without manual mounting
FencingFence each node in turn; confirm the intended node is isolated before recovery and no other node is affected
QuorumReview behaviour with pcs quorum status

Handy commands: pcs status --full, crm_mon -1, pcs resource status, systemctl status pcsd. (On RHEL 8, pcs resource show is deprecated in favour of pcs resource status.)

Rollback and recovery#

  • Before resource changes, capture pcs status --full, pcs config, quorum, fencing, multipath, LVM and mount output with timestamps. Put the cluster in the release-supported maintenance state only when the documented procedure calls for it; maintenance mode does not replace fencing or make concurrent manual mounts safe.
  • For a CIB-only error, stop the affected rollout and restore the reviewed CIB backup using the command documented by the installed pcs version, then revalidate constraints, fencing and quorum before clearing maintenance state.
  • For LVM metadata damage, stop all users on every node and engage storage support. Restore only from the matching vgcfgbackup after confirming the correct WWID and current on-disk state; an incorrect vgcfgrestore can worsen damage.
  • mkfs.gfs2, pvcreate and overwritten signatures have no configuration rollback. Stop I/O, preserve evidence and recover from a tested backup or snapshot supported by the storage design.
  • If a node loses lock membership or fencing is uncertain, do not manually mount the GFS2 filesystem or force VG activation. Fence/isolate the suspect node, establish quorum and lock-manager health, then use the Red Hat recovery procedure.

Vendor references#

Common mistakes#

  • No fencing, or untested fencing. The most serious one.
  • Too few journals for the number of nodes.
  • Device names that differ between nodes. Use stable IDs or multipath aliases.
  • Mounting through `/etc/fstab` instead of Pacemaker, so the cluster loses control of the filesystem.
  • One huge cluster for unrelated applications. A single quorum event then affects everything.
  • Forgetting that GFS2 is for shared access. It is not faster than a local filesystem for single-node workloads.

Verify package names, agent parameters and LVM requirements against the Red Hat Enterprise Linux 8 documentation for your minor release.

Building or rescuing a Pacemaker and GFS2 environment? See Managed Infrastructure Operations or start a project.