OpenShift etcd backup, handover checklist and DR cluster build notes

How to back up etcd on OpenShift, what to hand over to operations, and how to plan a disaster recovery cluster with a value sheet and restore decisions.

By Rami Chiha7 min read
Quick engineering answer

What this guide covers

What this solves
How to back up etcd on OpenShift, what to hand over to operations, and how to plan a disaster recovery cluster with a value sheet and restore decisions.
Applies to
OpenShift · OpenShift · etcd · backup · disaster recovery
Prerequisites
Review the guide's prerequisites, architecture, and decision sections before execution.
Risk
Follow every warning and validate environment-specific commands before a change window.
Expected result
A cluster is not finished when `install-complete` prints. It is finished when someone other than the builder can operate it, back it up and rebuild it. This article covers three things: the etcd backup, the evidence and artifacts that make up a proper handover, and the planning notes for a disaster recovery (DR) cluster.
How to verify
Use the guide's validation, test, acceptance, or handover steps.
Last verified
2026-10-04

A cluster is not finished when install-complete prints. It is finished when someone other than the builder can operate it, back it up and rebuild it. This article covers three things: the etcd backup, the evidence and artifacts that make up a proper handover, and the planning notes for a disaster recovery (DR) cluster.

This procedure was validated for the OpenShift 4 release-provided backup workflow described here. Backup scripts, output names, restore sequencing and supported recovery scope can change across version families. Before every upgrade or restore exercise, revalidate them against Red Hat's documentation for the cluster's exact release and record that release with the backup. Never use instructions from a newer cluster merely because both releases are OpenShift 4.

1. etcd backup#

etcd holds the cluster's state. Back it up on a schedule and store the copy outside the cluster. A backup that exists only on a control plane node is not a backup.

Take a backup after the cluster is healthy and after material configuration changes; do not wait for an assumed 24-hour certificate event. Certificate rotation timing and backup guidance are release-specific. A snapshot is only useful with its matching static-pod resources and a restore procedure supported for the exact OpenShift release.

bash
umask 077
kubeconfig="$HOME/secure-admin/cluster-kubeconfig"
test -f "$kubeconfig" && test "$(stat -c '%a' "$kubeconfig")" = 600
KUBECONFIG="$kubeconfig" oc get nodes -l node-role.kubernetes.io/master=

# Pick one Ready control plane node, record its name, then open the approved debug session.
control_plane_node=cp01
KUBECONFIG="$kubeconfig" oc get node "$control_plane_node"
KUBECONFIG="$kubeconfig" oc debug "node/$control_plane_node"
chroot /host
set -euo pipefail
umask 077
backup_id="$(date -u +%Y%m%dT%H%M%SZ)"
backup_dir="/home/core/assets/etcd-backup-$backup_id"
install -d -o root -g root -m 0700 "$backup_dir"
test -x /usr/local/bin/cluster-backup.sh
/usr/local/bin/cluster-backup.sh "$backup_dir"
test "$?" -eq 0
find "$backup_dir" -maxdepth 1 -type f -size +0c -printf '%f\n'
test "$(find "$backup_dir" -maxdepth 1 -type f -size +0c | wc -l)" -ge 2
(cd "$backup_dir" && sha256sum -- * > SHA256SUMS)
chmod 0600 "$backup_dir"/*
sync
printf 'Staged backup: %s\n' "$backup_dir"
exit; exit

For releases that document this script and path, it produces a snapshot and static-pod resources. The count check is only a fail-closed minimum: compare exact output types and names with the release-matched procedure. Any script error, empty/missing artifact, unexpected file, checksum failure or inability to identify the cluster and release makes the backup job fail; do not publish a "successful" status or delete the previous good generation.

Transfer the complete set through the approved backup mechanism, not a public web server or an administrator's ordinary home directory. Encrypt in transit and at rest with organisation-managed keys; grant read/restore access only to the recovery role, separate backup deletion from cluster administration where possible, enable immutable or object-locked generations, and audit access. Store cluster identity, cluster version, UTC timestamp, source node, script output, file inventory and checksums as protected metadata without embedding kubeconfig or other credentials. Recompute checksums at the destination and compare them before success is recorded; a checksum detects corruption but does not prove restorability.

After destination verification and catalog registration, remove the node staging directory using the approved sensitive-data cleanup process and verify it is absent. Do not put a kubeconfig, pull secret or backup encryption key in scripts, command arguments, environment exported to child processes, backup sets or logs. Retrieve credentials just in time from the approved vault, use a mode-0600 temporary kubeconfig on encrypted storage, avoid shell tracing, and revoke temporary access after the run.

Practical rules:

  • Schedule it with an automation identity limited to the documented backup actions and alert on command, artifact, transfer, checksum, catalog and cleanup failures.
  • Define retention before scheduling: minimum successful generations, daily/weekly/monthly periods, immutability period, legal holds, expiry owner and deletion evidence. Retention is a business and regulatory decision; automation must never prune the last known-good or held generation.
  • Test the release-matched restore procedure in an isolated recovery exercise. An etcd snapshot restores the originating cluster's state; it is not a supported way to seed an independently installed DR cluster. Restore is disruptive and must follow Red Hat's documented control-plane recovery sequence for that release.
  • Application data and persistent volumes need their own backup. etcd does not contain them.

Rollback for a failed backup run is to retain the previous verified generations, revoke temporary access, quarantine partial destination objects, and clean staging only after evidence has been captured. A restore has no generic rollback: define the release-specific recovery decision, outage approval, pre-restore evidence and abort points in a separately tested runbook before an incident.

2. Artifacts to keep#

ArtifactWhy
install-config.yaml backup (without secrets in tickets)Proves the original build values
Install directory metadataCluster identity and generated assets
Break-glass kubeconfigStored in a vault with access control
Load balancer configs from both nodesRequired for support and rebuild
DNS and IP tableOperations and DR planning
OAuth YAML without passwordsAuthentication record
Storage mapping (LUN ID, WWN, PV/PVC YAML)Ownership and restore
etcd backup (off-cluster)Control plane recovery

3. Final acceptance checklist#

CheckEvidence
DNS forward and reverse validateddig output
VIPs move between load balancersFailover test
HAProxy config valid on both nodeshaproxy -c -f ...
All nodes Readyoc get nodes -o wide
Cluster operators healthyoc get co
Machine config pools updatedoc get mcp
Console and OAuth routes respondHTTP 200 test
Directory login worksoc login with a real user
kubeadmin removedSecret not found
Storage visiblemultipath -ll per worker
PV/PVC workflow testedTest pod mount
etcd backup copied off-clusterComplete file list, source and destination checksums, and backup command success evidence
Backup controls testedEncryption, restore-role access, immutability, retention, alerting and staging cleanup evidence

Mark each as done, pending or partial with a date. Pending items are normal at handover; hidden pending items are not.

4. Minimum information per application#

  • Namespace and application owner
  • Deployment manifest, Helm values or Operator subscription
  • Image source and tag policy
  • Service, route and certificate needs
  • PVC names, sizes, mount paths and backup requirement
  • Runbook: restart, scale, logs, health checks, rollback

5. Daily quick checks#

bash
oc whoami
oc get nodes -o wide
oc get co
oc get mcp
oc get clusterversion
oc get pods -A | egrep -v 'Running|Completed' || true
oc get events -A --sort-by=.lastTimestamp | tail -50
oc adm top nodes || true

6. Building a DR cluster#

Use the same method, with new values. Never copy generated ignition files, kubeconfig, certificates or storage mappings from the primary cluster. Generate everything fresh.

Value sheet#

ParameterPrimaryDR site
Cluster namerecorded primary nameapproved unique DR name
Base domainrecorded primary domainapproved DR domain
API / API-INT / Apps VIPsrecorded two-VIP designnew API and Apps VIPs
Load balancers, bootstrap, bastionrecorded primary inventorynew DR addresses
Control planes and workersrecorded primary inventorynew DR addresses
Cluster / service networksdefaultsconfirm same or non-overlapping
NTP, proxy, directoryrecorded primary endpointsapproved DR endpoints

Build sequence#

  1. Confirm IP plan, DNS, NTP, directory reachability, registry access and firewall paths with the UPI prerequisites checklist.
  2. Build and failover-test the load balancers with the two-VIP HAProxy and Keepalived guide; prepare the matching installer and client.
  3. Create a new install-config.yaml and generate new manifests and ignition files using the UPI install guide.
  4. Boot bootstrap, control plane, then workers; complete bootstrap, then inspect and approve only CSRs that match the expected nodes, signer and request details.
  5. Configure authentication and RBAC with the LDAP/AD guide, and validate storage recovery with the static Fibre Channel PV guide where that pattern applies.
  6. Run the acceptance checklist before any application restore or failover test.

Decisions to make early#

TopicQuestion
Data sourceRestore from backup, replicated storage, or rebuild clean?
LUN mappingWho promotes or presents LUNs to DR workers?
PV/PVC namesSame as primary or site-specific?
Route names and DNSDo application names move on failover, or use separate names?
CertificatesReuse the wildcard or issue separately?
AuthenticationSame admin groups on DR?
BackupsWhere do control plane and application backups live?

The recovery time you can promise is set by these answers, not by how fast the cluster installs. Decide explicitly whether the objective is recovery of the original cluster from its etcd backup or deployment of a separate DR cluster followed by GitOps/application-data restoration; they are different procedures and the artifacts are not interchangeable.

Examples are generic. Verify commands against the Red Hat documentation for your OpenShift version, especially restore procedures.

Need a DR design reviewed or a restore test run? See Backup and Recovery and Kubernetes and OpenShift, or start a project.