OpenShift etcd backup, handover checklist and DR cluster build notes
How to back up etcd on OpenShift, what to hand over to operations, and how to plan a disaster recovery cluster with a value sheet and restore decisions.
What this guide covers
- What this solves
- How to back up etcd on OpenShift, what to hand over to operations, and how to plan a disaster recovery cluster with a value sheet and restore decisions.
- Applies to
- OpenShift · OpenShift · etcd · backup · disaster recovery
- Prerequisites
- Review the guide's prerequisites, architecture, and decision sections before execution.
- Risk
- Follow every warning and validate environment-specific commands before a change window.
- Expected result
- A cluster is not finished when `install-complete` prints. It is finished when someone other than the builder can operate it, back it up and rebuild it. This article covers three things: the etcd backup, the evidence and artifacts that make up a proper handover, and the planning notes for a disaster recovery (DR) cluster.
- How to verify
- Use the guide's validation, test, acceptance, or handover steps.
- Last verified
- 2026-10-04
A cluster is not finished when install-complete prints. It is finished when someone other than the builder can operate it, back it up and rebuild it. This article covers three things: the etcd backup, the evidence and artifacts that make up a proper handover, and the planning notes for a disaster recovery (DR) cluster.
This procedure was validated for the OpenShift 4 release-provided backup workflow described here. Backup scripts, output names, restore sequencing and supported recovery scope can change across version families. Before every upgrade or restore exercise, revalidate them against Red Hat's documentation for the cluster's exact release and record that release with the backup. Never use instructions from a newer cluster merely because both releases are OpenShift 4.
1. etcd backup#
etcd holds the cluster's state. Back it up on a schedule and store the copy outside the cluster. A backup that exists only on a control plane node is not a backup.
Take a backup after the cluster is healthy and after material configuration changes; do not wait for an assumed 24-hour certificate event. Certificate rotation timing and backup guidance are release-specific. A snapshot is only useful with its matching static-pod resources and a restore procedure supported for the exact OpenShift release.
umask 077
kubeconfig="$HOME/secure-admin/cluster-kubeconfig"
test -f "$kubeconfig" && test "$(stat -c '%a' "$kubeconfig")" = 600
KUBECONFIG="$kubeconfig" oc get nodes -l node-role.kubernetes.io/master=
# Pick one Ready control plane node, record its name, then open the approved debug session.
control_plane_node=cp01
KUBECONFIG="$kubeconfig" oc get node "$control_plane_node"
KUBECONFIG="$kubeconfig" oc debug "node/$control_plane_node"
chroot /host
set -euo pipefail
umask 077
backup_id="$(date -u +%Y%m%dT%H%M%SZ)"
backup_dir="/home/core/assets/etcd-backup-$backup_id"
install -d -o root -g root -m 0700 "$backup_dir"
test -x /usr/local/bin/cluster-backup.sh
/usr/local/bin/cluster-backup.sh "$backup_dir"
test "$?" -eq 0
find "$backup_dir" -maxdepth 1 -type f -size +0c -printf '%f\n'
test "$(find "$backup_dir" -maxdepth 1 -type f -size +0c | wc -l)" -ge 2
(cd "$backup_dir" && sha256sum -- * > SHA256SUMS)
chmod 0600 "$backup_dir"/*
sync
printf 'Staged backup: %s\n' "$backup_dir"
exit; exitFor releases that document this script and path, it produces a snapshot and static-pod resources. The count check is only a fail-closed minimum: compare exact output types and names with the release-matched procedure. Any script error, empty/missing artifact, unexpected file, checksum failure or inability to identify the cluster and release makes the backup job fail; do not publish a "successful" status or delete the previous good generation.
Transfer the complete set through the approved backup mechanism, not a public web server or an administrator's ordinary home directory. Encrypt in transit and at rest with organisation-managed keys; grant read/restore access only to the recovery role, separate backup deletion from cluster administration where possible, enable immutable or object-locked generations, and audit access. Store cluster identity, cluster version, UTC timestamp, source node, script output, file inventory and checksums as protected metadata without embedding kubeconfig or other credentials. Recompute checksums at the destination and compare them before success is recorded; a checksum detects corruption but does not prove restorability.
After destination verification and catalog registration, remove the node staging directory using the approved sensitive-data cleanup process and verify it is absent. Do not put a kubeconfig, pull secret or backup encryption key in scripts, command arguments, environment exported to child processes, backup sets or logs. Retrieve credentials just in time from the approved vault, use a mode-0600 temporary kubeconfig on encrypted storage, avoid shell tracing, and revoke temporary access after the run.
Practical rules:
- Schedule it with an automation identity limited to the documented backup actions and alert on command, artifact, transfer, checksum, catalog and cleanup failures.
- Define retention before scheduling: minimum successful generations, daily/weekly/monthly periods, immutability period, legal holds, expiry owner and deletion evidence. Retention is a business and regulatory decision; automation must never prune the last known-good or held generation.
- Test the release-matched restore procedure in an isolated recovery exercise. An etcd snapshot restores the originating cluster's state; it is not a supported way to seed an independently installed DR cluster. Restore is disruptive and must follow Red Hat's documented control-plane recovery sequence for that release.
- Application data and persistent volumes need their own backup. etcd does not contain them.
Rollback for a failed backup run is to retain the previous verified generations, revoke temporary access, quarantine partial destination objects, and clean staging only after evidence has been captured. A restore has no generic rollback: define the release-specific recovery decision, outage approval, pre-restore evidence and abort points in a separately tested runbook before an incident.
2. Artifacts to keep#
| Artifact | Why |
|---|---|
install-config.yaml backup (without secrets in tickets) | Proves the original build values |
| Install directory metadata | Cluster identity and generated assets |
Break-glass kubeconfig | Stored in a vault with access control |
| Load balancer configs from both nodes | Required for support and rebuild |
| DNS and IP table | Operations and DR planning |
| OAuth YAML without passwords | Authentication record |
| Storage mapping (LUN ID, WWN, PV/PVC YAML) | Ownership and restore |
| etcd backup (off-cluster) | Control plane recovery |
3. Final acceptance checklist#
| Check | Evidence |
|---|---|
| DNS forward and reverse validated | dig output |
| VIPs move between load balancers | Failover test |
| HAProxy config valid on both nodes | haproxy -c -f ... |
All nodes Ready | oc get nodes -o wide |
| Cluster operators healthy | oc get co |
| Machine config pools updated | oc get mcp |
| Console and OAuth routes respond | HTTP 200 test |
| Directory login works | oc login with a real user |
kubeadmin removed | Secret not found |
| Storage visible | multipath -ll per worker |
| PV/PVC workflow tested | Test pod mount |
| etcd backup copied off-cluster | Complete file list, source and destination checksums, and backup command success evidence |
| Backup controls tested | Encryption, restore-role access, immutability, retention, alerting and staging cleanup evidence |
Mark each as done, pending or partial with a date. Pending items are normal at handover; hidden pending items are not.
4. Minimum information per application#
- Namespace and application owner
- Deployment manifest, Helm values or Operator subscription
- Image source and tag policy
- Service, route and certificate needs
- PVC names, sizes, mount paths and backup requirement
- Runbook: restart, scale, logs, health checks, rollback
5. Daily quick checks#
oc whoami
oc get nodes -o wide
oc get co
oc get mcp
oc get clusterversion
oc get pods -A | egrep -v 'Running|Completed' || true
oc get events -A --sort-by=.lastTimestamp | tail -50
oc adm top nodes || true6. Building a DR cluster#
Use the same method, with new values. Never copy generated ignition files, kubeconfig, certificates or storage mappings from the primary cluster. Generate everything fresh.
Value sheet#
| Parameter | Primary | DR site |
|---|---|---|
| Cluster name | recorded primary name | approved unique DR name |
| Base domain | recorded primary domain | approved DR domain |
| API / API-INT / Apps VIPs | recorded two-VIP design | new API and Apps VIPs |
| Load balancers, bootstrap, bastion | recorded primary inventory | new DR addresses |
| Control planes and workers | recorded primary inventory | new DR addresses |
| Cluster / service networks | defaults | confirm same or non-overlapping |
| NTP, proxy, directory | recorded primary endpoints | approved DR endpoints |
Build sequence#
- Confirm IP plan, DNS, NTP, directory reachability, registry access and firewall paths with the UPI prerequisites checklist.
- Build and failover-test the load balancers with the two-VIP HAProxy and Keepalived guide; prepare the matching installer and client.
- Create a new
install-config.yamland generate new manifests and ignition files using the UPI install guide. - Boot bootstrap, control plane, then workers; complete bootstrap, then inspect and approve only CSRs that match the expected nodes, signer and request details.
- Configure authentication and RBAC with the LDAP/AD guide, and validate storage recovery with the static Fibre Channel PV guide where that pattern applies.
- Run the acceptance checklist before any application restore or failover test.
Decisions to make early#
| Topic | Question |
|---|---|
| Data source | Restore from backup, replicated storage, or rebuild clean? |
| LUN mapping | Who promotes or presents LUNs to DR workers? |
| PV/PVC names | Same as primary or site-specific? |
| Route names and DNS | Do application names move on failover, or use separate names? |
| Certificates | Reuse the wildcard or issue separately? |
| Authentication | Same admin groups on DR? |
| Backups | Where do control plane and application backups live? |
The recovery time you can promise is set by these answers, not by how fast the cluster installs. Decide explicitly whether the objective is recovery of the original cluster from its etcd backup or deployment of a separate DR cluster followed by GitOps/application-data restoration; they are different procedures and the artifacts are not interchangeable.
Examples are generic. Verify commands against the Red Hat documentation for your OpenShift version, especially restore procedures.
Need a DR design reviewed or a restore test run? See Backup and Recovery and Kubernetes and OpenShift, or start a project.