HAProxy and Keepalived for OpenShift UPI: load balancer template

A tested pattern for OpenShift user-provisioned installs: two HAProxy nodes, two Keepalived VIPs, bootstrap handling, SELinux and failover testing.

By Rami Chiha7 min read
Quick engineering answer

What this guide covers

What this solves
A tested pattern for OpenShift user-provisioned installs: two HAProxy nodes, two Keepalived VIPs, bootstrap handling, SELinux and failover testing.
Applies to
OpenShift · OpenShift · HAProxy · Keepalived · VRRP
Prerequisites
Review the guide's prerequisites, architecture, and decision sections before execution.
Risk
Follow every warning and validate environment-specific commands before a change window.
Expected result
OpenShift on user-provisioned infrastructure needs load balancing for the Kubernetes API, the Machine Config Server (MCS) used while nodes boot, and application ingress. This is an example for a release and platform whose documented requirements are satisfied by two RHEL 9 helper servers running HAProxy and Keepalived. Confirm support, health-check requirements and endpoint exposure against the exact OpenShift release before adopting it.
How to verify
Use the guide's validation, test, acceptance, or handover steps.
Last verified
2026-10-04

OpenShift on user-provisioned infrastructure needs load balancing for the Kubernetes API, the Machine Config Server (MCS) used while nodes boot, and application ingress. This is an example for a release and platform whose documented requirements are satisfied by two RHEL 9 helper servers running HAProxy and Keepalived. Confirm support, health-check requirements and endpoint exposure against the exact OpenShift release before adopting it.

This article gives the working template. Addresses are documentation-only values from the prerequisites checklist. It was validated for this two-node, two-VIP OpenShift 4 UPI pattern. Revalidate package support, ports, health checks and configuration syntax before a version-family, operating-system or network-topology change; do not infer support for every OpenShift 4 release.

Design#

EndpointPortBackend pool
API (api)TCP 6443Control plane nodes (+ bootstrap during install)
MCS (api-int)TCP 22623Control plane nodes (+ bootstrap during install)
Ingress HTTPTCP 80Worker nodes
Ingress HTTPSTCP 443Worker nodes

Both load balancer nodes run HAProxy. Keepalived moves two VIPs between them: one API VIP, to which both api and internal-only api-int DNS names resolve, and one Ingress VIP for *.apps. A separate api-int address is not inherently required; use one only when the release documentation and approved network design call for it. HAProxy runs in TCP mode, so TLS passes through untouched and certificates stay on the cluster.

Before configuration, reconcile the complete worker inventory against DNS and the build record. Every schedulable worker expected to run an ingress router must have exactly one port 80 and one port 443 backend line on both load balancers; omit dedicated non-ingress workers only when the IngressController placement explicitly excludes them. The three workers below are an example, not a complete production inventory.

Install and baseline both nodes#

Review the transaction on each load balancer before accepting it:

bash
sudo dnf install --assumeno haproxy keepalived
# After repository, package, dependency and change approval:
sudo dnf install -y haproxy keepalived
rpm -q haproxy keepalived
systemctl is-enabled firewalld
getenforce

Keep firewalld and SELinux enforcing. Save the package versions and repeat all configuration and validation steps on both nodes.

HAProxy configuration#

text
# /etc/haproxy/haproxy.cfg
global
  log /dev/log local0
  maxconn 20000
  daemon

defaults
  log global
  mode tcp
  option tcplog
  option dontlognull
  timeout connect 10s
  timeout client  1m
  timeout server  1m

frontend openshift_api
  bind *:6443
  default_backend openshift_api_backend

backend openshift_api_backend
  balance source
  option tcp-check                 # transport reachability only, not API readiness
  server bootstrap 10.42.10.30:6443 check      # remove after bootstrap-complete
  server cp01      10.42.20.11:6443 check
  server cp02      10.42.20.12:6443 check
  server cp03      10.42.20.13:6443 check

frontend machine_config_server
  bind *:22623
  default_backend machine_config_server_backend

backend machine_config_server_backend
  balance source
  option tcp-check                 # transport reachability only
  server bootstrap 10.42.10.30:22623 check     # remove after bootstrap-complete
  server cp01      10.42.20.11:22623 check
  server cp02      10.42.20.12:22623 check
  server cp03      10.42.20.13:22623 check

frontend ingress_http
  bind *:80
  default_backend ingress_http_backend

backend ingress_http_backend
  balance source
  option tcp-check                 # transport reachability only
  server wk01 10.42.20.21:80 check
  server wk02 10.42.20.22:80 check
  server wk03 10.42.20.23:80 check

frontend ingress_https
  bind *:443
  default_backend ingress_https_backend

backend ingress_https_backend
  balance source
  option tcp-check                 # transport reachability only
  server wk01 10.42.20.21:443 check
  server wk02 10.42.20.22:443 check
  server wk03 10.42.20.23:443 check

Add one server line per worker. Binding to * means HAProxy starts cleanly on the standby node that does not currently own the VIPs.

Install the reviewed configuration as root:root mode 0640, then validate before enabling it. Keep the previous configuration in a root-only rollback directory rather than overwriting the only known-good copy.

bash
sudo install -d -o root -g root -m 0700 /root/lb-config-rollback
sudo cp -a /etc/haproxy/haproxy.cfg /root/lb-config-rollback/haproxy.cfg.pre-change
sudo chown root:root /etc/haproxy/haproxy.cfg
sudo chmod 0640 /etc/haproxy/haproxy.cfg
sudo haproxy -c -f /etc/haproxy/haproxy.cfg
sudo systemctl enable --now haproxy
systemctl is-enabled haproxy
systemctl is-active haproxy
sudo ss -lntp | grep -E ':(6443|22623|80|443)[[:space:]]'

Bootstrap membership by phase#

PhaseBootstrap in API and MCS pools?
Before bootstrap bootsYes
During bootstrapYes
After bootstrap-completeNo. Remove or comment out, then reload
bash
sudo haproxy -c -f /etc/haproxy/haproxy.cfg && sudo systemctl reload haproxy

Leaving the bootstrap node in the pools after it has finished is a classic source of intermittent API errors.

Keepalived configuration#

On the primary node:

text
# /etc/keepalived/keepalived.conf (LB01, MASTER)
vrrp_script chk_haproxy {
  script "/usr/bin/pidof haproxy"
  interval 2
  fall 2
  rise 2
}

vrrp_instance OCP4_VIPS {
  state MASTER
  interface ens192                # the NIC that carries the VIPs
  virtual_router_id 50
  priority 120
  advert_int 1
  unicast_src_ip 10.42.10.21
  unicast_peer {
    10.42.10.22
  }
  virtual_ipaddress {
    10.42.10.10/24                # API
    10.42.10.12/24                # Apps
  }
  track_script {
    chk_haproxy
  }
}

On the second node, use state BACKUP, priority 110, unicast_src_ip 10.42.10.22 and the first node as unicast_peer. Keep virtual_router_id identical on both, and unique on the VLAN.

VRRP auth_type PASS is limited, transmitted in clear text and is not a meaningful security boundary, so this example omits it. Restrict VRRP/unicast peer traffic with network controls to the two known peers. If the threat model requires cryptographic peer authentication, use a network-layer mechanism supported by the operating system and network design rather than relying on VRRP PASS.

Confirm the interface name on every build (ip -br addr). Copying an interface name from another environment is a frequent cause of VIPs that never appear.

Protect, validate and enable Keepalived on each node after peer addresses, interface, prefix lengths and the two VIPs have been reviewed:

bash
sudo cp -a /etc/keepalived/keepalived.conf /root/lb-config-rollback/keepalived.conf.pre-change
sudo chown root:root /etc/keepalived/keepalived.conf
sudo chmod 0600 /etc/keepalived/keepalived.conf
sudo keepalived --config-test --use-file=/etc/keepalived/keepalived.conf
sudo systemctl enable --now keepalived
systemctl is-enabled keepalived
systemctl is-active keepalived
journalctl -u keepalived -n 50 --no-pager

If the installed Keepalived uses different validation options, stop and use that package version's documented config-test command; do not start an unvalidated configuration.

Host controls on RHEL 9#

If firewalld and SELinux are enforcing, allow exactly what is needed:

bash
# Load balancer nodes
firewall-cmd --permanent --add-port=6443/tcp --add-port=22623/tcp
firewall-cmd --permanent --add-service=http --add-service=https
firewall-cmd --permanent --add-protocol=vrrp
firewall-cmd --reload

# Allow HAProxy to connect to non-standard backend ports under SELinux
setsebool -P haproxy_connect_any 1

Do not leave SELinux permissive or the firewall off in production without a recorded approval. Both are easy to correct during build and much harder after handover.

Failover test#

bash
# On the active node: confirm it owns the VIPs
ip -br addr show | grep -E '10\.42\.10\.(10|12)(/|$)'

# Stop Keepalived and watch the VIPs move
systemctl stop keepalived
# On the other node
ip -br addr show | grep -E '10\.42\.10\.(10|12)(/|$)'

# Restore
systemctl start keepalived

Also stop HAProxy on the active node and confirm the tracking script moves the VIPs. A PID check only establishes that a process exists, while the backend tcp-check only establishes a TCP connection. Test authenticated or certificate-validated API readiness, MCS reachability from the node network, and an application route during each failover. Avoid normalizing curl -k as an operational health check because it hides certificate failures.

Rollback is node-by-node: move VIP ownership to the healthy peer, restore that node's reviewed files from /root/lb-config-rollback, validate both configurations, restart HAProxy and Keepalived, and repeat on the peer. Roll back if syntax validation fails, either VIP is simultaneously owned by both nodes, a required listener/backend is missing, or API and route checks fail after failover. Never restore both nodes concurrently.

Hardening ideas#

  • Where Red Hat documents an application-level readiness endpoint for the selected release, implement and test a version-compatible HAProxy check; do not invent an HTTP payload or disable TLS verification to make a check pass.
  • Ship HAProxy logs to your central logging.
  • Keep both haproxy.cfg and keepalived.conf under configuration management and restrict their ownership and permissions.

Next step#

With both nodes healthy, continue with the install walkthrough.

Examples use documentation-only names and addresses. Validate against the Red Hat and HAProxy documentation for your exact versions.

Want this designed and tested for your environment? See Kubernetes and OpenShift or start a project.