Establish the desired state before the window
vSphere Lifecycle Manager images use a declarative desired state for the ESXi version, vendor add-on, firmware and driver add-on, and additional components. A lifecycle review should compare that image with every host, identify drift, and resolve unsupported combinations before a maintenance window begins.
Broadcom documents that component or vendor-add-on mismatches can make a host incompatible with the cluster image, and automated remediation does not simply erase every mismatch. Validate hardware compatibility, server-vendor guidance, storage and network drivers, firmware dependencies, third-party VIBs, and the supported vCenter-to-ESXi path.
Review capacity and evacuation as one problem
Maintenance mode succeeds only when workloads can move safely. Review current and projected CPU and memory demand, reservations, admission control, datastore capacity, vSAN or storage policy state, DRS rules, affinity constraints, passthrough devices, latency-sensitive workloads, and any virtual appliances tied to a host.
Model the cluster with one host unavailable and enough headroom for workload variation. Confirm that backup, monitoring, DNS, identity, and management dependencies will remain reachable while each host is evacuated. If a cluster cannot tolerate the maintenance pattern, redesign the sequence or add temporary capacity before changing software.
- Confirm vCenter protection, configuration backups, and access to break-glass credentials.
- Check HA admission control, DRS health, active alarms, storage health, and failed hardware.
- List workloads that cannot vMotion and define an approved shutdown or exclusion path.
Stage remediation and preserve a decision point
Run the available pre-checks and compliance assessment before remediation. Start with a suitable host or lower-risk cluster, observe boot, driver, network, storage, management-agent, and workload behavior, then pause for review. Parallelism can shorten a window but also concentrates capacity and compatibility risk.
A rollback plan needs more than an old ISO. Record which changes are reversible, the supported rollback method, configuration dependencies, and the point at which continuing is safer than reverting. Assign who can stop the sequence and what evidence triggers that decision.
Validate service, not just compliance
Image compliance is necessary but not a complete exit criterion. Validate cluster alarms, HA and DRS state, host sensors, uplinks, storage paths, time, logging, backup integration, monitoring, and representative application transactions. Watch the environment after the window for delayed hardware, driver, or performance symptoms.
Close with an updated desired-state record, exceptions, actual timings, failed checks, and follow-up owners. That evidence improves the next cluster rather than making every maintenance window a fresh discovery exercise.