Setting Up vSphere High Availability for Reliable Host Recovery
vSphere High Availability (HA) restarts virtual machines on surviving ESXi hosts when a host fails. Its effectiveness depends on reliable management networking, correctly selected datastores, and an isolation response that matches the behavior of the workloads.
Datastore heartbeating provides an additional communication path when HA agents lose contact over the network. It helps vCenter and the HA master distinguish between a failed host, a network partition, and a host that remains connected to shared storage. This reduces unnecessary VM restarts during transient management network problems.
The configuration is straightforward, but the defaults should not be accepted without reviewing storage design, host isolation behavior, and application dependencies. A carefully tested HA cluster is much easier to operate during an actual infrastructure failure.
How vSphere HA Detects Host Failures
HA agents running on each ESXi host exchange network heartbeats, normally through the management VMkernel interface. The cluster master monitors these heartbeats and coordinates VM restarts when a host stops responding. A host that fails completely is handled differently from one that is running but cannot reach other HA agents.
Datastore heartbeating adds a secondary signal. When network heartbeats disappear, HA checks designated shared datastores for heartbeat files written by the affected host. If the host is still updating those files, the cluster may determine that the host is alive and reachable through storage rather than immediately treating it as failed.
This mechanism requires shared storage visible to multiple hosts. It does not replace redundant management networking, and it does not protect virtual machines from every storage outage. A datastore that becomes unavailable because of APD or PDL conditions cannot provide dependable heartbeat evidence.
Prerequisites For A Stable Cluster
All ESXi hosts should run compatible versions and be managed by the same vCenter Server. Confirm that each host has a consistent management network configuration, synchronized time, DNS resolution, and access to the same production datastores. HA can function with imperfect DNS in some environments, but troubleshooting becomes significantly harder.
Use at least two management network paths where practical. Separate physical adapters, redundant switches, and correctly configured VLANs reduce the chance that a single interface or switch failure creates a false isolation event. Verify that the HA network can carry traffic between every host, including across any routing or firewall boundaries.
Shared datastores should be accessible by multiple hosts and should have sufficient free space for heartbeat files. NFS and VMFS datastores can be used for datastore heartbeating, provided they are mounted consistently. Review storage multipathing and failure behavior before relying on a datastore as an HA communication path.
Configure Datastore Heartbeating In vCenter
Open the vSphere Client, select the cluster, and choose Configure. Under Services, open vSphere Availability, edit the HA configuration, and locate the datastore for heartbeating settings. The exact menu wording can vary slightly between vSphere releases, but the options are available within the cluster’s HA configuration.
The default policy generally allows vCenter to select suitable datastores automatically. This is appropriate for many environments, but administrators can specify particular datastores when automatic selection does not reflect the storage design. Select datastores that are visible to as many hosts as possible and avoid placing every heartbeat location on one storage system.
Use more than one datastore when the cluster has multiple shared storage systems. Heartbeat redundancy is valuable only when the selected datastores have independent failure domains. If all selected datastores depend on the same array, controller, fabric, or network path, the apparent redundancy may be misleading.
| Configuration area | Recommended approach | Operational concern |
|---|---|---|
| HA network | Use redundant management paths where possible | A single switch or NIC failure can create isolation |
| Heartbeat datastores | Select multiple shared datastores across storage paths | Avoid a single array or fabric dependency |
| Datastore policy | Use automatic selection unless storage requires manual control | Review which datastores vCenter actually selects |
| Host isolation response | Choose based on workload and network design | Incorrect action can cause duplicate VM execution |
| VM restart priority | Prioritize infrastructure and critical services | Dependent applications may start out of order |
Select The Host Isolation Response
The host isolation response controls what an ESXi host does when it determines that it cannot communicate with other HA agents. The primary choices are Leave powered on, Power off, and Shut down. The setting is configured under the cluster’s vSphere Availability options and can usually be overridden for individual virtual machines.
Leave powered on avoids an automatic power action, which can be useful when storage remains available and the administrator wants to prevent unnecessary shutdowns. However, if the isolated host is still running workloads while another host restarts the same VMs, duplicate execution or split-brain behavior can occur.
Power off is generally faster and more decisive than a guest shutdown. It may be appropriate for stateless services or workloads where rapid HA restart is more important than graceful operating system shutdown. Shut down allows the guest operating system to stop cleanly, but it depends on VMware Tools and can delay recovery if the guest is unresponsive.
The correct choice depends on whether the isolated host can still access production storage and whether the network failure affects client traffic. Review VM Overrides for databases, domain controllers, clustered applications, and appliances that have their own quorum or fencing mechanisms.
Validate Failure Detection Safely
Before testing, record the current cluster configuration and identify a maintenance window. Confirm that vMotion, DRS, vCenter, storage, and application monitoring are healthy. A controlled test should target a nonproduction VM or a host carrying workloads that can tolerate an interruption.
Test management network failure separately from a complete host failure. Disconnecting a management uplink can demonstrate isolation behavior, but do not remove every communication path without a recovery plan. Monitor the Recent Tasks, HA events, hostd and vpxa logs, and VM power state changes in the vSphere Client.
After restoring connectivity, verify that the host rejoins the cluster without duplicate VMs or unexpected power operations. Check the datastore heartbeat status and review alarms for network partition, host isolation, APD, or PDL events. A successful test confirms behavior under the tested conditions, not every possible switch, storage, or routing failure.
Operational Recommendations For HA
Document the selected heartbeat datastores, network paths, isolation response, and VM-level overrides. Include storage owners and application owners in the review, because a technically valid HA action can still conflict with database replication, licensing, or application clustering.
- Keep management networking redundant across separate physical paths.
- Use multiple heartbeat datastores with independent storage connectivity.
- Match isolation responses to workload recovery and fencing requirements.
- Review HA and storage alarms after every network or SAN change.
- Re-test failover behavior after major vSphere, switch, or array upgrades.
Regular validation is especially important in hybrid environments where vSphere hosts, cloud-connected services, and external load balancers share dependency chains. HA can restart a VM, but service recovery also depends on DNS, storage, network reachability, application ordering, and monitoring.
Configure vSphere HA in a maintenance window, verify datastore heartbeat visibility from every host, and test the selected isolation response with a controlled workload. A documented and exercised design provides far more confidence than an enabled cluster using unverified defaults.