Troubleshooting and recovering corrupted VMware VMFS datastores
VMFS corruption can stop a vSphere environment in its tracks, and the administrators who manage it often need answers within minutes. Whether you run a private cloud in a Sydney colocation facility or a smaller deployment in a regional office in Geelong, the underlying mechanics look the same. Recovery options depend on recognising what has failed and acting before the damage spreads.
A damaged datastore rarely affects a single host. Shared storage such as iSCSI LUNs or Fibre Channel volumes feed multiple ESXi servers, and a single corrupted metadata block can take dozens of workloads offline. Australian organisations in finance, healthcare and higher education rely on consistent uptime, and even a brief outage can interrupt research at Monash, halt retail systems in Melbourne, or delay online services across the APAC region.
Recovery is rarely a single click. It usually requires a measured walk through diagnostics, controlled changes and verified restoration from backup. This article lays out the practical sequence Bryan and Dan have used in production, including the commands, log locations and decision points that matter when a datastore is failing.
The goal is not to encourage heroic fixes but to provide a clear path. A structured approach reduces the chance of compounding the problem and gives you a defensible record of the steps — something auditors and service desk leads in Canberra and Adelaide will appreciate.
Root causes of VMFS datastore corruption
VMFS corruption rarely appears out of nowhere. It is usually a downstream symptom of a fault higher in the stack: a storage controller issue, a fibre channel switch reboot, an HBA firmware mismatch, or a sudden power loss during a write. Each can leave the on-disk metadata — heartbeat, file allocation tables and resource directory — inconsistent.
Local conditions also play a part. Australian data centres, including the NextDC campuses in Sydney and Melbourne and Macquarie facilities in the Sydney CBD, run hot aisles at high density. Component wear in those environments can accelerate the failure rate of HBAs and SSDs, particularly in tier-2 arrays. Heat-related HBA faults have been observed to scramble outstanding I/O, which ESXi then interprets as datastore corruption.
Operator error is another common trigger. A LUN being extended online, an accidental unmap during storage vMotion, or a misconfigured storage multipathing policy can all cause transient inconsistency. Sometimes a vendor migration tool writes a new signature over the VMFS header. Identifying the trigger is the first step in choosing the right recovery method.
Warning signs worth investigating early
The earliest indicators are subtle. Storage I/O latency climbs on a single host well before the datastore becomes inaccessible, and vCenter alarms may start firing for "Datastore usage threshold" or "Storage path redundancy lost" without an obvious cause. VMs may report "I/O error" inside the guest, or fail to power on with a message about a locked virtual disk.
A short list of signals that should always trigger a deeper look includes:
- Sudden, unexplained latency spikes above 50 ms on a previously stable LUN
- Hosts losing access to a datastore while storage remains reachable from the array side
- "Lost access to volume" entries in the vmkernel log
- vSphere HA heartbeats failing because the heartbeat directory cannot be read
Any one of these in isolation might be a false positive, but a combination often points directly at VMFS metadata trouble. A quarterly review in your monitoring platform — Veeam ONE, PRTG or a Grafana stack — tends to catch issues before the datastore is fully lost.
Diagnostic steps before attempting recovery
The first rule of datastore recovery is to make the situation no worse. Snapshot the LUN at the storage layer if your array supports it, then start gathering evidence. The esxcli storage core path list and esxcli storage vmfs extent list commands give you a quick view of what the host sees, and vmkfstools -V reports the state of any mounted VMFS volume.
The vmkernel and hostd logs under /var/log/ are essential. Search for "NMP", "FS3.0", "Heartbeat" and "Corruption" to see whether the kernel detected an inconsistency or simply lost contact. Cross-referencing these timestamps with the array's event log tells you whether the fault lives on the host or the SAN. Australian shops running Dell PowerStore, Pure FlashArray or NetApp AFF usually have vendor contracts that can pull dumps quickly — use them.
If the datastore is unmountable but the LUN is still presented, try esxcli storage filesystem rescan followed by a controlled vmfs resignature of the affected volume. This rewrites the metadata pointer but leaves the user data blocks intact, which recovers most logical corruption. Physical damage to the metadata region is rarer but more serious, and usually requires restoring from a storage-level snapshot or backup.
Recovery procedures that actually work
When the datastore is still partially accessible, the safest path is to evacuate the running VMs through vMotion to a healthy datastore, then repair the volume offline. Unmount the datastore, run vmkfstools --fix if the version supports it, and attempt a remount. ESXi 7.0u3 and 8.0 environments typically repair the journal cleanly.
For more severe cases, a structured recovery sequence looks like this:
- Export a full list of affected VMs and their VMDK locations from vCenter
- Take a storage-level snapshot of the LUN if the array allows it
- Resignature the VMFS volume and remount, restoring any corrupt VMDKs from backup
When the metadata cannot be rebuilt, the last resort is to recreate an empty VMFS on the LUN and restore everything from backup. This is faster and more predictable than manually patching a half-corrupted filesystem. Australian IT teams usually schedule these full restores during the Sunday maintenance window, which lines up with the AEST/AEDT timezone and keeps the change outside business hours.
Preventing the next incident
Recovery is expensive, prevention is cheaper. Keeping HBA firmware current, applying storage vendor patches promptly, and running with at least two independent paths to every LUN removes most of the failure modes that lead to corruption. RRDM (Round Robin with path ranking) is usually a safer default than Fixed for shared workloads.
Backups still matter. A consistent snapshot taken by Veeam, Commvault or Nakivo, replicated offsite to a secondary site or to AWS S3 in the Sydney region, gives you an out when other options fail. Test the restore annually and document the procedure so a junior admin in Perth or a regional office can follow it without paging the team at 2am.