Configuring vSphere Replication for RPO-Based Disaster Recovery
Australian organisations operating across Sydney, Melbourne, Brisbane, and Perth face a familiar reality: critical workloads are rarely housed in a single metro area. With data sovereignty obligations, the Australian Cyber Security Centre's Essential Eight guidance, and APRA CPS 234 mandates for financial entities, the bar for recovery assurance sits higher than ever. Pair that with the geographic distances between capital cities, where a Sydney-to-Perth link can balloon to 60 milliseconds of latency, and the case for a well-tuned replication strategy becomes hard to ignore.
vSphere Replication offers a hypervisor-level mechanism to ship virtual disk deltas from one vCenter environment to another, without requiring shared storage. Because the engine runs as a virtual appliance embedded in the management stack, it integrates cleanly with existing ESXi clusters and avoids the licensing overhead of array-based replication. Administrators can tailor the Recovery Point Objective on a per-VM basis, which makes it well-suited to mixed environments where some workloads are weekly and others demand sub-15-minute RPOs.
The catch is that the default configuration is rarely aligned with the recovery commitments written into business continuity policies. Mismatched RPO settings, undersized recovery networks, and forgotten MPIT retention windows have ended more DR drills in Australia than any hardware failure ever did. The remainder of this guide walks through the moving pieces so a vSphere Replication deployment lines up with the RPO the business expects.
Preparing the Sites and the Replication Appliance
Before touching any RPO numbers, both the protected and recovery sites need a clean foundation. Each vCenter must be reachable on TCP 8043 and 8044 for VRMS traffic, and the recovery site's ESXi hosts require outbound access to VRMS on TCP 8043 during initial pairing. In a typical Sydney primary / Melbourne secondary layout, firewall rules often need to be carved between the corporate WAN edge and the management VLAN, since replication traffic should never share a saturated business subnet.
Deploy the vSphere Replication appliance OVF onto the recovery vCenter and register it as an extension. During deployment, allocate at least 12 GB of RAM and provision a thick-provisioned disk for the embedded PostgreSQL database, particularly when larger fleets of replicated machines are planned. Once registered, log into the VRMS interface at https://<vrms>:8043, pair the protected vCenter using its SSO credentials, and confirm that both sites appear as healthy.
A common oversight is assuming the appliance can be installed and forgotten. The VRMS database holds replication metadata, recovery point indexes, and bookmark information. Without a backup of this appliance, a full re-seed of every protected VM becomes necessary after a rebuild, which defeats the original point of computing it.
Mapping RPO Settings to Business Workloads
The RPO field in a replication configuration is where business intent meets network reality. A 5-minute RPO for a SQL cluster means roughly five minutes of delta data to capture, compress, and acknowledge across the WAN. When the underlying link between, say, a Brisbane production cluster and a Canberra recovery site is provisioned at 50 Mbps with 25 ms RTT, that 5-minute window will frequently slip into "behind" status.
Treat RPO as a tiered commitment. Group VMs by criticality: Tier 1 for transactional databases and identity providers, Tier 2 for line-of-business applications, Tier 3 for dev and test. Assign RPOs of 5, 15, and 60 minutes respectively. This zoning approach prevents an over-eager Tier 3 workload from consuming all available bandwidth on a constrained link, a scenario that has been characterised in Australian Telecom-related reshoring guides as a frequent cause of replication backlog.
Enable network compression on every protected VM. Compression alone can reduce outbound traffic by 40-60%, which has measurable impact on monthly cross-connect costs for organisations paying per-megabit transit between data centres. Pair compression with scheduled replication windows for non-critical workloads so that Tier 1 traffic always has headroom.
Network Optimisation and Bandwidth Planning
Bandwidth throttling is the most under-used lever in vSphere Replication. The appliance supports QoS-style limits measured in Mbps, which lets administrators cap a noisy VM without affecting its neighbours. As a rule of thumb, allocate no more than 70% of the inter-site WAN capacity to replication traffic, leaving room for live VMotion migrations, backup windows, and user-driven cross-site activity.
For organisations connecting Sydney and Perth, the long distance plus limited carrier routes means that WAN optimisation appliances from vendors such as Riverbed or Silver Peak remain relevant. vSphere Replication speaks standard TCP, so any device that accelerates long-haul TCP sessions will shrink replication windows. Where SD-WAN has been deployed across branches feeding back to a central vCenter, replication traffic should be pinned to a dedicated underlay class of service.
Retention policy is the silent RPO multiplier. Multiple Point in Time recovery relies on the appliance keeping hourly, hourly, and daily snapshots for the configured duration. A 24-hour retention means 24 recovery points to ship, store, and maintain. Australian compliance regimes that demand multi-day recovery testing often push retention beyond the 24-hour default, which in turn drives storage planning on the recovery side.
Operational Hygiene for Production
Operational checks are critical. Monitor replication health through the vSphere Client dashboard on a daily schedule, and configure SNMP forwarding to email alerts to a distribution list that an Australian-focused care manager owns. Lag indicators above the configured RPO should trigger a pager, not a morning email. Pair that with weekly drills that promote a single VM at the recovery site and validate application boot, which is how Australian Practitioners Advisory Council recommendations translate into operational practice.
Documentation of the runbook matters as much as the configuration. Record the sequence to pause replication, the network commands to verify reachability, and the recovery steps that reviewers like me or my peers would follow during a real event. Store the document in a system that lives outside both vCenters, since a regional outage could wipe out the very environment that holds the recovery procedure.
Practical Recommendations
Build a configuration baseline that survives staff turnover and audit scrutiny.
- Define three RPO tiers and apply them consistently across every protected workload, with justification tied to business impact analysis
- Enable compression on all replicated VMs and cap replication traffic at 70% of available inter-site capacity
- Back up the vSphere Replication appliance database as part of the standard nightly backup cycle, with offsite storage in a different region
- Schedule quarterly recovery drills that include DNS failover, application-level validation, and a documented RTO measurement
- Review firmware version policy on the VRMS appliance quarterly to stay aligned with VMware compatibility and Australian Signals Directorate hardening guidance