Creating Safe Nomad Releases with Canary Deployments and Rollbacks
HashiCorp Nomad can release a new application version gradually instead of replacing every running allocation at once. Its deployment settings allow operators to start a small number of canary allocations, observe their health, and then promote or revert the deployment.
This approach is useful for teams running services across Australian cloud regions such as AWS Sydney or Melbourne. It reduces the risk of taking an entire production service offline when a container image, configuration value, or application dependency behaves differently under real traffic.
A reliable canary process combines Nomad’s update stanza with service health checks, deployment monitoring, and an external load balancer or service mesh. The scheduler manages allocations, while the traffic layer determines how users reach the new version.
Prepare The Nomad Job
A canary deployment begins with a normal service job. The important difference is the update block, which defines how many allocations Nomad may replace, how many canaries to create, and what should happen when health checks fail.
The example below uses three application allocations and creates one canary during each deployment. The datacenter names are examples for an Australian environment; use the names configured in the Nomad server and client agents.
job "payments-api" {
datacenters = ["syd1", "mel1"]
type = "service"
update {
max_parallel = 1
canary = 1
auto_promote = false
auto_revert = true
health_check = "checks"
min_healthy_time = "30s"
healthy_deadline = "5m"
progress_deadline = "10m"
stagger = "30s"
}
group "api" {
count = 3
network {
port "http" {
to = 8080
}
}
service {
name = "payments-api"
port = "http"
check {
name = "HTTP readiness"
type = "http"
path = "/ready"
interval = "10s"
timeout = "2s"
}
}
task "server" {
driver = "docker"
config {
image = "registry.example.com/payments-api:2025.03.1"
ports = ["http"]
}
resources {
cpu = 500
memory = 512
}
}
}
}
Understand Canary Allocation Behaviour
When the image or job specification changes, Nomad compares the new version with the currently deployed version. With canary = 1, it starts one new allocation while the existing allocations continue running. max_parallel = 1 limits further replacement activity, which keeps the release controlled.
auto_promote = false leaves the deployment waiting for an operator or automation pipeline to approve it. This is valuable for production services that handle payments, personal information, or workloads subject to the Australian Privacy Act and Australian Privacy Principles.
A canary allocation is not automatically a percentage-based traffic split. Nomad schedules the allocation and reports its health, but traffic distribution depends on Consul, Fabio, Traefik, an API gateway, or another load-balancing layer. Configure that layer deliberately if the canary must receive only selected requests.
Add Meaningful Health Checks
A process-level check only proves that a container is listening. A useful readiness endpoint should verify that the application can serve requests and has completed startup tasks such as loading configuration, opening required connections, and applying compatible database migrations.
The /ready endpoint should return a successful status only when the application is ready for service. Keep liveness and readiness separate where possible. A temporary dependency failure may make an allocation unready, while restarting the container repeatedly could make an incident worse.
min_healthy_time prevents Nomad from treating a briefly successful allocation as stable. healthy_deadline limits how long an allocation has to become healthy. progress_deadline controls how long the deployment may make no meaningful progress before Nomad marks it failed and, with auto_revert = true, returns to the previous version.
Deploy And Inspect The Release
Submit a changed job with the Nomad CLI:
nomad job plan payments-api.nomad
nomad job run payments-api.nomad
nomad deployment list
nomad deployment status <deployment-id>
The plan command identifies changed resources before scheduling begins. After submission, deployment status shows healthy, unhealthy, running, and placed allocations. Check both Nomad events and application logs rather than relying on allocation health alone.
For teams operating across Sydney and Melbourne, compare results in both datacenters. A release may work in one availability zone while failing in another because of regional networking, secrets, firewall rules, or a different dependency endpoint. Schedule observation periods using AEST or AEDT consistently so operators in Perth, Brisbane, and eastern states interpret deployment windows correctly.
Promote A Successful Canary
After validating logs, latency, error rates, resource consumption, and business transactions, promote the canary:
nomad deployment promote <deployment-id>
Promotion tells Nomad that the canary is accepted. The scheduler can then replace the remaining old allocations according to max_parallel. Continue monitoring during this phase because problems may appear only when the new version handles more connections or a wider range of requests.
Automation can promote a deployment after checks from Prometheus, Datadog, CloudWatch, or an internal release controller. Keep approval gates for high-impact services. In regulated Australian environments, retaining the job specification, deployment ID, approval record, and health evidence supports operational reviews and controls such as APRA CPS 234.
Trigger And Test A Rollback
If the canary fails, leave the deployment unpromoted or mark it failed:
nomad deployment fail <deployment-id>
With automatic reversion enabled, Nomad uses the previous stable job version. Verify the result with:
nomad job status -verbose payments-api
nomad deployment list
For a deliberate rollback to a known job version, inspect the job history and revert to the selected version:
nomad job history payments-api
nomad job revert payments-api <version>
Test rollback before relying on it in production. Confirm that the previous container image remains available, secrets and configuration are backward-compatible, and database changes do not prevent the old application from starting. A migration that removes columns or changes data formats may make an otherwise correct Nomad rollback ineffective.
Keep deployment events and logs for the period required by your organisation’s security and compliance policies. A rollback is most useful when operators can quickly identify the failed version, affected allocations, observed symptoms, and exact recovery action.