Skip to main content
  • Cloud
  • Migration
  • DevOps

What a zero-downtime migration actually costs you

Zero downtime is achievable on almost any system. The question worth asking is whether the price of achieving it is one your programme should pay.

Priya Nandakumar2 min read

"Zero downtime" sits in most migration briefs as though it were free. It is not. It is a constraint you buy, and the currency is parallel infrastructure, duplicated write paths, and engineering hours spent on machinery you will throw away the week after cutover.

What you are actually paying for

A true parallel-run migration means both systems are live simultaneously. Writes go to both. Reads are compared. Discrepancies are investigated rather than ignored. For the duration, you are running two production estates and a reconciliation process between them.

On one recent engagement that period lasted seven weeks. The infrastructure cost roughly doubled for those weeks, and two engineers did nothing but triage comparison failures. That was the right call — the workload processed regulated transactions and a bad hour would have been a reportable event. It would have been the wrong call for an internal reporting tool.

The question to ask instead

Rather than asking whether you can achieve zero downtime, ask what a two-hour maintenance window on a Sunday morning actually costs. For a meaningful share of systems, the honest answer is close to nothing, and a scheduled window removes most of the complexity above.

We now ask clients to price the window before we design the migration. The conversation is short and it reliably changes the plan.

When the answer really is zero

Some systems earn it: payment rails, clinical systems, anything with a regulatory uptime commitment, and anything where the cutover window would land in a different timezone's business hours. For those, phase the migration workload by workload rather than attempting one heroic weekend. Each phase proves itself against live traffic before the old path is retired, and each phase is independently reversible.

The failure mode we see most often is not a botched cutover. It is a team that committed to zero downtime for the whole estate, ran out of appetite somewhere around workload nine of fourteen, and finished the job with exactly the maintenance window they could have scheduled at the start.

Let's build something that ships

Tell us what you're working on and we'll come back with a practical plan, not a sales deck.