Skip to main content
Corevia Technologie
Infrastructure & Cloud

Disaster recovery: why most recovery plans fail on the day

An RTO negotiated in a meeting remains an intention until the first timed drill. What separates the plan from reality: restore arithmetic, backup immutability, forgotten dependencies, and the drill as the only admissible proof.

7 min readThe Corevia team

The plan exists. It is signed, filed, cited in tender responses, and it promises recovery within four hours. On the day it has to be used, recovery takes three days. In nearly every case the technology held. The plan described an objective, whereas recovery is a sequence of interdependent operations whose real end-to-end duration nobody had ever measured.

RTO and RPO are objectives, not observations

RTO is the maximum acceptable outage duration for a service; RPO is the amount of data you accept losing, expressed in time. Both have counterparts rarely found in the documents: the duration and the data loss actually observed during a drill. Until those observations exist, the RTO remains an intention. The vocabulary of ISO 22301 adds two further useful notions, the maximum tolerable period of disruption and the minimum acceptable level of service during a crisis, which force an answer to a question often avoided: what can the business run on in degraded mode, and for how long?

The real duration is a sum of serial timings: detecting the event, deciding that it is a disaster, mobilising people, rebuilding the foundation, restoring the directory, restoring the data, validating applications with the business, reopening the service to users. The most underestimated link in that chain is the decision itself. Who declares the disaster, above what threshold, with what authority, and on the basis of what information? If declaration takes four hours because it climbs three management levels on a Sunday, no technology will meet a two-hour RTO.

Then comes an arithmetic that is almost systematically skipped. Restoring 20 TB at a sustained effective throughput of 400 MB/s takes roughly fourteen hours of pure transfer, before any validation. And that effective throughput is rarely the datasheet figure: rehydrating deduplicated data, reading the catalogue, random access on the target and the number of permitted parallel jobs often cut it in half. Sizing an RTO on theoretical throughput mechanically produces an error of between two and fivefold. The only defensible figure is the one you have measured on your own infrastructure, with your own volume.

Immutable backups and the 3-2-1-1-0 rule

The threat model has changed. The reference disaster is no longer a server room fire but the simultaneous encryption of production and backup, because the backup server belonged to the same domain as the rest of the estate and a service account held elevated privileges. Attackers look for the backup console before encrypting anything: it is what determines whether the victim will pay. Any recovery architecture designed before that shift needs re-examining, even if it performs perfectly in a technical drill.

  • Three copies of the data, including production.
  • Two different kinds of media, so as not to depend on a single failure mode.
  • One copy off-site, beyond the reach of a local disaster.
  • One copy offline, disconnected or made immutable, unreachable from the production domain.
  • Zero errors: restorability is verified automatically, and a failed verification opens an incident instead of adding one more warning in the console.

Real immutability holds up against a legitimate administrator, which no checkbox in a console guarantees on its own. Object lock in storage, hardened repository, write-once media: the mechanism matters less than the guarantee, namely that no account, whatever its privilege level, can shorten retention or delete a restore point during the locked period. The identity plane must be separated too, with dedicated local accounts, multi-factor authentication distinct from production and no domain join.

The drill is the only proof

Drills come in degrees: document review, tabletop exercise with decision-makers, restoring an isolated component, a full restore into a segregated environment, a real failover of a service with its users. Each level answers a different question and none replaces the next. ISO 22301 requires an exercise programme and periodic evaluation of documentation and capabilities, because an untested plan ages faster than the infrastructure it describes.

A drill must produce a record: a timestamped chronology of operations, the recovery time and data loss actually observed against the stated objectives, the list of blocking points, the decisions taken, and corrective actions with an owner and a date. A drill that runs without a single difficulty is closer to a demonstration than to a test. The organisations that improve are those willing to play uncomfortable scenarios: the key person is unreachable, the primary site cannot be accessed, the most recent backup is corrupted.

The recurring mistakes

  1. 1Forgotten dependencies. The directory and identity provider are restored before anything else, and their recovery procedure is specific. Then come DNS, address distribution, the certificate authority, licence servers, secret vaults, the multi-factor authentication service, time synchronisation and monitoring itself.
  2. 2Circular dependency. The password manager hosted on the infrastructure being restored, the recovery plan stored on the encrypted file share, the procedure reachable only through the intranet: provide an out-of-band copy and a tested emergency access path.
  3. 3Everything is critical. When every application is ranked top priority, none really is, and the recovery order gets decided under pressure. Prioritisation must be arbitrated and validated by the business, since IT cannot set it alone.
  4. 4Outdated runbooks. Hostnames, addresses, versions and contacts drift within months. A runbook is revised at every architecture change and verified at every drill.
  5. 5Monitoring backup instead of restore. A successful job indicator says nothing about restorability. And only yesterday’s restore point ever gets tested, almost never the oldest point still under retention.
  6. 6No out-of-band communication. If email and telephony rest on the stricken system, the crisis team cannot even convene. Keep a printed contact list and an independent channel.
  7. 7Notification obligations missing from the plan. Regulatory deadlines run during recovery: 72 hours for a personal data breach, 24 hours for an early warning under NIS2 where the entity is in scope. The crisis team needs to have them in mind from the opening hours.
  8. 8Snapshots mistaken for backups. A snapshot kept in the same account, subscription or region as production disappears along with it, and a compromised administrator deletes it within minutes.

What the cloud shifts

The shared responsibility model is often read backwards: the provider’s commitment covers the availability of its infrastructure, and stops where the recoverability of your data begins, after a deletion or an encryption event. The retention offered by a software-as-a-service platform protects against mishandling for a few weeks, rarely against an administrative compromise, so it does not play the role of a backup. A copy in a separate account, with separate credentials, remains the rule. Finally, in infrastructure described as code, the repository and the integration pipeline are critical dependencies too: restoring a cloud platform means replaying code and then reinjecting data, which is to say running a software project against the clock.

  • Disaster recovery
  • Business continuity
  • Backup
  • Resilience
Back to insights

A project, an audit, an emergency?

Describe your situation in a few lines. We come back within one business day with an initial read and the questions that matter.

contact@coreviatechnologie.com