Backups and recovery: what companies think vs. what they can restore

A deposit is not security. Without recovery, it's just a save file. In infrastructure design, it is essential to know what we can restore and how quickly.

A functional backup is measured by the ability to recover, not by "the job running". You need defined RPO/RTO, separate access and regularly tested recovery procedures - for cloud and on-prem.

Key questions

  • What exactly are we backing up (data, configurations, infra state, secrets)?
  • What can we restore automatically and what is manual improvisation?
  • What is the RPO/RTO for each service?
  • When was the last time we tested recovery and with what result?
  • Where are access rights and recovery keys stored?

The most common reality: there is a backup job, but the restore has never been attempted. In the event of an incident, this will result in loss of time and uncertainty.

A good infrastructure design defines a backup policy (what is backed up, how often and for how long) and, above all, a recovery procedure (steps, responsibilities, testing).

It is true for cloud and on-prem: a recovery test is cheaper than an incident without recovery.

Typical risks in practice

  • Backup without recovery: backups exist, but recovery has never been really tested.
  • Shared access: the backup system uses the same accounts as production.
  • Unbacked Identities: IAM, AD or cloud accounts are not part of the recovery.
  • Missing runbook: the incident is improvised, the responsibilities are not clear.
  • A false sense of security: "backup is green" is confused with the ability to restore.

Most of these problems are not technological. These are missing decisions, ownership and testing.

Related:

Do you need help?

If you need to design a backup and recovery strategy for your infrastructure, get in touch with us – we will advise you how to set it up correctly.

Contact WOV Tech