Almost every company has backups. Far fewer know how long it would take them to get back to business if the entire environment went down tomorrow. That gap between having copies and being able to recover is what separates backup from cyber resilience.
Ransomware has made that gap obvious. Modern attackers are not content with encrypting production servers: they go after the backup console, delete snapshots and try to encrypt the repositories themselves before launching the attack. Anyone who finds out on incident day that the backups were reachable with the same compromised admin account finds out too late.
Backup is not resilience
In SP 800-160 Vol. 2 Rev. 1, NIST defines cyber resiliency as the ability to anticipate, withstand, recover from and adapt to adverse conditions, attacks or compromises. Backups serve one of those four goals well — recover — and only if they are built to survive the attack itself.
In practice, a backup that exists but does not restore, or restores in three weeks when the business can only tolerate three days, does not deliver resilience. Restoring complex infrastructure is rarely as simple as it looks: dependencies between systems, startup order, credentials, licenses, DNS and identity services tend to surface only during a test — or a crisis. So the right question is not "do we have backups?" but "how long will it take, and how much data will we lose, before we are operating again?"
The other goals depend on controls that reduce the chance of an attack ever reaching the backups: phishing awareness, patching operating systems and software, multifactor authentication and email filtering. CISA's #StopRansomware Guide lists these same measures as baseline prevention. For small and midsize companies without a dedicated security team, relying on a specialized partner is often the realistic way to keep these controls and the monitoring of the environment in place.
RTO and RPO: the numbers the business must define
Two metrics, described in NIST SP 800-34 Rev. 1 and also used in ISO 22301, capture that question:
- RTO (Recovery Time Objective). The maximum time a system can be unavailable before the business impact becomes unacceptable.
- RPO (Recovery Point Objective). The point in time to which data must be recovered — in practice, how much data loss is tolerable. A four-hour RPO requires copies at least every four hours.
These values are not an IT decision. They come from the business impact analysis (BIA), which identifies critical processes, their supporting systems and the maximum tolerable downtime. Without a BIA, backup frequency and recovery architecture are chosen by habit rather than need — and it is common to discover that the billing system has the same backup policy as a forgotten file server.
3-2-1 and 3-2-1-1-0: the rule and its update
The 3-2-1 rule is the classic redundancy reference: three copies of the data, on two different types of media or storage, with one copy offsite. It protects well against hardware failure, human error and physical disasters.
Against ransomware, it falls short if every copy is reachable over the network with the same credentials. Hence the 3-2-1-1-0 variant, popularized by the industry:
- +1 offline, isolated or immutable copy. A copy the attacker cannot change or delete, even with administrator privileges in the environment.
- 0 errors on verification. Backups checked automatically and restores tested, with no outstanding failures.
CISA's #StopRansomware Guide points the same way: maintain offline, encrypted backups of critical data and regularly test their availability and integrity in a disaster recovery scenario.
Immutable, offline, isolated: what changes in practice
The terms are often used interchangeably, but they are not the same:
- Immutable. The storage prevents changes or deletion during a retention period (for example, object lock on object storage or repositories with a WORM policy). Logical protection.
- Offline (air gap). The copy has no network connection to the production environment — tape out of the library, disconnected media. Physical protection.
- Isolated (vault). A separate environment with its own network, identity and administration, accessed in a controlled way.
Whatever the choice, three safeguards always apply: backup credentials separate from the corporate domain and protected with MFA; monitoring of mass deletions, retention changes and job failures, with alerts treated as security events; and encryption of the copies, so that a repository leak does not become a data leak.
A clean copy matters as much as an available one. Attackers often stay in the environment for a while before acting. Restoring a backup that already contains the intruder's access restarts the incident. Retention has to cover windows long enough to return to a point before the compromise, and restores should be verified before going back into production.
Restore testing is the most neglected control
A job that completed successfully is not the same as recoverable data. A mature testing program includes:
- Automatic verification. Integrity of the copies checked on every run, with an alert on failure.
- Periodic sample restores. Files, databases and virtual machines restored in a controlled environment, confirming that the application actually works.
- Full recovery test. At least for critical systems, simulate total loss and measure the real time to operation — then compare it with the defined RTO.
- Exercises with the business. Tabletop exercises in which managers make decisions under an outage scenario, as recommended by NIST SP 800-34 and ISO 22301.
Each test should lead to adjustments: updated runbooks, documented dependencies, a revised restore priority. Classifying and segmenting sensitive data helps set that order — what comes back first is what supports critical processes.
From backup to a continuity plan
Backup is a technical component; continuity is an organizational capability. ISO 22301 treats business continuity as a management system, with roles, policies, BIA, strategies, plans, exercises and continual improvement. NIST SP 800-34 details the information system contingency plan, with activation, recovery and reconstitution phases.
In a ransomware incident, the recovery plan also has to work with incident response: before restoring, the attack must be contained, the entry point understood and the attacker confirmed to be out of the environment. Likewise, continuous detection by a SOC over the backup infrastructure increases the chance of spotting a sabotage attempt before it succeeds.
Frequently asked questions
Does a cloud backup count as an offsite copy?
It counts as a copy in another location, but not necessarily as a protected copy. If the cloud is reachable with the same credentials as the environment and allows unrestricted deletion, an attacker with privileges can wipe it. Enable immutability, separate identities and monitor retention changes.
What is the difference between RTO and RPO?
RTO measures how long a system can be down; RPO measures how much data loss is tolerable. A system can have a long RTO and a short RPO (it can take a while to come back, but cannot lose transactions), or the other way around.
How often should I test restores?
There is no single number. Integrity checks should be automatic and continuous; sample restores, periodic; and full tests of critical systems should follow at least a planned cycle and always happen after significant infrastructure changes. Frequency should be proportional to the criticality defined in the BIA.
Do immutable backups solve the ransomware problem?
They solve part of it: they guarantee there will be a recoverable copy. They do not prevent data exfiltration, which now accompanies many attacks, nor do they replace containment, eradication and checking that the copy is clean.
References: NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems (BIA, RTO, RPO and contingency planning); NIST SP 800-160 Vol. 2 Rev. 1, Developing Cyber-Resilient Systems (resiliency goals: anticipate, withstand, recover, adapt); CISA, #StopRansomware Guide (offline, encrypted and tested backups); ISO 22301:2019, Security and resilience — Business continuity management systems — Requirements. The 3-2-1-1-0 rule is an industry convention, not a normative standard. This article is educational and does not describe a specific product.