We had a disaster recovery plan.
The production platform was already designed for high availability, operating across multiple Availability Zones within its primary cloud region.
That protected the service from a number of infrastructure failure scenarios without requiring a full disaster recovery event.
We also had backups, and restoration from those backups had been tested successfully.
But there was another scenario we wanted to examine.
What happens if the primary region itself is no longer available?
For that, our recovery strategy relied on establishing the required capability in a secondary region and restoring from backup.
On paper, the pieces were there.
But we wanted evidence that they would work together.
So we scheduled an exercise.
The Plan Was to Test Regional Recovery
The scenario was deliberately more severe than the failures our normal high-availability architecture was designed to handle.
Assume the entire primary region was unavailable.
Could we establish the platform in the secondary region, restore the required data and services, and return the business service to operation within our recovery objectives?
This wasn’t happening during an incident.
It was a planned exercise intended to test the recovery process end to end.
We already knew that individual parts of the recovery strategy worked.
Now we wanted to know whether the recovery capability worked as a whole.
And that distinction became important very quickly.
We Didn’t Get as Far as the Restore
Early in the exercise, we encountered a dependency in the recovery process that wasn’t available in the secondary region.
The deployment capability required to provision part of the recovered environment had itself been established within the environment we were assuming was unavailable.
That meant we couldn’t continue with the practical recovery in the way the procedure expected.
This wasn’t a question of whether a backup could be restored.
We had already tested that.
It raised a broader question:
Could all of the capabilities required for recovery themselves survive the disaster scenario?
That was something we needed to understand before the recovery exercise could continue as originally planned.
We Changed the Exercise
We could have worked around the problem simply to continue with the technical restore.
But that would have weakened the purpose of the exercise.
The scenario said the primary region was unavailable.
Anything our recovery depended on had to be considered from that starting point.
So rather than treating the exercise as complete, we changed approach.
We continued as a tabletop exercise and walked through the recovery process step by step.
If the primary region really had disappeared, what would we do next?
What would we need?
Where would it come from?
What depended on something else being available first?
And how would we know when the platform was genuinely ready to operate again?
Those questions moved the discussion beyond individual infrastructure components.
We started looking at recovery as a complete system.
Recovery Wasn’t One Thing
Infrastructure needed to be available.
Data needed to be restored.
Services needed to start in the correct sequence.
Dependencies needed to be reachable.
But even if all of those things happened successfully, we weren’t necessarily finished.
Related systems needed to represent a state we could trust.
Activity occurring outside the recovered state needed to be understood.
And before returning the platform to service, we needed confidence that the technology and the business state agreed.
Our mental model of recovery therefore started to change.
It wasn’t simply:
Restore → Start → Done
It was closer to:
Restore → Reconcile → Validate → Resume
That difference matters.
A technically running platform isn’t necessarily a recovered business service.
Then We Looked at the Recovery Objectives Differently
We had defined recovery objectives.
They gave us targets for how much data loss the business could tolerate and how quickly the service needed to return.
Those targets were important.
But the exercise made us ask a different question:
What evidence do we have that we can achieve them?
A documented recovery objective doesn’t demonstrate recovery capability.
Neither does having backups.
Neither does operating across multiple Availability Zones.
Neither does having a secondary region.
Each addresses part of the resilience and recovery strategy.
But none independently proves that the complete service can be recovered within the required objective.
That requires exercising the capability as a whole.
The Exercise Hadn’t Failed
We hadn’t completed the practical recovery exercise we originally planned.
It would have been easy to describe that as a failed DR test.
We came away thinking about it differently.
The exercise had done exactly what we needed it to do.
It had exposed a dependency.
It had challenged assumptions in the recovery process.
And by continuing through the scenario, it had made us think about recovery as an end-to-end business capability rather than a collection of infrastructure components.
Most importantly, it gave us a clearer understanding of what still needed to be proven.
That was more useful than completing an exercise that simply confirmed what we already believed.
What Changed for Us
We became more careful about the distinction between resilience, restore capability and recoverability.
A highly available architecture can protect against many infrastructure failures.
Tested backups can demonstrate that data can be restored.
A disaster recovery plan can describe how the organisation intends to respond to a larger failure.
All three matter.
But the question we ultimately need to answer is different:
Can we recover the service when the failure we’re designing for actually happens?
The only way to build confidence in that answer is to exercise the capability and follow the evidence.
Sometimes that evidence confirms what you expected.
Sometimes it changes the questions you need to ask.
And sometimes the most valuable outcome of a disaster recovery exercise isn’t proving that everything works.
It’s discovering what you still need to prove.