The restore succeeded. Then someone tried to sign in.
The database is available in the recovery environment. The restore job succeeded, the application servers started, and the team is preparing to declare the exercise complete. Then someone tries to sign in.
The application cannot reach its identity provider. A recovery account lacks access to a required secret. A background worker still points to a production endpoint. The data has returned, but the customer cannot use the service.
This is an illustrative recovery exercise, not a customer incident. It exposes a problem that backup-level reporting alone cannot resolve: a recovered component is only one part of a working business service. The remaining dependencies need a recovery path, an accountable owner, and a test that demonstrates their role in the outcome.
Recovery dependency mapping makes those relationships explicit before an incident. It gives the team a way to turn “the service must recover” into a sequence it can execute and evaluate.
What is recovery dependency mapping?
Recovery dependency mapping connects a business service to the resources, access, configuration, people, and external systems required to restore and use it. Each relationship should explain what must be available, where it will come from, who owns it, and how readiness will be checked.
A production architecture diagram is a useful starting point. The recovery map asks additional questions. Can the team access the backup platform if the normal identity system is unavailable? Are encryption keys usable from the recovery environment? Can the application resolve its dependencies through the recovery network? Which steps require a person whose normal communications channel may also be affected?
Forttic’s guide to ResOps starts with critical services and their dependencies. The practical next step is to make each required relationship testable. A map that names a database without explaining how the recovered application reaches it is still incomplete for that purpose.
Begin with one business journey
Choose a service whose disruption has a clear business consequence. Then describe one journey that demonstrates its minimum acceptable operation. “The customer can authenticate, view an order, and submit an update” is more useful than “the application is up.” It gives the team observable behavior to test.
Agree which functions can be temporarily unavailable. A service may support essential transactions while recommendations or analytics remain offline. That can be an acceptable degraded state if the business has approved it in advance. Record its limits and duration so an improvised workaround does not silently become the recovery standard.
Keep recovery time and data-loss requirements attached to that journey. Starting the database within its target does not establish that the customer journey met the service’s recovery-time objective. Likewise, a recent database copy does not establish consistent application behavior if another required data store was recovered to an incompatible point.
The service owner should approve the acceptance criteria, with engineering explaining how they will be demonstrated. This creates a shared definition of success before teams start optimizing individual restore steps.
Map both the service and the recovery process
There are two useful views of the dependency map. The first describes what the service needs to function. The second describes what responders need to recover it. They overlap, but treating them separately helps expose circular dependencies.
For example, a recovery runbook stored behind the affected identity service may be inaccessible when needed. An orchestration workflow may rely on a secrets store located only in the failed environment. The team needs a tested way to obtain the required access without bypassing its security controls.
Use a compact record for each dependency. The following is an illustrative starting point, not an exhaustive architecture standard.
| Dependency | Recovery question | Evidence to seek |
|---|---|---|
| Identity and access | Can operators recover the service, and can intended users authenticate? | Tested operator access and application sign-in under the chosen scenario |
| Data and keys | Can the application read the selected recovery points and reconcile required data? | Restore results, authorized key access, and scoped integrity checks |
| Network and naming | Can recovered components find and reach each other? | DNS, routing, and connectivity checks from the recovery environment |
| Configuration and secrets | Do endpoints, certificates, and credentials match the recovery design? | Versioned configuration and successful access to required secrets |
| Queues and integrations | Can work resume without unintended duplication or external side effects? | Controlled message processing and integration test results |
| Recovery tooling | Can responders reach runbooks, artifacts, and orchestration tools? | Access and execution checks using the planned recovery route |
Attach an owner and a fallback decision to each unresolved dependency. Some external services cannot be restored by your team. Their map entries should identify the dependency, the expected provider behavior, and the service’s agreed response if it is unavailable.
Sequence recovery around readiness conditions
A list of resources does not explain their startup order. Describe the condition each recovery step needs before it can proceed. “Database process started” and “database accepts authenticated application requests” are different conditions. The latter is a more useful prerequisite for starting a dependent application.
Some tasks can run concurrently; others must wait. Microsoft’s Azure Site Recovery documentation describes recovery groups and ordered startup, with parallel execution within groups. That is one product-specific example of making dependencies executable. Your sequence still needs to reflect the actual application, including checks beyond machine startup.
Look for cycles. If two components each require the other during initialization, the runbook needs a tested bootstrap procedure. If a service depends on a shared platform, agree how competing recovery requests will be prioritized. Otherwise, each application team may assume the shared dependency will be available first.
Avoid treating the map as a promise of one universal sequence. Regional loss, identity compromise, and accidental deletion can leave different resources available. Preserve the scenario assumptions so operators can tell when a procedure applies.
Test the customer journey after the restore
Separate resource restoration, application readiness, and business acceptance in the exercise record. These checkpoints answer different questions, and a failure at one should remain visible even if earlier checkpoints pass.
AWS Backup, for example, supports a validation workflow after a restore-testing job completes. Completion can trigger a separate check whose result is recorded. The organization still defines what that check proves; a reachable health endpoint alone need not establish a functioning customer journey.
For the illustrative order service, use controlled test identities and data to sign in, retrieve the expected record, submit an update, and confirm the downstream result. Keep the test environment from sending real payments, messages, or customer notifications. If a dependency is simulated, record that substitution and the behavior it leaves untested.
Retain the scenario, environment, configuration version, chosen recovery points, timestamps, checks, and exceptions. A useful result might state that order viewing and editing passed in an isolated environment while the payment integration remained simulated. That gives the next reviewer a precise claim to evaluate.
Keep evidence relevant as the service changes
A passed exercise is evidence about a particular system under particular conditions. When those conditions change, identify which part of the evidence needs reconsideration.
A new authentication provider may affect sign-in and operator access. A changed encryption-key policy may affect restore permissions. A new queue may change processing order or introduce another data-consistency question. These changes should route to the person responsible for the affected recovery relationship.
The response can be proportional. A targeted permission check may resolve one uncertainty; a material change to the service’s dependency chain may justify a broader recovery exercise. The decision should explain why the selected check is sufficient for the change being assessed.
This extends the discipline in Forttic’s article on backup coverage after cloud changes: protection and recovery assumptions need attention as the environment evolves. Preserve earlier test results as historical evidence while clearly identifying any current scope that awaits revalidation.
Make unresolved dependencies part of the operating review
Give the service an accountable recovery owner while retaining specialist ownership of individual tasks. The service owner ensures the acceptance criteria remain clear and unresolved dependencies progress. They do not need unrestricted access to every system.
A useful review starts with the last tested journey and asks what changed afterward. Discuss dependencies that lack a tested recovery route, scenarios that rely on unverified assumptions, and failures that remain open. Count verified journeys within a defined scope if that helps, but keep material exceptions beside the count.
When a test reveals a gap, define the evidence needed to close it. Updating a connection string can complete the engineering task; a subsequent successful application connection establishes the relevant outcome. Forttic’s verified-closure guide explains how to preserve that distinction in the remediation record.
Where continuous enforcement fits
Forttic’s CRE framework connects discovery, assessment, governed remediation, verification, and evidence across connected backup and cloud systems. Within a service-oriented ResOps practice, that loop helps address protection conditions that drift from the required state.
Business acceptance still needs a service-specific definition. A generic control check cannot decide whether an order workflow is acceptable for your customers. Use the protection evidence and the application test results together, with explicit scope for each.
Start with one service and one journey. Build the dependency record, test the planned recovery route, and bring the first unresolved relationship to an operations review. You will leave with a clearer next action than a general request to improve disaster recovery.
Take the Forttic CRE assessment to examine how discovery, corrective action, verification, and reporting connect in your current operating process.
A restored database is not a recovered service
A restore job can succeed while identity, secrets, networking, or a downstream integration still block the customer journey. Map those dependencies, test the journey, and keep the evidence scoped to what was actually proven.
Frequently asked questions
What is recovery dependency mapping?
Recovery dependency mapping identifies the resources, access, configuration, people, and external services required to restore and operate a business service. It also records prerequisites, ownership, and evidence needed to validate the recovery route.
Why can a database restore succeed while the application remains unavailable?
The application may still lack working identity, networking, secrets, configuration, keys, or downstream services. A database restore demonstrates one component’s recovery; application and business-journey checks establish the broader outcome.
What is the difference between an infrastructure recovery test and a service recovery test?
An infrastructure test establishes whether selected technical resources can be recovered. A service recovery test evaluates defined business behavior across the required dependencies, such as a user signing in and completing a controlled transaction.
Does every system change require a full recovery exercise?
No. Assess which recovery assumptions the change affects. Use targeted checks where they adequately address the uncertainty, and broader exercises when the change materially affects the service recovery path. Record the reasoning and remaining limitations.
Who should own the dependency map?
A named service or recovery owner should maintain accountability for the overall outcome. Application, identity, cloud, network, and backup teams should own their dependency records and recovery tasks, with a clear route for unresolved gaps.
Does a passing recovery test prove continuous recoverability?
A passing test provides evidence for its recorded scope, scenario, and time. Ongoing monitoring, change assessment, and appropriate revalidation are needed to maintain confidence as the environment changes.