← All success stories
Cloud Solutions
Canadian Group of Companies

A Recovery Plan the Team Can Run

A Canadian group of companies had Azure DR tools across two sites. BITSUMMIT assessed the gaps, designed the hybrid recovery architecture, and left their IT team with runbooks and training to run it.

Industry
Facility Management
Service line
Cloud Solutions
Timeline
~6 weeks
A Recovery Plan the Team Can Run
10
servers
Critical servers covered by the DR design and runbooks
2
sites
Brought into one hybrid Azure DR architecture
4
phases
Assessment, design, runbooks and knowledge transfer

The challenge

The tools were in place. A recovery the team could run was not.

The client is a Canadian group of companies with operations across two sites and a hybrid footprint spanning on-premises infrastructure and Microsoft Azure. Azure Backup, Azure Site Recovery, and Commvault were already in place.

What was missing was confidence those pieces added up to a recovery the team could carry out. Recovery time and recovery point objectives for critical systems had not been validated with the business. Procedures were not standardized. Gaps sat unmapped and unprioritized.

They needed a practical DR design for 10 critical servers across both sites, and procedures their own people could follow — not a full enterprise business continuity program.

Our approach

An untested DR plan is a hope, not a plan. The work focused on two things: a design grounded in how the environment actually behaves, and runbooks written for the people who will use them. BITSUMMIT delivered it in four phases, each reviewed before the next began.

1. Discovery and assessment

Reviewed both sites and Azure: inventory, application dependencies, Azure Backup and Commvault configuration, and Azure Site Recovery replication for the 10 servers in scope. Output: a prioritized gap analysis, and recovery objectives validated with stakeholders instead of assumed.

2. Azure DR architecture

Designed a hybrid DR architecture on Azure Site Recovery and Azure Backup for the 10 prioritized servers, including:

  • Replication approach and policies
  • On-premises-to-Azure and Azure-to-Azure failover paths (ExpressRoute, VPN, DNS)
  • Break-glass and role-based recovery access
  • Backup and storage recommendations that keep Commvault in place
  • Sizing, cost projections, and a phased implementation roadmap

3. Runbooks

Runbooks for planned and unplanned failover, failback, re-protection, and restore — from a single item to a full VM — plus RACI, escalation and communication templates, and tabletop exercise templates the team can run on its own schedule.

4. Knowledge transfer

Walked IT through architecture and runbooks, briefed leadership on findings and roadmap priorities, handed over the full documentation package.

Platforms in scope

  • Azure Site Recovery: replication and orchestrated failover for the critical servers
  • Azure Backup: backup and restoration in Azure, including storage redundancy choices
  • Commvault: the client's existing backup platform, factored into the design rather than replaced
  • ExpressRoute: hybrid connectivity for the DR network paths between the sites and Azure

What this did not include

This phase covered assessment, design, and runbooks. Building out the design and running live failover tests against production were scoped as a later phase. The runbooks and tabletop templates are written to support that first exercise — testing is where a plan becomes a capability.

The outcome

The client left with a recovery design their team understands and procedures written before they are needed.

  • A team trained to run recovery itself — a recovery event does not depend on BITSUMMIT being in the room
  • Documented Azure DR architecture for 10 critical servers across two sites, on the platforms they already run
  • Prioritized gap analysis so the next spend hits the highest-risk gaps first
  • Standardized runbooks for failover, failback, re-protection, and restore, with roles and escalation paths
  • Phased roadmap for implementation and live failover testing, with sizing and cost guidance

That is resilience without the chaos: design the team understands, procedures written down before they are needed, people already trained.

If DR sits inside a wider platform change, see how we build Azure Backup and Site Recovery into VMware-to-Hyper-V migration from the start.

Not sure how your recovery plan would hold up under a real outage? Book a DR readiness review — 30 minutes with a senior engineer on objectives, runbooks, and testing gaps.

Your story could be the next one.

Tell us what you're trying to modernize, secure or migrate. We'll bring a plan and a named senior engineer.

Schedule a call →