← All articles

The Plan That Lived in a Binder

A thick three-ring binder labeled "Disaster Recovery Plan" sitting on a dusty shelf, a shaft of light across its spine

On the morning of June 6, 1944, the invasion of Normandy did not go according to plan. It never does. The weather forecast was wrong. The Channel was rougher than the models predicted. Paratroopers from the 82nd and 101st Airborne were scattered across the French countryside, some of them miles from their drop zones. Landing craft put men ashore at Omaha in the wrong sectors, into fire the naval bombardment was supposed to have cleared. Radios failed. Units couldn't find their officers. Officers couldn't find their units. The operational plan, two years in the drafting, ran out of instructions in the first hour.

What carried the invasion wasn't the plan. It was that thousands of officers and NCOs had spent those two years doing the planning. They had walked mock beaches, rehearsed the sequencing, argued over contingencies, war-gamed the alternatives. When the document fell apart in the surf, they already had the reflexes the document was supposed to substitute for.

Thirteen years later, Eisenhower put it into a single sentence at the National Defense Executive Reserve Conference: "Plans are worthless, but planning is everything."

Every mid-market company I've walked into has it exactly backward. They have a disaster recovery plan. Somewhere. It's a PDF in SharePoint, or a three-ring binder on a shelf in a server closet, or a Confluence page that hasn't been edited since 2022. It was written by a consultant, approved by the CIO, filed away, and considered done. When I ask who has rehearsed it, the room goes quiet. When I ask who has read it in the last twelve months, it usually gets quieter.

A disaster recovery plan isn't a document. It's a rehearsed capability. The document is the artifact that falls out of the process, not the other way around. Companies that build the document first almost always end up with a plan that doesn't survive contact with reality, because it was written from the top down, by the people who know the systems, for an audience that will never use it.

The plan gets written in the wrong order

The most common mistake is letting IT draft the plan alone. IT knows what systems exist. They don't always know which ones matter, and they almost never know what an hour of downtime on each one actually costs the business. So the plan ends up organized around infrastructure, the file server, the email environment, the CRM, the backup tier, instead of around revenue.

A real DR plan starts with a Business Impact Analysis, run with finance and operations at the table. The question isn't "what servers do we have" but "what does this company fail to do if system X is down, and for how long, before somebody important gets very unhappy." The answers surprise people. The system IT has been spending the most money protecting is rarely the one that tops the list. The system nobody worries about is frequently mission-critical because the controller uses it to run payroll.

If you can't put a dollar-per-hour number next to every critical system, you aren't ready to write a DR plan. You're ready for a different conversation first.

Objectives are commitments, not aspirations

Recovery Time Objective and Recovery Point Objective sound like technical terms. They aren't. They're business commitments with a cost curve attached. A 15-minute RPO costs an order of magnitude more than a 24-hour RPO. A one-hour RTO costs more than a same-day RTO. Which one you buy is a business decision, driven by what the BIA told you about the cost of downtime.

Most of the DR plans I see set objectives after the tools have been chosen. That means the plan describes what was already bought, instead of what the business actually needs. The contract promises a 15-minute RPO and the infrastructure can't deliver it. Nobody finds out until a Friday at 4 PM.

The dependency graph includes what you don't own

The plan doesn't cover only your servers. It covers your DNS provider, your SSO vendor, your payment processor, your email host, your CDN, the SaaS tools your customer service team lives inside. When Dyn went down in October 2016, or when Fastly dropped a third of the internet in June 2021, companies that considered themselves multi-region found out their definition of "region" didn't include their vendors' regions.

An outage at a vendor you didn't map is identical to an outage in your own rack. You just have fewer options, because their recovery timeline is not under your control. A plan that doesn't enumerate third-party dependencies, and doesn't have a fallback position for the critical ones, isn't finished.

Write the runbooks for the humans, not the auditors

A DR plan written for compliance reads like a policy document, paragraphs of passive-voice prose about roles and responsibilities. A DR plan written for use reads like a pilot's emergency checklist: short, sequenced, unambiguous, testable. If X, do Y. If Y fails, escalate to Z. Contact A at the number on page one. Phrase the customer email as follows.

If the incident commander at 2 AM can't execute the plan without calling the author, the plan is already broken. The test isn't whether the binder is thorough. The test is whether somebody who wasn't in the room when it was written can follow it under pressure.

Rehearse. Tabletop first, live second.

Tabletop exercises find the gaps in the thinking. A conference room, a scenario, ninety minutes, and the team walking through what they would do. You don't need working infrastructure. You need the questions: who calls whom, what decisions get made, who has authority to spend money, who talks to customers, who talks to regulators. The gaps in the thinking are always more interesting than the gaps in the infrastructure, and they're free to find.

Live drills find the gaps in the infrastructure. An actual restore, an actual failover, an actual phone tree. These are more expensive and more disruptive, which is why most companies skip them. Skipping them doesn't mean the gaps aren't there. It means the gaps get discovered during the real event, when the cost of finding them is denominated in hours of downtime instead of a planned Saturday morning.

The bar is quarterly tabletops and annual live drills. The floor is annual tabletops. Never is what most companies do.

The plan rots the day it's signed

Vendors change. Staff turn over. Systems get replaced. Contracts get renegotiated. A DR plan that hasn't been edited in a year is almost certainly wrong in a dozen small ways, and a handful of those will matter during an incident. The phone number that doesn't work anymore. The contact who left the company. The tool that was decommissioned. The cloud account nobody has the credentials for.

The plan needs an owner, a review cadence, and a version history. Put it under version control. Schedule a quarterly walk-through. Treat it like code, because that's what it is, instructions that have to execute correctly the first time, every time, under conditions nobody has control over.

How Lanovix handles this

Every managed IT engagement we take on starts with a Business Impact Analysis, not a tools audit. We sit down with the business side and build the cost-of-downtime picture first, before we touch infrastructure. That conversation usually reveals priorities that nobody had written down, and it sets the budget for everything that follows.

From there we build the plan in layers. Objectives, grounded in what the business actually needs and can afford. Dependency maps, including the third-party vendors and the out-of-band access paths. Runbooks written for use, not for audit. A named incident commander and a named backup, with authority defined in advance. And a rehearsal schedule, on the calendar, with owners, so the plan gets exercised before it gets tested.

We also version the plan and own the maintenance. When a vendor changes, when a system gets replaced, when a key person leaves, the plan gets updated inside the same change window. The binder doesn't end up on a shelf, because we don't maintain a binder. We maintain a living document, and we run it regularly with the client's team so that the thinking, not just the paperwork, stays current.

Eisenhower's point wasn't that plans don't matter. It was that the plan alone isn't the capability, the practice is. The company that survives the outage isn't the one with the thickest binder. It's the one whose people have done the planning often enough that when the binder is wrong, which it always will be, they already know what to do.

If your disaster recovery plan is a document you haven't touched in a year, get in touch. We'll help you make it something you've actually run.

Share LinkedIn