Why do most organizations fail when it comes to DRP?

Analysis
Dec 10, 20083 mins

Over the past couple of years, many companies have attempted to plan for and implement a disaster recovery plan (DRP).  In fact, most of our clients insist upon have architecture designs take DRP into account.  And, thanks to these practices, “most organizations” will attempt to deploy mission critical business functions in a redundant manner.

But, is just embracing DRP from a technology stand-point truly enough to sustain an organization’s operations in the event of a disaster?  In today’s economic state, filled with corporate layoffs, banking collapses, and god knows what else.  I often pose this type of question to clients trying to plan their infrastructure to meet any possible future events or needs.

Interestingly enough, most organizations do not see the point of such a question.  After all, in their mind, they have addressed DRP from an architecture stand-point, and in some cases even drafted an actual DRP document that specifies resources, procedures, contact points, and such.

However, what I’m really trying to draw attention to is the lack of planning for day-to-day operations in relation to knowledge retention, role rotation, and (yes) resource allocation.  It may sound odd, but these items are often just not fully considered or deeply pondered by most organizations.  Instead, “people” resources are often viewed as disposable or transient to meet business plans that are often designed to follow a fleeting market direction or driven by perceived cost reductions.

To define what this picture actually looks like… I’ve seen origination after organization that have an “IT guy” or number of people (PC statement) which have deep and extensive knowledge about a particular system or system(s).  In other words, they are often a key cog in a complex widget which is designed to keep the entire business going.  More than often, their knowledge and skills tend to be silo’ed and without these individuals present and working 24/7 an organization’s business operations might possibly be impacted.

In other words, these individuals are also a single point of failure in that their role is not redundant, the knowledge is not shared, and they are often so busy fighting fires they are also poorly allocated resources.  Making this scenario even more complicated is that an organization’s management layer often does not fully understand what these individuals do, the technologies that they support, nor depth to which their loss would impact business.

Naturally, when framed in this light… you should be able to easily picture the direction my thought stream is going.  Yes, what tends to happen either ends up in disaster or costing a bit of money.  But, this also seems to be a systemic problem in IT.  How can it be addressed?  Not really sure…