Figuring out the cause of an IT failure is more complex than the problem itself

Opinion
Sep 6, 20054 mins

* Getting snowed under by a blizzard of failure alerts

Solving problems in complex IT environments is a time-consuming, people-intensive operation. In fact, most IT organizations do a very poor job when it comes to identifying and resolving such issues as why an IT or business process fails.

Solving problems in complex IT environments is a time-consuming, people-intensive operation. In fact, most IT organizations do a very poor job when it comes to identifying and resolving such issues as why an IT or business process fails.

And yet fail they do. And at some sites, they fail far too frequently for anybody’s liking. Just about every site has some sort of alerting system associated with applications or IT processes that send out a distress call when a process fails. The sad truth is that most such systems contribute very little to solving the problem, and in fact in some cases they may even contribute to making the situation worse. Here’s why:

Typically when a failure occurs the support team gets snowed under by a “blizzard” of alerts. This blizzard however, does not give much guidance as to the source of the fault because in almost all cases the information the alert supplies is simply a listing of symptoms. Unfortunately, this list of symptoms is accompanied by zero indication of what has caused them.

In a moderately complex IT infrastructure, many parts of the overall system may be affected by a single fault and each, when it encounters a problem, spawns an alert. In a moderately large shop, it is not unusual for a single problem to generate hundreds of alerts, each describing a symptom but none, alas identifying the problem itself or showing at which point along the data path the problem had taken place.

In such cases, all the support team knows is that a fault has occurred and a process (or perhaps a great number of processes) has ended prior to completion. It is at this point that the various subject matter experts go out and examine almost every issue described in an alert. Eventually, they usually do find the trouble, but at what cost?  Let’s consider a bit of what is involved.

Repairing a problem often occupies numerous technical personnel who likely as not will first turn to topology maps, Visio diagrams, or Excel spreadsheets. As none of these is capable of being updated in real time, they frequently find themselves referring to outdated documents. This is of course an even greater problem when what is being managed is systems spread out over a wide area.

How do you economically dispatch personnel to go out to physically examine the potential problem points?  The answer of course is that you don’t. Such an approach is clearly a time consuming and an inefficient use of staff and other resources.

Worse yet, while the problem is waiting to be identified and solved, the IT team typically has no way to understand which business processes are being hit. This quite naturally tends to strain relations between IT managers and their peers on the business side of the house. Lacking knowledge of what is going on, IT management can’t inform line of business managers how long their revenue systems may be down.  Service level agreements (SLAs) are violated, a fire drill ensues, and the corporate ship is not a happy one.

The most appealing answer, quite candidly, is to retire to a beach house in the Caribbean. As for those of us who can’t do this (shame on all of us for having lacked the foresight to be born independently wealthy), well we really should pay greater attention when someone eventually does come up with a tool that provides root cause analysis of our problems.