* SAN modeling tools help you understand systems changes prior to implementation
endif; ?>Sometimes we have to reach the end of a run of good fortune to appreciate fully how lucky we have been. In Las Vegas, it’s is a localized, personal problem (unless you are gambling with the company’s funds), but when it happens to a production system it can have severe impact for an IT manager.
Most of us realize that even when an enterprise is operating normally, it is only a matter of time before some component will halt unexpectedly. Sometimes it’s the 50-cent fuse that breaks to save the machine. At other times, Murphy’s Law asserts itself and the $100,000 storage system dies to save the fuse.
Such disruptions occur because, hiding just beneath the surface, lurks a system based on patches, recently added assets from a company your boss just acquired, on-the-fly fixes, and a long list of workarounds. Fix one part of the system and, unless you have the time and resource to test it thoroughly, you have no idea what the fix’s impact may be on other, seemingly unrelated pieces of the infrastructure.
Companies often have some idea of how to marshal resources in response to a single department’s needs – if a particular application stops performing normally, they know what to do. But they rarely exhibit much understanding of how their response may have affected their entire IT operation from top to bottom.
As a result, almost anything can happen. How much downtime results from having implemented a “fix” to another system is anybody’s guess (who would admit to such a thing?), but anecdotal evidence says the number is substantial. Unfortunately, lacking a proper test environment and the necessary time to use it, what may happen is often anybody’s guess.
Calculating the second-order effects (the “ripple” effects) of an outage event – and the effects caused by its remediation – can be challenging, but doing so is absolutely necessary if you want to understand downtime costs in any complex environment. After all, if one of the results of your fix is to bring down an important production system, you may wind up impacting the company’s revenue-generating operations.
The issue of operational downtime thus moves beyond simple IT importance to being a measure of a corporation’s capacity to do business. Seen in this light, downtime has strategic significance for the entire enterprise.
All of this contributes to the feeling of unease that goes home with many IT execs when they leave the office – they can sense a problem is brewing somewhere in their management substructure, but they don’t truly know what the time and dollar impact will be when something finally does hit the fan. And most IT managers do not consider bad things happening on the IT floor to be a career-enhancing situation.
An increasing number of vendors now offer some nice storage-area network modeling tools that enable you to see what happens as you play mix-and-match with your SAN’s contents, bringing systems and data paths online and offline. These enable you to do “what if?” analyses prior to implementing changes, and to see the potential effects across the system. Some may prove to be a very good investment.
A number of tools are available for general SAN validation and design. Most also provide availability and interoperability information, and allow you to use “what if” scenarios that reflect local storage management policies as you bring devices online and offline.
Not all such tools offer equal value, and not all will necessarily be appropriate for you. In many cases, for instance, they lack the ability to look beyond a proprietary set of products, so if the tool can’t discover and analyze your SAN elements just pass it by. Increasingly, of course, the inability to auto-discover SAN elements almost borders on the inexcusable.
If this is of interest to you, you might want to check out the following currently available products: CA’s BrightStor SAN Designer, EMC’s ControlCenter San Advisor, and Onaro’s SANscreen.
Strategic issues can’t be solved with piecemeal approaches. Tools such as these can inspect and provide analysis across a whole system, and should provide a wider view of what is likely to happen when you change something.




