* Event automation and real-time network systems
endif; ?>Once upon a time, in a galaxy close to here, system uptime was pretty much dependent on your answer to two questions: how good were your back-up and recovery services, and how athletic were your operators?
Good backups often meant that you could recover from an error caused by corrupted data, assuming that the backed up data was actually restorable. Operator athleticism was not only important when it came time to dash across the IT room to reset devices after a machine check or a hang, but was also a crucial part of the recovery process (run to the back room, identify the correct rack full of tapes, find the right tape, grab it and run it out to the tape drive, slap it on the drive, reboot if necessary, run back to the office to send out a message to the users, get back to your coffee, etc.).
Nowadays back-up and recovery tools and procedures are typically better than they ever were, and – due to improved hardware, software and automation – operators get much less exercise (tell me that’s not obvious in most computer rooms).
Also however, we have new categories of products to keep our systems running.
While several sorts of products now help ensure better system uptime, one category – sometimes referred to as “event automation and real-time network systems” – is becoming increasingly important.
If we look beyond that obviously awkward and not particularly descriptive name, we see a potential for a set of solutions that can help us manage across the entirety of the IT infrastructure, thereby allowing storage managers to understand the impact of network and server events on storage performance. This is the technology that, in the truest sense, will enable IT managers to monitor and manage across the entirety of their IT system.
The companies that play in this segment build software that combines alarm and event management capabilities with rules-based correlation, application impact analysis, and root-cause analysis. This combination enables IT staff to identify quickly and efficiently the devices and device interactions that cause performance and availability issues. Once intelligence about the local computing environment has been discovered, the software lets managers identify applications impacted by a device that is generating faults or alarms. In the more sophisticated applications, managers are also informed which applications will be affected by a service action planned for a given device.
This field was pioneered in the 1980s by Belcore, the centralized R&D group that provided operations support systems (OSS) and other technical services to the Bell operating companies. When I worked there in the 1990s, the company’s root cause analysis software was considered one of the corporate crown jewels, but like most jewelry in that category, it was extremely expensive and was only affordable for very large enterprises. And of course, it pretty much only looked at telephony-related issues.
Now new candidates have taken up the mantel, and a new generation of products shows both advanced technology and greatly improved affordability. Offerings from CentrePath (http://www.centrepath.com) and SMARTS (http://www.smarts.com, but remember that EMC is in the process of acquiring this company) are already seeing significant user acceptance. As more IT centers move to incorporate information lifecycle management, utility computing and distributed environments, the services provided by these vendors (and presumably, others as well) will become central to the support of the new generation of inherently complex IT environments.
Watch this space for more on this topic during the coming months.




