An analyst tells why understanding the value of corporate data must come before your no-fail business-continuity plan can soar.
An analyst tells why understanding the value of corporate data must come before your no-fail business-continuity plan can soar.
|
IT managers today are caught in an endless three-way tug of war when it comes to storing and protecting corporate data. Pulling them one way is their legal accountability for protecting intellectual property. Pulling them another way is their need to prove IT effectiveness by meeting performance metrics laid out in service-level agreements. The increasing need to justify how every dollar invested in IT assets supports business processes pulls them in a third direction.
The importance of data protection – and of managing that data once it has been protected – has set in motion several storage initiatives. One of the most interesting is the concept of the information life cycle, which assumes that the value of data (and of the information it provides) will change according to a generally predictable set of circumstances. Because the circumstances are known, managing the data proactively throughout each stage of this life cycle becomes IT’s job, taking into account any application and legal requirements.
The term information life-cycle management (ILM) might be relatively new, but the concept has been around for many years. Decades ago, even the largest storage devices were small compared with what we expect to see today. Mainframe vendors, under pressure from IT managers facing space constraints on their primary (and very expensive) storage devices, began to develop methods for migrating data from expensive disk drives such as Direct Access Storage Devices (DASD) to lower performance (but much less expensive) tape drives and, eventually, to optical media. IT managers viewed storage as a hierarchy, with fast/expensive devices at the top and slower/less-expensive devices at the bottom. And so the process of migrating data from DASD to tape became known as hierarchical storage management (HSM).
Typically, IT managers manually initiated this data migration to less-costly devices. The migrations ran as batch jobs, and first-in first-out operations through which older data was “demoted” to the lower performing devices to make way for newer data irrespective of its actual value. Managers looked for high watermarks, points at which it appeared that disk utilization was in danger of being maxed out. When this happened, they usually backed up or archived everything on a disk to tape. They managed relatively little, however, and as a result, tape media companies flourished, whole facilities were given over to tape storage, and the time to re-acquire data via a tape restore often challenged even the most patient IT managers.
Nowadays, ILM focuses on understanding the value of the reference data that supports business needs, and on providing policies for proactively managing that data in accordance with its value. This correlation is important. It means the management system moves data to devices that provide levels of support that are appropriate to the data’s value and use. High-value data that must be accessed frequently and that demands high availability goes on high-performance devices, and gets access to high-performance backup and recovery services. Data of lower value (or high-value data that does not require high-speed access) is allocated to slower, less-expensive devices.
One difference between ILM and HSM is that, with the ILM system, data is tracked and maintained at an availability level defined by IT management. So if a certain dataset requires availability within 30 minutes under any circumstances, an ILM policy will prohibit its transfer to devices that cannot provide this level of service.
|
Yet another advantage of migrating data to a policy-defined set of storage hardware is increased optimization of the overall storage system. Because IT can establish policies that make sure lower value data does not take up expensive space on high-performance devices, it receives the double benefit that comes with optimizing disk usage in favor of data value. IT can make sure high-performance storage space is available for data that deserves it, while deferring investments in additional high-performance storage because it is more efficiently allocating the assets currently in operation.
Also inherent within ILM is the ability to track data throughout its life cycle, and to make an audit trail available to whomever might need it. Today, the “whomever” might be a corporate officer, a stockholder or, with increasing frequency, a federal or industry regulatory body that has established formal rules regarding data retention and access.
Significantly, managing the data life cycle begins the moment the data takes residence on a system, and lasts until such point as the data is permanently retired. In every sense, ILM requires ongoing supervision of data and of the storage assets that protect it. In many cases, this is more than just a good idea, it is the law.
When it comes to managing data, every company has its own set of rules that gets translated into formal IT policy. Whatever those rules are, underlying them must be these basic assumptions:
Assumption 1: In an IT environment with limited resources, each piece of data must be given every bit of protection it deserves … but not one iota more. If information is a company’s most important asset, it must be given all the protection it needs. But unless a site has unlimited time, storage assets and personnel, resources allocated to one set of data – disk space, tape media, operators’ time and so forth – can be re-allocated only to other data if managers are willing to accept the effect this will have on protecting the first dataset.
In IT departments where budgets are limited but demands are not, managers know maintenance windows do not get larger just because more data needs backing up. If no more hours can be added to the day and no more assets can be provided to make the work more efficient, and if all available assets are already in use, then protection given to one dataset most likely must be withdrawn from another. As a result, IT managers must balance their resources and make tough decisions about which of multiple jobs should take precedence.
Assumption 2: All information is not of equal value to an organization. All information is not created equally, and under conditions of limited resources, it follows that data of different values cannot be afforded the same level of corporate resources for maintenance and protection. At its simplest, this means the data for the departmental football pool is less deserving of corporate resources to protect it than are the files containing the accounts receivable. At a more complex level, this might mean examining the value of data at a more detailed level. Such a situation might incorporate the assumptions that most e-mail is valuable, but that the CEO’s e-mail is the most valuable. It would be assigned the highest value and be protected accordingly.
Assumption 3: The value of information to an organization changes over time. For several reasons, data might shift in value. Understanding this change, and adapting the IT support structure to accommodate it, lies at the heart of ILM.
High-value data is entitled to one set of services, lower-value data to another. With ILM, data migrates from one set of assets to another when warranted by a change in its value. Value changes occur for three main reasons:
Predictable (cyclical) depreciation. Because of the way certain applications are used, the data associated with them depreciates at what are essentially predictable rates. The value of data associated with an accounts-payable program is time-driven. For example, it will change in value in direct relation to the way bills are paid. Data about bills to be paid in March will have lesser value in April, but will still be valuable as reference data for many business processes. Months later, the same dataset can be demoted to less-costly storage as the likelihood of the data being referenced decreases further.
Application event triggers. Certain applications lend themselves to less predictability than do others during some phases of their life cycles. Medical records, for example, get a high degree of use (and consequently, have high value) during patient visits, but might lie dormant for years between visits and in most cases are of little reference value after a patient’s death. Because of recent regulatory requirements that cause medical records to be kept for 30 years, however, their value continues (although at a much lower level) well into the future.
- Regulatory compliance issues. Within the last few years, U.S. regulatory agencies have imposed formal requirements on IT regarding data retention and accessibility. The two best-known among these are the Sarbanes-Oxley Act of 2002, which requires financial records be maintained and kept accessible for at least seven years, and The Health Insurance Portability and Accountability Act, which mandates that patient health records be kept private and secure for six years. As a result of such regulatory requirements, some kinds of data might continue to maintain a definable value for many years after it has stopped being used.
Improving business continuity with ILM
Because ILM seeks to manage data through the most efficient use of storage resources, it typically contains policies that address the issues of data protection, replication, archiving and deletion, as well as media management. Depending on local policies, a company might move around data with surprising frequency. Fortunately, all this need not complicate business continuity plans.
|
Implicit in managing all this data movement is the ability of the management software to track each step of the data’s migration. Now the same capability that provides regulatory agencies with auditing data can, when integrated with corporate disaster-recovery programs, provide the pathway to the most valid set of the most valuable data. And of course, because the ILM policies have allocated the most valuable data to the most robust storage hardware and services, the likelihood of vital data being preserved locally and remotely can be greatly increased.
Hardware and management software play vital roles when it comes to efficient ILM. Fortunately, corporations are learning to treat data availability, performance, recovery objectives and retention policies differently based on data’s business value at a given time. Vendors too have been paying attention. IT managers looking to implement ILM will find helpful products from a range of storage software and hardware vendors, as well as traditional systems management providers.
Additionally, two relatively recent categories of storage hardware offer great value potential when it comes to ILM. Along with the traditional offerings of tape, optical storage and high-performance disk, managers should consider serial-attached ATA (SATA ) disks and virtual tape as candidates for their site’s hardware mix.
SATA takes the same inexpensive parallel disk technology used in desktop computers (commonly known as Integrated Drive Electronics [IDE] or enhanced IDE) and gives it a serial interface. Some vendors have seen this as an opportunity to provide relatively inexpensive storage systems for use with data that is not business-critical. It is now available in both Just a Bunch Of Disks (JBOD) and RAID configurations, and in some cases has a much higher level of reliability than does the IDE technology we have all used for the last decade.
|
A virtual tape library (VTL) operates like a standard tape library, but is disk-based. Because VTLs use disks, they remove the opportunity for both robotic failure and media error. Also, because disk I/O in most cases is much faster than tape I/O, VTLs offer faster backup and restore speeds, which can be a big help to managers who must cope with shrinking maintenance windows and have a need for rapid restores. VTLs, some of which are available as appliances, are offered by a range of storage vendors including FalconStor and StorageTek. VTLs are likely to be particularly well-suited to distributed companies and remote locations where IT skills are at a minimum.
Which data deserves your team’s best efforts and your highest levels of service? At what point does that dataset change in value, and become less (or more) deserving of your business-continuity efforts? Understanding the value of your data through each stage of its life cycle will go a long way toward more intelligent assignment of the software, hardware and services that you need to ensure continuity of crucial business processes.
Karp is a senior analyst with Enterprise Management Associates. He can be reached at mkarp@emausa.com.




